Leading the Research to Improve the Developer's Experience
Developers can be captive customers, too. In large enterprise organizations they rarely get to choose the internal tools, platforms, and processes their job hands them, any more than an employee chooses the app they use to enter their time — they just have to use them, every day, whether the experience is good or not.
Founded and led the practice
45 champions → 450 participants
3,000+ developers, 6 technology orgs
DX Core 4, Snapshot self-assessment
BUILT FROM THE BOTTOM UP
Everyone knew developers were hitting friction. Nobody could say where the worst of it was.
Division of 3,000+ developers
Multi-Family Tech
Capital Markets Tech
Single Family Tech
Tech, Products & Eng Services
Technology Infrastructure
Data Engineering & Platforms
Developers from all across CIO
DevInsights Team
Before forming the DevInsights Team, we only had access to directors, principals, and VPs. That leadership proudly described the bulk of our 3,000+ developers as “full-stack.” One population, one label, one set of assumptions.
The top-down approach wasn't producing anything we could act on. I knew we had to reach the “front-line” coders — the developers who lived, ate, and breathed code all day. Their insight wouldn't just guide our team's roadmap. It could guide an entire division of 3,000+ developers.
So I went sideways. I asked designers and project managers — the people who worked closest to developers — who they most liked working with. Then I asked those developers the same question, and kept pulling the thread until I had a pool of 45.
What I found inside the “full-stack” label was far more nuanced than leadership believed. Using active directory information, 1:1 interviews, and focus groups, I built a detailed dataset on every team member: full-stack or front-end or back-end by preference, tooling (VS Code vs. JetBrains), years at Fannie, years as a developer, assets supported, manager, team, director, and VP.
That dataset became the instrument. Forty-five developers, precisely characterized, speaking for 3,000+.
Referral sampling surfaces the visible and the well-connected. I used active directory data to check coverage across all seven technology organizations and corrected for gaps as the panel grew.
A community of passionate voices
A panel is only useful if it stays warm. We ran DevInsights in a dedicated Microsoft Teams channel with monthly meetings and standing activity cycles, and we were explicit about the deal up front:
Attend DevInsights activities and monthly meetings — roughly five hours a month
Speak up in group activities; no one is silent
Maintain access to the tools we work in: Mural, Figma, and others
Talk to fellow developers about their friction points and bring them back to the team
Cameras on during recorded sessions
Five hours a month from 45 developers bought us a standing line into how thousands of engineers actually worked. Public-product teams pay a fortune for access like that. We had it down the hall.
A research practice, not a research project
The panel's real value was that it let us match the method to the question instead of defaulting to a survey.
When a product owner needed to know whether a workflow was broken or just unfamiliar, I could pull five back-end developers on JetBrains who supported that specific asset. When we needed to understand what the division as a whole was struggling with, we ran the Snapshot.
Every activity ran the same shape: a facilitator who owned the room and an observer who watched and took notes, followed by analysis that ended in recommendations rather than findings.
Research that stops at a findings deck doesn't survive contact with an engineering roadmap.
Activities in rotation: 1:1 interviews + focus groups and empathy sessions + design reviews + usability testing + card sorting and core modeling + surveys
Planning.
Scheduling the activity, determining which developers to include, emailing their managers, determining the kind of activity to run, and other preparation tasks.
The Activity.
This is the core of the research intiative. An activity could be a design review, a series of customer interviews, a survey, usability test, card sorting exercise, etc.
Analysis.
After the activity, the team must take the time to review the feedback data. The analysis should include recommendations for product optimization.
Optimization.
Using the analysis, the team works to optimize the product based on the feedback analysis - updating the design, revising the workflow, updating the code, etc.
Activities that require direct input from customers should have a facilitator and an observer. The facilitator is the "face" of the activity and will interact directly with the customers. The observer will take notes and watch users while they interact with the facilitator.
Measuring what we couldn't see
Forty-five developers could tell us what was wrong. However, having more developers would help us pinpoint to where improvements were needed down to the unit level and yield true insights.
So we grew the panel. DevInsights went from 45 to 450 developers — roughly one in seven people in the division — recruited across all seven technology organizations rather than by referral this time, so the larger pool corrected the sampling bias the original 45 carried. The original 45 stayed on as champions: the group we pulled from for interviews, focus groups, and design reviews.
Division of 3,000+ developers
450
The New DevInsights Team
45 : Champions
The original 45 DevInsights team, now promoted to Champion
450 : Snapshot participants
Recruited all across CIO, picked from teams using AI assistants — roughly one in seven developers
3,000 : All the developers in CIO
Recruited all across CIO, picked from teams using AI assistants — roughly one in seven developers
Two tiers, two jobs. The champions gave us depth and outreach. They would rally the dev teams they worked in, listened to their pain points, jotted down their own pain points, and encouraged developers on their team to take the quarterly Snapshot survey.
The 450 gave us a survey population that represented 1 in 7 developers. Imagine having a product where you could speak to one-seventh of your customer base.
That kind of representation made this a true DevInsights team. We stood up the DX Intelligence Platform on the DX Core 4 framework and ran the division's first developer self-assessment — the Snapshot — alongside driver deep-dive surveys and product assessments.
The April Snapshot pointed hard at two things. Deep work scored low — developers weren't failing at their jobs, they were being prevented from doing them, fragmented across meetings, tools, and context switches. Build and test scored low alongside it: the loop between writing code and knowing whether it worked was slow enough to break concentration on its own.
That reframed the problem for leadership from “developer productivity” to “developer focus,” and it set the priorities for the year.
What changed
Leadership took the April findings into 2025 planning. Deep work and build-and-test became the priorities the platform organization worked against — the DevEx North Star metrics the division judged itself by, translated into OKRs and roadmap commitments.
Two interventions followed. For deep work: No Recurring Meeting Thursdays, a division-wide protected day. For build and test: GitHub Copilot in VS Code and OpenAI's Codex, rolled out at scale on the back of the pilot assessments.
The July Snapshot went back to the same 450 developers.
Intervention: No Recurring Meeting Thursdays, a division-wide protected day. Bar lengths are indicative until the real driver scores go in; the labeled delta is the measured result.
Intervention: GitHub Copilot in VS Code and OpenAI's Codex, chosen on the back of the pilot assessments. Delta stated numerically alongside bar length — the comparison never depends on the colour difference.
The improvements happened because our team made these areas a priority. One of the programs we initiated was a “No Recurring Meetings Thursday”. This meant that developers had a least one day where they could work without
The division attributes the build-and-test gain to the assistants and the assessment process that picked them, together. Deep work's gain maps to one program cleanly. But one cycle isn't a trend, and no survey wave separates a protected Thursday — or a well-targeted pilot — from everything else that happened in the same quarter. We'd want several more waves before calling either durable.
The panel as standing infrastructure
DevInsights was a research initiative, not a design one. Its output was insight, and insight is only worth what someone does with it.
What made it durable is that it stayed available. When our teams built Chassis CodeGen and DragOn, we didn't go recruit participants and wait three weeks — we already had a characterized panel and could pull the exact developers who'd use the thing. Design decisions on both products ran through developers who did that work every day.
That's the captive-user advantage. The people who depend on what you build are down the hall, and if you do the work to reach them once, they stay reachable.