Measuring What Copilot Actually Changes for Developers

When GitHub launched a technical preview of Copilot in 2021, the working hypothesis was that AI pair programming would improve developer productivity. Early users reported exactly that, but anecdotal feedback is not measurement. To understand Copilot's real effects, GitHub's researchers had to first confront a harder question: what does productivity even mean for software developers?

There is little consensus on how to measure developer productivity. Open questions remain about which metrics matter, how much weight to give self-reports, and whether the traditional output-versus-input view fits work that is fundamentally about complex problem solving. GitHub's own 2021 research suggested developers conceptualize a productive day less in terms of raw output and more as a day where they stayed focused, made meaningful progress, and felt good at the end of it. That finding aligns with broader academic work showing that satisfied developers perform better.

A Three-Part Research Approach

Because AI-assisted development is a young field, GitHub had little prior research to build on. Instead, the researchers designed their own studies around three principles:

  • Look at productivity holistically. Using the SPACE productivity framework, they selected dimensions beyond simple speed, including satisfaction and well-being.
  • Combine perception with observation. They ran both qualitative surveys for first-hand experience and quantitative experiments for measured outcomes, cross-checking the two.
  • Test realistic scenarios. Recruitment targeted professional developers, and tasks were designed to reflect everyday development work.

Productivity Extends Beyond Speed

A large-scale survey of more than 2,000 developers using Copilot revealed benefits that have little to do with velocity:

  • Job satisfaction improved. Between 60–75% of users reported feeling more fulfilled in their job, less frustrated while coding, and better able to focus on satisfying work.
  • Mental energy was conserved. Seventy-three percent said Copilot helped them stay in flow, and 87% said it preserved mental effort during repetitive tasks. Prior research shows context switches and interruptions are among the most draining aspects of development work, so this is a meaningful distinction.

Survey responses measuring dimensions of developer productivity--perceived productivity, satisfaction and well-being, and efficiency and flow--when using GitHub Copilot

The qualitative findings consistently pointed to the same conclusion: Copilot absorbing the boring, repetitive side of coding reduces cognitive load. That frees developers to spend their mental capacity on the complex, critical thinking they find more rewarding.

(With Copilot) I have to think less, and when I have to think it’s the fun stuff. It sets off a little spark that makes coding more fun and more efficient.

- Senior Software Engineer

Speed Still Matters

The survey also confirmed an expected but striking result: more than 90% of users perceived that Copilot helped them complete tasks faster, especially repetitive ones. To verify that perception with hard data, GitHub ran a controlled experiment.

The setup was designed to minimize confounds. Ninety-five professional developers were randomly split into two groups. All were familiar with JavaScript, received identical instructions, and were asked to write an HTTP server. One group used Copilot; the other did not. GitHub Classroom automatically scored submissions against a test suite for correctness and completeness.

Summary of the experiment process and results (described in following paragraph)

The results showed two distinct effects:

  • Completion rate was higher with Copilot. Seventy-eight percent of the Copilot group finished the task, versus 70% without it.
  • Speed gain was dramatic. The Copilot group finished in an average of 1 hour 11 minutes; the control group took 2 hours 41 minutes. That is a 55% reduction in time, with a 95% confidence interval of [21%, 89%] and statistical significance at P=.0017.

Further analysis is underway, including examination of heterogeneous effects and code quality differences, with academic publication planned.

What the Findings Mean

The combined evidence points to a dual benefit. Copilot supports faster completion times while also conserving developers' mental energy and allowing them to focus on more satisfying work. Engineering leaders who ran early trials are reportedly starting to weigh these factors from a perspective of holistic developer well-being rather than raw throughput.

The engineers’ satisfaction with doing edgy things and us giving them edgy tools is a factor for me. Copilot makes things more exciting.

- CTO, Large Engineering Org

GitHub's research is not happening in isolation. Independent work includes an evaluation with 24 students and Google's internal assessment of ML-enhanced code completion. The broader research community is also exploring Copilot's implications for education, security, the labor market, and developer practices and behaviors. As more studies emerge across these settings, the understanding of AI-assisted development's effects will continue to evolve.