Why “Developer Productivity” Is the Wrong Conversation

When a CFO asks whether more budget is the only lever, or whether AI is being used to cut the cost per story point, tech leaders hear a familiar subtext: prove that engineering effort is paying off. The problem is that these questions optimize for the wrong unit. Cost per story point and similar metrics live entirely in the realm of output. They say nothing about whether the output changes customer behavior, revenue, or cost structure.

Knowledge work is not factory work. In manufacturing, output is tangible, repeatable, and easy to benchmark. In knowledge work, more hours and more delivery do not reliably produce more business value. The link between output and outcome is often weak. Fixating on productivity metrics that ignore that link can backfire, pushing teams toward measurable-but-irrelevant activity rather than meaningful results.

The pressure is understandable, though. When business impact is unclear, the response is usually to ship more ideas and hope something lands. That spray-and-pray approach produces a growing estate of features and capabilities. The cost of keeping everything running climbs, and the share of budget left for new change shrinks. When tech leaders ask for more, they get asked to justify themselves in terms of impact—and too often, the organization lacks the intelligence to provide that justification.

The constructive way out is not to argue productivity is unmeasurable. It is to reorient the conversation toward business impact and to build the organizational capability to track it.

Closing the Impact Feedback Loop

Most delivery organizations lack impact-feedback loops. Work gets prioritized, built, and shipped, but the downstream consequences are rarely traced back to the original effort. Without that loop, the correlation between delivery and business results stays invisible. A CTO who wants to shift the conversation from productivity to impact needs to close that gap systematically.

That means treating impact intelligence as a discipline. It involves knowing what demand is genuinely worth taking on, what measurement debt the organization is carrying, and whether delivered work actually validated its intended impact hypothesis. Every part of that discipline is something a tech executive can push forward without waiting for a mandate from the CFO—and having a product CXO on board makes it easier still.

Start With Demand, Not Delivery

Before worrying about how work is delivered, clarify why it exists. New initiatives should be framed as investments, which makes them subject to the same scrutiny any investment deserves. That scrutiny begins with demand. A robust demand management practice tests whether a proposed effort is worth the cost, aligns with strategic intent, and goes through an explicit decision to proceed. It is the point of maximum leverage: the cheapest moment to reject or reshape an idea is before it enters the delivery pipeline.

Pay Down Measurement Debt

Just as codebases accumulate technical debt, organizations accumulate measurement debt. The longer impact-tracking instrumentation is postponed, the harder it is to reconstruct what happened. Decisions that once looked clear become harder to validate. Paying this down means investing in the plumbing that connects delivery output to downstream business outcomes.

The book Impact Intelligence (impactintel.net) goes deeper into the mechanics and provides worked examples of simple impact attribution. It also addresses the cultural friction between business stakeholders and data teams, which often blocks the adoption of such practices.

Validate Impact, Not Just Completion

A project finishing on time is not the same as an initiative succeeding. Yet most governance processes stop at delivery milestones. The missing step is impact validation: checking, after release, whether the presumed business outcome materialized. Fixing a loop here turns delivery from an end in itself into a hypothesis-testing cycle.

Equip Teams to See Impact

Delivery teams usually have little visibility into how their work translates into business results. Equipping them means giving them access to the data and context that connect their output to downstream effects. When teams can see the impact chain themselves, they can make better prioritization choices without waiting for top-down direction.

Reorienting the Executive Conversation

A tech CTO, CIO, or CDO does not need to accept productivity pressure as a given. There is a better position: learn how to talk about business impact more fluently than the person asking the productivity question. That requires actively managing the measurement and feedback loops described above. Over time, this shifts the question away from “what is the cost per story point and is AI lowering it” toward “what is the impact of this portfolio of investments?”

It is a defensive move as much as an offensive one. When work is delivered without demonstrable impact, budgets get scrutinized, and the conversation reverts to productivity micro-metrics. When the link between effort and outcome is visible, the discussion changes. For tech leaders increasingly caught between delivery pressure and impact ambiguity, building impact intelligence may be the most strategic investment available.

Proximate vs. Downstream Impact

The central idea here is Impact Intelligence: a continuous awareness of how tech and business initiatives influence key business metrics, not just the low-level metrics closest to the work. The visual tool for this is the impact network, a KPI-tree-like map of the links between factors that drive business results. Unlike a simple tree, it often forms a network, following conventions where green, red, blue, and black arrows show desirable effects, undesirable effects, rollup relationships, and expected functional impact, respectively. Solid and dashed arrows distinguish direct from inverse relationships. These are best treated as a probabilistic causal model rather than a deterministic one. The bottom row of the network—specific features or initiatives—is a temporary overlay; the KPI structure above it tends to be stable, even as the book of work changes. (This formulation is distinct from OKR cascades or impact mapping, tracing back to earlier work on Alignment Maps.)

New capabilities typically move lower-level product or service metrics first. That direct, first-order effect is proximate impact, which is easy to see and easy to take credit for. The real business value, however, shows up as downstream impact, further up the chain, where multiple factors muddy attribution. The examples below zoom into specific slices of an impact network.

Example: Customer Support Chatbot

Consider an AI chatbot meant to reduce contact center call volume without hurting customer satisfaction. The immediate win is measurable as virtual assistant capture—the number of satisfactory chatbot sessions. That is proximate impact, the typical focus of a build-measure-learn loop. But the downstream impact that matters is the resulting call savings. Proximate impact is not always a reliable leading indicator of that outcome, so it is worth mapping the full chain.

Example: Regulatory Compliance Assistant

A compliance analyst handles a case by studying it, checking relevant regulations and recent amendments, and using judgment to form a recommendation. Reviews and approvals follow, and the total Time to Decision can stretch from hours to weeks. Slow decisions hurt the business. If analysts are the constraint, an AI assistant to interpret the changing regulations is a plausible fix. The proximate impact might be measured by adoption, such as prompts per analyst per week. A better measure is the time saved per case. But the downstream impact, a real improvement in Time to Decision, only materializes if the assistant actually works and if the time to an initial recommendation was genuinely the bottleneck.

Example: Email Marketing SaaS

A SaaS email marketing vendor’s revenue depends on new subscriptions and renewals. Renewals hinge on how useful the product is, along with pricing. To drive usefulness, the product team adds features to lift email engagement—say, personalizing send times based on a recipient’s past open/click behavior. The feature reads behavioral heuristics to find each contact’s peak engagement windows, then feeds the campaign scheduler. Is success measured by Email Open Rate or Click Through Rate? An A/B test can verify those, but that is still just proximate impact. The downstream question is whether those higher engagement rates translate into more revenue for the customer and, in turn, stickier renewals.

The Idea-to-Impact Cycle

An impact network is a useful starting point—a shared visual, akin to the ubiquitous language in domain-driven design. Improving impact intelligence, though, means finding weaknesses in the full idea-to-impact cycle. Displaying it as a sequence is fine, but in practice it is iterative. Any segment can be weak, but two matter most for impact intelligence: idea selection and impact measurement & iteration. Those are the leverage points. The middle segments—execution and delivery—contribute to the outcome but not to the intelligence about it. Sloppiness at these two ends feeds the spray-and-pray approach shown earlier.

The challenge is ownership. Idea selection and impact measurement typically sit with business leaders, relationship managers, or product chiefs, not with the technology organization. Yet it often falls to the tech CXO to absorb the productivity pressure from poor business impact. Approaching the responsible leaders can feel futile: if they were able and willing to bring rigor to those steps, they might have fixed it already. In organizations where Product and Engineering grew into separate functions with their own senior leadership, this conversation can easily hit political resistance. Smaller or younger companies can design around that split, but anyone well past that decision has to contend with its current shape. For them, the leverage is in the practices the tech side controls.

Making Impact Intelligence Operational

The quickest route to better impact intelligence runs through the C-Suite Core — COO, CFO, or CEO. They approve funding, set governance, and carry the mandate to fix how investments are validated. The recommended play is to bring them the case, perhaps through a summary at a leadership offsite or a copy of the book. But if they won't engage, a blessing to try targeted reforms on your own is often enough. You are, after all, the one who lives with the consequences of the status quo.

Demand Management with Teeth

Product owns prioritization, but the rationale behind idea selection is often poorly documented. A solid business case or justification should answer every question in the Robust Demand Management Questionnaire. While a shared impact network helps, answers to that questionnaire are non-negotiable. They turn ideas into SMART ones (Specific, Measurable, Achievable, Relevant, Time-bound). Without them, ideas remain VAPID (Vague, Amorphous, Pie-in-the-sky, Irrelevant, Delayed), making it impossible to validate impact after delivery — the source of the problems shown in Figure 1.

Your lever is team bandwidth. Assert the right to allocate it only to adequately documented work — for significant efforts, not every story or bug. Set your own threshold, such as two person-weeks. Crucially, separate prioritization from scheduling. Product keeps the former; the latter already respects dependencies and team availability. It should now also require documentation. If those answers were produced during triage, ensure Engineering has access.

Adequate documentation is not the same as well-justified. Engineering isn't judging whether an idea is worth doing; it is verifying that projected benefits and timelines are plausible on paper. Product still sets priority. If Product answers “we don’t know,” that’s a fact worth recording — it shows how much capacity goes to well-informed versus ill-informed demand. One client, Travelopia, implemented an early version of this; a conference video documents their experience.

Expect pushback. Common objections include: it slows us down, it invites self-censorship, it's not agile, innovation isn't predictable, the PMO/VMO handles it, it's not collaborative, and — the most serious — “we don't have the data.”

When Data Is the Obstacle

The “we don't have the data” objection is a showstopper when true. Begin by reporting the situation: one client rated qualifying requests on a scale from inadequate to excellent, then shared monthly visibility into how much engineering effort was spent on hunches. Awareness is the first step, and it surfaces the root cause — often inadequate measurement infrastructure. Name it measurement debt so it competes for attention and funding alongside technical debt.

An organization takes on measurement debt when it implements initiatives without investing in the measurement infrastructure required to validate the benefits delivered by those initiatives.

Paying Down the Debt Yourself

Measurement debt deserves a measurement improvement program, but that needs separate budget. If the CFO won't fund it, start within your own remit. Have teams instrument application code to emit structured, impact-relevant events at meaningful points. Store them and build dashboards to validate proximate and downstream impact. Build this alongside new functionality, filling only the gaps your existing third-party analytics tools miss. A product operations team, if present, can help developers target the most important gaps. Treat it as another form of observability — the observability of business impact — and scope it to important or effort-intensive functionality first.

From Projection to Validation

With measurement in place, you can produce a projection-versus-performance report. Product may already do this; Engineering should ask to participate. If they don't, take the lead — otherwise you have no alternative when asked why delivery isn't faster. The questionnaire's answer to “by how much and in what time frame” sets a date for a proximate impact retrospective. Compare projection to actuals, learn without blame, and feed the findings back into future demand management.

A Sample Report of Proximate Impact
Feature/Initiative Metric of Proximate Impact Expected Value or Improvement Actual Value or Improvement
Customer Support AI Chatbot Average number of satisfactory chat sessions per hour during peak hours. 2350 1654
“Regu Nerd” AI Assistant Prompts per analyst per week > 20 23.5
Time to initial recommendation -30% -12%
Email Marketing: Personalized Send Times Email Open Rate 10% 4%
Click Through Ratio 10% 1%

Downstream impact validation works differently because results stem from many contributing factors, not a single initiative. In the chatbot example, call volume rose only 2.4% despite 4% customer growth — but isolating the chatbot's effect requires further analysis. That analysis, contribution analysis, attributes improvement across factors and initiatives both inside and outside Engineering. It often needs a business owner to convene participants monthly or quarterly, which may be beyond a reformist CTO's reach. Still, make sure the measurement infrastructure exists so that such a retrospective is possible if a business leader chooses to hold one.

A sample report of downstream impact
Feature/Initiative Metric of Downstream Impact Expected Improvement Observed Improvement (Unattributed) Attributed Improvement
AI Chatbot Call Volume (adjusted for business growth) -2% -1.6% ?
“Regu Nerd” AI Assistant Time to Decision -30% -5% ?
Email Marketing: Personalized Send Times MQL 7% 0.85% ?
Marketing-Attributed Revenue 5% Not Available ?

For that example, assume the 1.6 percentage-point gap (160 basis points) between call volume growth and customer-base growth breaks down as follows: your analysts attribute 60 bps to seasonality, leaving 100 bps to self-service channels. After a round of contribution analysis, you allocate credit across the chatbot and other self-service improvements. This is Simple Impact Attribution, a heuristic-based contrast to the controlled experiments data scientists might prefer but cannot always run.

Return on Projection

Strict ROI is rarely calculable at approval time. But once you conduct impact validation, you can compute the next best thing: Return on Projection (ROP) — the benefits realization ratio. If a metric was projected to improve 5% and improved 4%, ROP is 80%. That is more informative than “it was delivered correctly.”

Championing ROP with the COO or CFO gives them a basis for the next funding round. Investment decisions that rest on projections alone reward whoever inflates them most — after all, nobody checks later. The point is not perfect forecasting; product development is not deterministic. ROP disciplines demand by discouraging unrealistic projections and reduces the spray-and-pray pattern. The book covers portfolio-level aggregation.

Enlisting the Teams

Your campaign need not be a solo effort. Help delivery teams understand that software delivery — even feature adoption — does not equal business impact. A hierarchy of outcomes, with business impact at the top, clarifies which outcomes matter and which merely support higher ones. Impact intelligence verifies those links work. Teams that internalize this picture will support demand management, welcome measurement debt reductions, and begin asking Product and business leaders about the impact of what they shipped.

Answering the Pushback on Demand Management

Introducing robust demand management is the keystone of the other recommended actions, but it will likely meet resistance from those on the receiving end. Here is how to respond to the most common objections.

“We can’t slow down”

This trades accuracy for speed. Accuracy here means preparing well to achieve the desired impact, and skipping it for speed is exactly the spray-and-pray dysfunction—a scattershot approach that relies on luck rather than skill. Anything requiring skill and strategy must be learned for accuracy first and speed later. If accuracy is missing, slowing down to gain it serves business impact.

None of these actions require reducing efforts around productivity or flow. The reformist CTO does not neglect efficiency; they balance it with effectiveness, recognizing that the Classic Enterprise has over-indexed on delivery agility while neglecting business agility.

“Let’s put our house in order first”

Waiting until DORA metrics reach elite status before adopting demand management is misplaced sincerity. Multiple deploys per day mean little without impact intelligence—it is just another variant of the speed-over-accuracy fallacy. This hesitation may also signal a siloed organization, where Engineering is expected to build fast and correctly while Product handles building the right thing. But without impact intelligence, accuracy is unknown; it becomes an article of faith in the idea-triage process. If that faith has led to a feature factory, delay only makes things worse.

“It’s not agile”

The questionnaire documents the hypothesis—it does not dive deep into the solution. Agile doesn’t mean jumping out of the plane and figuring out the landing mid-air; planning and then iterating is perfectly valid. Given that engineering bandwidth is expensive and the backlog is full of competing ideas, careful shortlisting matters when first-round selection has been lax.

AI-enabled productivity may ease bandwidth constraints, but churning out more features without impact intelligence reinforces the vicious cycle. The Agile Manifesto’s preference for working software over documentation is not about skipping the rationale for building that software. Working software doesn’t always produce business impact, and the questionnaire is not a plan in the sense that the principle about responding to change refers to.

“Innovation isn’t predictable”

If benefits cannot be assured early on, stop pretending otherwise at prioritization. Don’t inflate projections to get in line. If the projections are believed, document them and revisit post-delivery. If functionality is being built with no credible evidence of benefit, record that too—those funding the work deserve to know how much is a shot in the dark.

This is not about eliminating failure; failure is part of innovation. The problem is that the Classic Enterprise often doesn’t even notice when an initiative failed to produce adequate impact, so it neither decommissions nor avoids the associated run costs.

“Our PMO/VMO already takes care of this”

They likely don’t. An idea justification template lacks both the means and the mandate to verify impact after delivery. The template may lack pointed questions, or vague answers are accepted, and there is a tendency to report benefits realized merely because work was completed or money was spent. If they truly have an equivalent questionnaire and it is properly filled out before work arrives, use it—no need to duplicate effort.

“This isn’t collaborative”

Those accustomed to getting their priorities scheduled may cry gatekeeping. That is why the reformist CTO should seek the backing of the COO or CFO before starting. Also, consider the naming: the phrase Robust Demand Management might carry baggage. Calling it Verifiable Ideas or Ideas with Full Disclosure may make socialization easier.

Taking the Lead

If technology leaders outside your function aren’t pushing for better impact intelligence, it is in your and the company’s interest to do so. Institute robust demand management, pay down measurement debt, introduce impact validation, and publish projection-versus-performance reports. Equip your teams to aim for business impact, and the developer productivity pressure eases—more importantly, you can lead on the business impact of discretionary spend.

This path is harder than simply addressing the productivity challenge. Without it, you may never speak to true business impact and will remain trapped in the vicious cycle. The C-Suite will always see your role as executional: focused on delivery, infrastructure, and operations. There is no shame in that—unless you believe you can do better.