Blog · DX for executives

Is AI paying off? Measure it like an executive.

Licenses are easy to count. Impact is not. What two DX papers tell senior leaders about measuring AI in engineering – utilization, impact and cost – and how to lead the rollout without losing your teams’ trust.

Philipp Göllner · · 8 min read

The question on every board agenda

Your teams have AI tools. The licenses are paid. And at some point someone on the board asks the obvious question: is it working? Counting licenses or active users does not answer it. Neither do the vendor claims you read online.

Two papers from DX, the engineering intelligence company, are the best answer we know. Measuring AI code assistants and agents, by Abi Noda and Laura Tacho, sets out the DX AI Measurement Framework. The AI strategy playbook for senior executives, by Justin Reock, turns it into a leadership plan. Here is what they say, in our words, and what it means for you.

One finding should shape every AI budget. The playbook shows that the average gains from AI are positive but modest – and that the averages hide the real story. Broken down by company, some companies see change confidence rise by twenty points with AI, while others see it fall by twenty. Some organizations get much better with AI. Others get worse. The tool is the same; the leadership around it is not.

The DX AI Measurement Framework

Three dimensions. In this order.

The framework follows how AI adoption actually unfolds: first people start using the tools, then you look for the effect, then you optimize the spend. Utilization on its own proves nothing; it matters because you can correlate it with the productivity metrics you already trust.

  1. Utilization

    Are people using the tools?

    • Daily and weekly active users of AI tools
    • Share of pull requests that are AI-assisted
    • Share of committed code that is AI-generated
    • Tasks handed to agents
  2. Impact

    Is engineering getting better?

    • Time saved per developer and week
    • Developer satisfaction
    • DX Core 4: PR throughput, perceived rate of delivery, Developer Experience Index
    • Quality: maintainability, change confidence, change failure rate
    • Hours of work completed by agents
  3. Cost

    Is the spend worth it?

    • AI spend, in total and per developer
    • Net time gain: time saved minus spend
    • What an agent hour costs

Summarized from the DX AI Measurement Framework. The agent metrics are new, and DX expects them to change as agentic tools mature.

Where the data comes from

No single number tells you the truth

Both papers recommend a mix: direct signals such as time saved for a quick read on each tool, and the longer trend in your core productivity metrics to find hidden benefits and hidden risks.

  • Telemetry

    System data from your tools: active users, accepted suggestions, AI-generated code. Continuous, but it tells part of the story – an accepted suggestion may still have been rewritten.

  • Surveys

    Regular, for example quarterly: satisfaction, perceived productivity, confidence in changes, maintainability. What system data cannot see, at the cost of a snapshot in time.

  • Experience sampling

    A short question at the moment of work – after a merge: “Did AI help with this PR, and how much time did it save?” The best way to find the use cases with real returns.

Five things only leadership can do

Measure teams, never individuals

Metrics such as lines of AI-generated code are easy to game. Use them to judge people and you get compliance instead of results, and the data becomes worthless. DX recommends saying plainly that AI metrics stay out of performance reviews, that the purpose is to understand what helps, and that the data decides where the money goes.

Count agents as part of the team

An agent’s pull request belongs to the team that directs it. That is how the framework counts throughput, and it hints at where the work is going: every developer will increasingly lead a small team of agents, and be measured more like a manager is today.

Balance speed with quality

Speed gains in the first months can turn into maintenance debt later. Track change failure rate, maintainability and confidence alongside throughput. Keep code review mandatory on production paths, and keep your tests and linting running unchanged or stronger.

Make it safe to learn

The playbook puts psychological safety first, and grounds it in Google’s Project Aristotle: people who fear for their jobs don’t experiment. Frame AI as a way to augment your engineers, give them time to learn inside the sprint, and expect a dip before the gains. Pilots are safe-to-fail experiments, not mandates.

Look past code generation

For most organizations, writing code was never the bottleneck. The biggest returns in the playbook’s examples come from code review agents, incident handling, legacy refactoring and documentation – wherever the team’s time is actually stuck. Find the constraint first, then apply AI to it.

From the playbook

A plan in three phases

  1. Months 1–3

    Foundation

    • Get the leadership team on one data-driven picture
    • Say why: augment engineers, not cut heads
    • Set up the measurement: utilization, impact, cost
    • Write simple AI rules with Security, Legal and Compliance
  2. Months 4–9

    Enablement

    • Remove barriers: licenses, approved tools, private models where IP is at stake
    • Give people time to learn, inside the sprint
    • Pilot agents on bottlenecks the teams name themselves
    • Measure the pilots and show the teams that win
  3. Month 10 on

    Scale

    • Roll out what the pilots proved
    • Track AI against your core engineering KPIs, continuously
    • Keep a living AI guide for every engineer
    • Look for the next bottleneck, and the next

Where XALT comes in

XALT is an Atlassian Platinum Solution Partner, for DX too, with specializations in Software Development, Service Management (ITSM) and Cloud Migration. We introduce DX in your engineering organization and set up the measurement the papers describe: the baseline before the rollout, the framework’s three dimensions, and the communication that keeps your teams’ trust.

The measurement then tells you where to go next. It shows which teams already benefit, where the bottleneck sits, and which step on the journey to agentic engineering pays off first: a developer self-service, a DevEx platform, or agents that write the code and the tests and wait for a person’s approval.

Atlassian Platinum Solution Partner Atlassian Software Development Specialization, EMEAAtlassian Service Management Specialization, EMEAAtlassian Cloud Migration Specialization, EMEA

Sources

  • Abi Noda and Laura Tacho, Measuring AI code assistants and agents, DX, 2025.
  • Justin Reock, The AI strategy playbook for senior executives, DX, 2025.

Both papers are summarized here in our own words. Learn more at getdx.com.

Do you know what AI is changing in your engineering organization?

Seven questions show where you stand, including cost control and upskilling. Or talk it through with Philipp.

This page as Markdown

blog-measure-ai-impact.md text/markdown · for AI agents Static file: /blog/measure-ai-impact/index.md

This is what an AI agent reads: the same content as the designed page, as plain Markdown. Switch back with “Human” at the bottom left.