Blog · DX for executives
Is AI paying off? Measure it like an executive.
Licenses are easy to count. Impact is not. What two DX papers tell senior leaders about measuring AI in engineering – utilization, impact and cost – and how to lead the rollout without losing your teams’ trust.
The question on every board agenda
Your teams have AI tools. The licenses are paid. And at some point someone on the board asks the obvious question: is it working? Counting licenses or active users does not answer it. Neither do the vendor claims you read online.
Two papers from DX, the engineering intelligence company, are the best answer we know. Measuring AI code assistants and agents, by Abi Noda and Laura Tacho, sets out the DX AI Measurement Framework. The AI strategy playbook for senior executives, by Justin Reock, turns it into a leadership plan. Here is what they say, in our words, and what it means for you.
One finding should shape every AI budget. The playbook shows that the average gains from AI are positive but modest – and that the averages hide the real story. Broken down by company, some companies see change confidence rise by twenty points with AI, while others see it fall by twenty. Some organizations get much better with AI. Others get worse. The tool is the same; the leadership around it is not.
The DX AI Measurement Framework
Three dimensions. In this order.
The framework follows how AI adoption actually unfolds: first people start using the tools, then you look for the effect, then you optimize the spend. Utilization on its own proves nothing; it matters because you can correlate it with the productivity metrics you already trust.
-
Utilization
Are people using the tools?
- Daily and weekly active users of AI tools
- Share of pull requests that are AI-assisted
- Share of committed code that is AI-generated
- Tasks handed to agents
-
Impact
Is engineering getting better?
- Time saved per developer and week
- Developer satisfaction
- DX Core 4: PR throughput, perceived rate of delivery, Developer Experience Index
- Quality: maintainability, change confidence, change failure rate
- Hours of work completed by agents
-
Cost
Is the spend worth it?
- AI spend, in total and per developer
- Net time gain: time saved minus spend
- What an agent hour costs
Summarized from the DX AI Measurement Framework. The agent metrics are new, and DX expects them to change as agentic tools mature.
Where the data comes from
No single number tells you the truth
Both papers recommend a mix: direct signals such as time saved for a quick read on each tool, and the longer trend in your core productivity metrics to find hidden benefits and hidden risks.
Telemetry
System data from your tools: active users, accepted suggestions, AI-generated code. Continuous, but it tells part of the story – an accepted suggestion may still have been rewritten.
Surveys
Regular, for example quarterly: satisfaction, perceived productivity, confidence in changes, maintainability. What system data cannot see, at the cost of a snapshot in time.
Experience sampling
A short question at the moment of work – after a merge: “Did AI help with this PR, and how much time did it save?” The best way to find the use cases with real returns.
Five things only leadership can do
Measure teams, never individuals
Metrics such as lines of AI-generated code are easy to game. Use them to judge people and you get compliance instead of results, and the data becomes worthless. DX recommends saying plainly that AI metrics stay out of performance reviews, that the purpose is to understand what helps, and that the data decides where the money goes.
Count agents as part of the team
An agent’s pull request belongs to the team that directs it. That is how the framework counts throughput, and it hints at where the work is going: every developer will increasingly lead a small team of agents, and be measured more like a manager is today.
Balance speed with quality
Speed gains in the first months can turn into maintenance debt later. Track change failure rate, maintainability and confidence alongside throughput. Keep code review mandatory on production paths, and keep your tests and linting running unchanged or stronger.
Make it safe to learn
The playbook puts psychological safety first, and grounds it in Google’s Project Aristotle: people who fear for their jobs don’t experiment. Frame AI as a way to augment your engineers, give them time to learn inside the sprint, and expect a dip before the gains. Pilots are safe-to-fail experiments, not mandates.
Look past code generation
For most organizations, writing code was never the bottleneck. The biggest returns in the playbook’s examples come from code review agents, incident handling, legacy refactoring and documentation – wherever the team’s time is actually stuck. Find the constraint first, then apply AI to it.
From the playbook
A plan in three phases
-
Months 1–3
Foundation
- Get the leadership team on one data-driven picture
- Say why: augment engineers, not cut heads
- Set up the measurement: utilization, impact, cost
- Write simple AI rules with Security, Legal and Compliance
-
Months 4–9
Enablement
- Remove barriers: licenses, approved tools, private models where IP is at stake
- Give people time to learn, inside the sprint
- Pilot agents on bottlenecks the teams name themselves
- Measure the pilots and show the teams that win
-
Month 10 on
Scale
- Roll out what the pilots proved
- Track AI against your core engineering KPIs, continuously
- Keep a living AI guide for every engineer
- Look for the next bottleneck, and the next
Where XALT comes in
XALT is an Atlassian Platinum Solution Partner, for DX too, with specializations in Software Development, Service Management (ITSM) and Cloud Migration. We introduce DX in your engineering organization and set up the measurement the papers describe: the baseline before the rollout, the framework’s three dimensions, and the communication that keeps your teams’ trust.
The measurement then tells you where to go next. It shows which teams already benefit, where the bottleneck sits, and which step on the journey to agentic engineering pays off first: a developer self-service, a DevEx platform, or agents that write the code and the tests and wait for a person’s approval.


Sources
- Abi Noda and Laura Tacho, Measuring AI code assistants and agents, DX, 2025.
- Justin Reock, The AI strategy playbook for senior executives, DX, 2025.
Both papers are summarized here in our own words. Learn more at getdx.com.