What DevOps Delivery Metrics Measure, and What They Miss

A sticker reading DevOps held up in front of a blurred background

Most articles about DevOps measurement still describe four metrics, one of them MTTR. That set is out of date. DORA, the research programme Google Cloud runs, currently publishes five, and it replaced mean time to restore with failed deployment recovery time. Getting the list right matters less than knowing what the list can carry: these are self-reported survey correlations from a voluntary sample, published by a company that sells the tooling, and the researcher who led the original work has since co-authored the argument for not reading them in isolation.

The Set Is Five Metrics Now, and MTTR Is Gone

Google Cloud says the original key metrics were introduced in 2013. Since then the programme has retired the mean time to restore framing in favour of failed deployment recovery time and added a fifth measure, so a dashboard still showing four boxes with MTTR in one of them is tracking a model DORA itself has moved past.

The addition is the interesting part, because it catches work the older four hide. A team can post excellent deployment frequency while a large share of those deployments are emergency repairs of the ones before. The current published set, grouped into throughput and instability, runs as follows:

  1. Change lead time, measured from commit to production.
  2. Deployment frequency.
  3. Failed deployment recovery time.
  4. Change fail rate, the share of deployments that need immediate intervention, a rollback or a hotfix.
  5. Deployment rework rate, the share of unplanned deployments caused by production incidents.

A Voluntary Survey Run by the Company Selling the Tooling

DORA describes itself as a programme run by Google Cloud, with research reports published annually since 2014. The most recent is the 2025 DORA Report, titled State of AI-assisted Software Development, published on 23 September 2025 and drawing on nearly 5,000 technology professionals worldwide plus more than 100 hours of qualitative interviews.

That provenance sets the limits. The findings are self-reported survey associations from a non-random sample of people who elected to take a DevOps survey, published by a cloud vendor. They are correlations rather than causal measurements, and the population self-selects towards organisations already invested in the practices being studied. None of that makes the numbers useless; it makes them evidence about a particular population’s reported experience.

Within those limits the 2025 figures are pointed. Ninety percent of respondents report using artificial intelligence at work and more than 80 percent believe it has increased their productivity, while 30 percent report little or no trust in AI-generated code. Ninety percent of organisations have adopted at least one platform. AI adoption shows a positive relationship with delivery throughput and product performance, and continues to show a negative relationship with software delivery stability.

The 2024 report, published on 22 October 2024, put numbers on that tension. Google Cloud reported that a 25 percent increase in AI adoption was associated with a 7.5 percent increase in documentation quality, 3.4 percent in code quality and 3.1 percent in code review speed, against a 1.5 percent decrease in delivery throughput and a 7.2 percent reduction in delivery stability. In the same report, more than 75 percent relied on AI for at least one daily responsibility and 39 percent reported little to no trust in what it produced.

The Original Research Lead Co-Wrote the Rebuttal

The sharpest criticism of delivery metrics is not an outsider’s swipe. Nicole Forsgren, lead author of the original DORA research and later at GitHub, co-wrote The SPACE of Developer Productivity with Margaret-Anne Storey of the University of Victoria and Chandra Maddila, Thomas Zimmermann, Brian Houck and Jenna Butler of Microsoft, published in Communications of the ACM in June 2021.

The paper opens by naming the belief it wants to kill, calling it a myth that one productivity metric can tell you everything, when productivity spans several dimensions of work and is heavily shaped by the context the work happens in. On activity counts specifically, and deployment frequency is an activity count, the authors write that such metrics alone do not reveal which explanation applies and should never be used in isolation either to reward or to penalise developers.

SPACE proposes five dimensions to set against DORA’s five metrics: satisfaction and well-being, performance, activity, communication and collaboration, and efficiency and flow. Taken together the two frameworks make a reasonable working rule. Delivery metrics describe the delivery system. They do not describe the people in it, and the person who established them says so.

Easily Abused by Your Leaders

Nathen Harvey, who leads DORA at Google Cloud, told The New Stack in November 2023 that DORA is a lot more than four metrics and that the research programme itself is much more than four metrics, adding that the metrics can be easily abused by your leaders. He is explicit that they are meant for team-level improvement rather than individual performance evaluation, and cautions against turning them into surveillance instruments.

That warning is a design constraint, not a footnote. A change fail rate attached to an individual engineer will be optimised, and the cheapest way to optimise it is to deploy less and batch more, which degrades every other metric in the set. The five metrics are only coherent when they move together at the level of a team that owns a service.

Both Named Transformations Are Insider-Told or a Decade Old

The most cited DevOps turnaround is HP’s LaserJet FutureSmart firmware programme, a team of 400 to 800 developers worldwide whose change began around 2006 and 2007 and was documented through 2013. Gary Gruver, then HP’s director of engineering, reported through IT Revolution in February 2014 that the build cycle went from one week to three hours with 10 to 15 builds a day, commits from one a day to 100 a day, and full regression testing from six weeks to 24 hours. Engineering capacity spent on new features rose from 5 percent to 40 percent, against a previous state in which 80 to 90 percent of effort went on porting and qualifying existing code, and the team was changing 75,000 to 100,000 lines of code daily afterwards.

Every one of those figures originates with the people who ran the programme, published first in their own book with Mike Young and Pat Fulghum in 2012. No independent audit of them exists. They are specific, attributable and plausible, which is a different standing from measured.

The second commonly cited example, Etsy, has the same problem plus age. InfoQ reported in March 2014 on Daniel Schauenberg’s QCon London talk describing roughly 50 deploys a day, more than 14,000 test suite runs daily on a continuous integration cluster serving 150 engineers, at 60 million monthly visits and 1.5 billion monthly page views. The practices named were Jenkins, Chef, a one-click deploy tool called Deployinator, feature flags for A/B testing, blameless post-mortems recorded in a tool called Morgue, and developers on call roughly one week a month. That is a valuable historical account of what the practice looked like in 2014, and it should be presented as history rather than as current benchmark data.

The practical upshot for a team adopting this is unglamorous. Track all five metrics rather than the two that flatter, treat DORA’s percentages as one vendor’s survey population rather than physics, pair them with something that measures the people doing the work, and be honest that the famous before-and-after stories were written by their own participants.

Sources: Google Cloud (2025 DORA Report announcement) · DORA · Communications of the ACM (Forsgren et al., SPACE) · The New Stack · IT Revolution · InfoQ

More articles to read