On September 25, GitHub added a new metric to its enterprise Copilot report: it splits the total wait time of a PR (pull request — a code change submitted by a developer) from 'ready for review' to 'merged' into three separately-timed segments, with each segment reporting the median (P50 — half of PRs are faster, half slower) and the 90th percentile (P90 — the slowest 10%). We note: the significance isn't 'one more data point,' but that it decomposes the vague complaint 'we ship slow' into three problems that can each be addressed individually.

What this is

The three-phase split is straightforward:

  • Phase 1 — 'Marked Ready to First Reviewer' — measures team responsiveness; the bottleneck may simply be that nobody opens it.
  • Phase 2 — 'First Review to Last Review' — measures collaboration complexity, including back-and-forth revisions, supplementary tests, and scope changes.
  • Phase 3 — 'Final Approval to Actual Merge' — measures release and permissions workflow; often a machine or process is the thing waiting, not a code discussion.

Key caveats to note: the current implementation counts only PRs created by humans and merged after review by at least one other human; Copilot and other bot reviews are excluded. There is no historical backfill — data before September 21 is not included in the stats.

Industry view

GitHub's own framing: these metrics are symptoms of a queuing system, not developer performance scores. Using P90 directly to rank individuals incentivizes rubber-stamping 'approvals,' splitting PRs for no good reason, and even bypassing reviews — clearly contrary to the tool's design intent.

We care more about another methodological trap: naively averaging P90 across multiple repos does not yield the organization's P90. One slow PR in a small repo will be weighted the same as hundreds of PRs in a large repo — meaning you might hold a full week of meetings over a trivial side project. Without raw durations, at minimum weight by merge volume, or break out the repo dimension separately.

There's another risk the source flags but easily misses: faster reviews don't equal better quality. Rollback rate, production defects, follow-up fixes — these 'quality signals' must be read alongside speed metrics, otherwise it's easy to overlook problems like CI (continuous integration — the automated pipeline that runs tests and builds) queues disguising human waits as 'slow collaboration.'

Impact on regular people

For enterprise IT: Teams with enterprise Copilot enabled can pull this data directly and localize bottlenecks by repo and change type — slow response means assigning rotating reviewers, slow collaboration means splitting PRs, slow merges means updating auto-merge rules. Far more useful than 'dev is slow.'

For individual careers: Developers won't be directly scored on this data in the short term, but be aware management can see it — slow review response will surface as a team-level problem, not as an individual KPI.

For consumer markets: No direct consumer impact, but the move reflects an industry trend: AI tool vendors are filling the 'outcome measurement' gap, shifting from 'selling tools' to 'selling results.'