Anthropic’s 25× CI Spike: How They Redesigned Test Impact Analysis
Coffee Summary
- FACT: Anthropic engineer Sachin Malhotra (blog dated September 14, 2026) reports CI job volume rose 25× over six months as agentic coding accelerated shipping and tests.
- FACT: Engineers ship 8× as much code per quarter vs 2021–2025 averages; Claude authors ~80% of that code and heavily assists PR review.
- FACT: Test-impact “listener + selector” hit scaling walls; three patches bought ~70 days, ~29 days, then <1 day before a journal/DB redesign enabled horizontal scale.
- FACT / advice: Malhotra urges teams to assume ~25× CI load within two quarters, keep state out of process, and instrument listener lag.
- OPINION (AIImpish): If coding agents are already on your team, CI capacity planning is a product risk — not a weekend ops chore.
What happened
On September 14, 2026, Anthropic published an engineering post by Sachin Malhotra: *“Agentic coding is straining CI. Here’s how we scaled test impact analysis at Anthropic.”* FACT: The post states CI job volume increased 25× over six months. Drivers cited: engineers on average ship 8× as much code per quarter as they did from 2021–2025; Claude authors 80% of that code and plays a large role in reviewing and approving PRs; tests across the codebase grew ~10× while headcount rose only nominally.
Anthropic’s test impact analysis (test selection) service decides which tests run per change using past performance and package relevance — so not every test runs on every PR. The service depends on two components staying in sync: a listener that records results from every CI run, and a selector that uses that history to choose tests for opened PRs.
Why it matters
Writing code is no longer the bottleneck; once PR review speeds up, CI becomes the choke point. FACT: When many CI jobs run every second, the listener can fall behind the PR queue. Malhotra notes that ~20 minutes of listener lag can mean tens of thousands of test updates not applied to the selector — which leads to stale selection: flaky/failing tests kept in the critical path, fixed tests delayed, and noisy investigations when bad changes merge.
Agents also change PR shape: Claude prefers smaller, more granular PRs, raising daily CI job count, with overnight/weekend agent activity lifting the floor. Full-suite-on-every-PR teams feel this first.
What changed
Patch timeline (FACT — Anthropic blog)
| Patch | Approach | How long it lasted (stated) |
|---|---|---|
| 1 | Bigger machine (doubled cores) | ~70 days |
| 2 | Shard listener by package (single writer per package) | ~29 days |
| 3 | Daily restarts / memory workarounds | Less than a day of relief; restarts also caused further lag |
FACT: The v0 design kept a running history per test in a single process / single writer, which blocked horizontal sharding. After patches failed to keep up, the redesign gave the service an in-memory data store: listeners append to a journal and stay largely stateless; a small consumer rolls the journal into per-test history every few seconds; the selector reads that history quickly.
FACT: Redesign took ~three weeks for one engineer (vs ~a quarter a year earlier); stable since cutover. Distributed design is more expensive but easier to scale and profile than a shaky singleton.
Lessons labeled for builders
- FACT (author advice): Assume architecture must handle ~25× load within two quarters; plan for 10–20× perceived scale in v0 if budget allows.
- FACT: Instrument so CI jobs in ≈ jobs out; treat listener lag as a first-class SLO.
- FACT: Keep state out of the process; avoid critical singletons you cannot measure or canary.
Who should care
- CI / DevEx / platform engineers whose suites already slow merges when agents open more PRs.
- Teams adopting Claude Code or other coding agents who have not resized test selection, runners, or caches.
- Eng managers tracking “agent velocity” without tracking CI queue depth and flake rate.
- Vendors/buyers of test-impact products needing a realistic AI-native load story.
Limitations
- Load multipliers (25×, 8×, 10×, 80%) are Anthropic-stated — not third-party audited.
- Claude Tag anecdotes are illustrative; some conversations are recreated in the post.
- No vendor names, exact store product, or open-source blueprint published.
- “Horizontally scaled test selection will become industry standard” is author anticipation (CLAIM/OPINION).
What to do next
- Measure current CI jobs/day and listener/selector lag before your next agent rollout wave.
- If every test still runs on every PR, time-box a test-impact pilot — agents amplify full-suite cost fastest.
- Prefer designs that keep selection state out of process (journal + consumer) so you can shard listeners.
- Alert on lag thresholds (Anthropic’s internal example used ~50k jobs behind) before pages become daily.
- Budget for 10–25× CI growth over ~two quarters if agent authorship is rising — half-measures expire faster than last year.
AIImpish Take
Anthropic’s post is a concrete FACT trail: agentic coding drove a 25× CI job spike, quick patches bought shrinking windows, and a stateless journal + horizontal listeners redesign restored stability. For any team shipping with coding agents, the actionable lesson is planning load and instrumentation early — not copying Anthropic’s exact stack.
AIImpish