Anthropic’s 25× CI Spike: How They Redesigned Test Impact Analysis

Coffee Summary

  • FACT: Anthropic engineer Sachin Malhotra (blog dated September 14, 2026) reports CI job volume rose 25× over six months as agentic coding accelerated shipping and tests.
  • FACT: Engineers ship 8× as much code per quarter vs 2021–2025 averages; Claude authors ~80% of that code and heavily assists PR review.
  • FACT: Test-impact “listener + selector” hit scaling walls; three patches bought ~70 days, ~29 days, then <1 day before a journal/DB redesign enabled horizontal scale.
  • FACT / advice: Malhotra urges teams to assume ~25× CI load within two quarters, keep state out of process, and instrument listener lag.
  • OPINION (AIImpish): If coding agents are already on your team, CI capacity planning is a product risk — not a weekend ops chore.

What happened

On September 14, 2026, Anthropic published an engineering post by Sachin Malhotra: *“Agentic coding is straining CI. Here’s how we scaled test impact analysis at Anthropic.”* FACT: The post states CI job volume increased 25× over six months. Drivers cited: engineers on average ship 8× as much code per quarter as they did from 2021–2025; Claude authors 80% of that code and plays a large role in reviewing and approving PRs; tests across the codebase grew ~10× while headcount rose only nominally.

Anthropic’s test impact analysis (test selection) service decides which tests run per change using past performance and package relevance — so not every test runs on every PR. The service depends on two components staying in sync: a listener that records results from every CI run, and a selector that uses that history to choose tests for opened PRs.

Why it matters

Writing code is no longer the bottleneck; once PR review speeds up, CI becomes the choke point. FACT: When many CI jobs run every second, the listener can fall behind the PR queue. Malhotra notes that ~20 minutes of listener lag can mean tens of thousands of test updates not applied to the selector — which leads to stale selection: flaky/failing tests kept in the critical path, fixed tests delayed, and noisy investigations when bad changes merge.

Agents also change PR shape: Claude prefers smaller, more granular PRs, raising daily CI job count, with overnight/weekend agent activity lifting the floor. Full-suite-on-every-PR teams feel this first.

What changed

Patch timeline (FACT — Anthropic blog)

Patch Approach How long it lasted (stated)
1 Bigger machine (doubled cores) ~70 days
2 Shard listener by package (single writer per package) ~29 days
3 Daily restarts / memory workarounds Less than a day of relief; restarts also caused further lag

FACT: The v0 design kept a running history per test in a single process / single writer, which blocked horizontal sharding. After patches failed to keep up, the redesign gave the service an in-memory data store: listeners append to a journal and stay largely stateless; a small consumer rolls the journal into per-test history every few seconds; the selector reads that history quickly.

FACT: Redesign took ~three weeks for one engineer (vs ~a quarter a year earlier); stable since cutover. Distributed design is more expensive but easier to scale and profile than a shaky singleton.

Lessons labeled for builders

  • FACT (author advice): Assume architecture must handle ~25× load within two quarters; plan for 10–20× perceived scale in v0 if budget allows.
  • FACT: Instrument so CI jobs in ≈ jobs out; treat listener lag as a first-class SLO.
  • FACT: Keep state out of the process; avoid critical singletons you cannot measure or canary.

Who should care

  • CI / DevEx / platform engineers whose suites already slow merges when agents open more PRs.
  • Teams adopting Claude Code or other coding agents who have not resized test selection, runners, or caches.
  • Eng managers tracking “agent velocity” without tracking CI queue depth and flake rate.
  • Vendors/buyers of test-impact products needing a realistic AI-native load story.

Limitations

  • Load multipliers (25×, 8×, 10×, 80%) are Anthropic-stated — not third-party audited.
  • Claude Tag anecdotes are illustrative; some conversations are recreated in the post.
  • No vendor names, exact store product, or open-source blueprint published.
  • “Horizontally scaled test selection will become industry standard” is author anticipation (CLAIM/OPINION).

What to do next

  1. Measure current CI jobs/day and listener/selector lag before your next agent rollout wave.
  2. If every test still runs on every PR, time-box a test-impact pilot — agents amplify full-suite cost fastest.
  3. Prefer designs that keep selection state out of process (journal + consumer) so you can shard listeners.
  4. Alert on lag thresholds (Anthropic’s internal example used ~50k jobs behind) before pages become daily.
  5. Budget for 10–25× CI growth over ~two quarters if agent authorship is rising — half-measures expire faster than last year.

AIImpish Take

Anthropic’s post is a concrete FACT trail: agentic coding drove a 25× CI job spike, quick patches bought shrinking windows, and a stateless journal + horizontal listeners redesign restored stability. For any team shipping with coding agents, the actionable lesson is planning load and instrumentation early — not copying Anthropic’s exact stack.