OpenAI Agents API Public Beta: Managed Codex Harness Explained
Coffee Summary
- OpenAI launched the Agents API in public beta on September 10, 2026 — a managed Codex-style harness over a simple API.
- OpenAI hosts sessions, orchestration, context compaction, and recovery; you choose the compute environment.
- Environments: OpenAI-hosted sandbox, self-hosted, or sandbox partners; hosted = Linux with Python/Node/CLI at `/workspace`.
- No separate Agents API fee — pay model tokens, tools, and hosted sandbox/container rates; beta header `OpenAI-Beta: agents=v1`.
- Caveat: data residency is currently US-only; the Agents API does not support Zero Data Retention, even with a self-hosted sandbox.
What happened
On September 10, 2026, OpenAI put the Agents API into public beta. The pitch is straightforward: the same harness and infrastructure that powers Codex — sessions, orchestration, context management, recovery, tools, and subagents — is now available through an API instead of only through Codex-style products.
OpenAI hosts and maintains that harness. Your application supplies the task, tools, and knowledge, then chooses where the agent actually runs.
Why it matters
Most “agent” projects fail on plumbing, not prompts. Teams rebuild session durability, context compaction, tool search, artifact handling, and subagent orchestration — then redo it when a new model ships.
The Agents API productizes that layer. Docs describe a managed Codex harness that can run commands in a sandbox, apply skills, connect via MCP, steer mid-run, compact context across long sessions, delegate to subagents, and resume where it left off.
Early customer quotes on the launch post (company-reported, not independent benchmarks):
- **Ciridae** (Jack Weissenberger, CTO): evaluation score **0.71 → 0.85**, with a claimed **4x latency reduction** tied to subagent flows.
- **SafetyKit** (Bhavyansh Sabharwal): claimed **60% reduction in cost per case** after migrating a case-review workflow, plus lower latency and better token efficiency while maintaining existing performance.
Treat those as launch-post CLAIMs until you reproduce them on your own evals.
What changed
Managed harness vs building your own
| Layer | Build yourself | Agents API (public beta) |
| — | — | — |
| Session / recovery | Custom | OpenAI-managed |
| Context compaction | Custom | Built into harness |
| Subagents / orchestration | Custom | Multi-agent support |
| Tools | Your stack | MCP, functions, built-ins (e.g. web search) |
| Compute | Your choice | OpenAI-hosted, self-hosted, or partners |
| Pricing for the API itself | N/A | No separate Agents API fee |
Example model in OpenAI’s docs snippets: `gpt-6-astra`. Requests use the beta header `OpenAI-Beta: agents=v1`.
Hosted vs self-hosted vs partners
You pick the environment:
1. OpenAI-hosted sandbox — Linux workspace with Python, Node.js, and CLI tools; working directory `/workspace`. OpenAI provisions it. Files under `/workspace/outputs` can be published as artifacts when a turn completes. Hosted sandboxes bill at standard container rates; model usage is separate.
2. Self-hosted — your image, compute, or private network when you need tighter control.
3. Sandbox partners — launch post names ecosystem integrations including Blaxel, Cloudflare, Daytona, DigitalOcean, E2B, Modal, Oracle, Runloop, and Vercel.
Hosted is the fast path. Self-hosted / partner sandboxes matter when you need VPC placement, custom compute, or specific storage — not when you merely want Zero Data Retention (see caveats).
Who should care
- Teams shipping long-running coding, incident, research, or ops agents who do not want to own harness internals.
- Platforms that already have tools and UX, but need durable sessions and subagent orchestration.
- Builders comparing “glue an LLM to tools” vs a versioned, model-co-evolving harness.
Limitations
- **Public beta** — expect API surface and operational limits to move.
- **Data residency:** currently **US-only** per Agents API docs.
- **Zero Data Retention:** the Agents API **does not support ZDR**. Choosing a self-hosted sandbox **does not** make it ZDR-eligible.
- Hosted sandbox idle timeout: docs note deletion after about an hour without activity/keep-alives (not configurable) — plan for artifact export.
- Customer numbers on the launch page are CLAIMs; do not treat Ciridae/SafetyKit figures as industry benchmarks.
What to do next
1. Skim the Agents API overview and OpenAI-hosted sandbox docs; confirm US residency / non-ZDR fit your compliance bar.
2. Prototype one task end-to-end in a hosted sandbox (files → code → `/workspace/outputs`).
3. Re-run your internal eval suite before/after migrating orchestration (do not rely on Ciridae/SafetyKit quotes alone).
4. Decide environment strategy: hosted for speed, self-hosted/partner for network and compute constraints.
5. Budget tokens + tools + container time — there is no separate Agents API fee, but hosted sandboxes are not free.
AIImpish Take
The Agents API is less “a new model” and more “stop rebuilding Codex.” If your bottleneck is harness reliability — sessions, compaction, subagents, artifacts — this beta is worth a spike. If your bottleneck is compliance (non-US residency or ZDR), read the data-controls caveats first: self-hosting the sandbox does not buy Zero Data Retention. Price the whole stack (model + tools + containers), then judge the API on your own evals, not launch-page CLAIMs.
AIImpish