Category: AI & Tools
-

Prompt Cache Diagnostics GA: Raise Hit Rate Without Rewriting Everything
Prompt Cache Diagnostics GA: Raise Hit Rate Without Rewriting Everything Coffee Summary FACT (changelog Sep 10, 2026): Prompt Cache Diagnostics is generally available in the Responses API for GPT-5.6 and later supported models. Compare cache reuse against a previous response, see miss reasons, and follow troubleshooting guidance. Goal: improve reuse (latency + cost) without inventing…
-

OpenAI Project API Key Expiration and Org/Project Max Lifetime
OpenAI Project API Key Expiration and Org/Project Max Lifetime Coffee Summary FACT (changelog Sep 10, 2026): You can set expiration dates when creating project API keys. FACT: Admins can enforce a maximum key lifetime at organization or project level in Platform settings; new keys must expire within that limit. FACT (production best practices): OpenAI strongly…
-

GPT-6 Astra at Critical Cyber: Builder Migration Checklist
GPT-6 Astra at Critical Cyber: Builder Migration Checklist Coffee Summary CLAIM (OpenAI): GPT-6 Astra is the first broadly deployed OpenAI model to reach Critical cybersecurity capability under the Preparedness Framework (safety overview dated Sep 3, 2026; changelog Sep 8). FACT (changelog): Migration gotchas — no none reasoning effort; no custom temperature / top_p / logprobs;…
-

Amodei “Pace the Frontier”: Altman Matches Evaluators, Calls 2026 IPO Ill-Advised
Amodei “Pace the Frontier”: Altman Matches Evaluators, Calls 2026 IPO Ill-Advised Coffee Summary FACT: On September 12, 2026, Anthropic CEO Dario Amodei published the essay *We Must Pace the Frontier*, proposing a three-step plan to slow capability growth without a full halt. FACT: Anthropic unilaterally committed to step 1 — permanent, employee-like access for third-party…
-

Claude Fable 5.1: Same Model as Mythos, 75% Cheaper Cache Reads
Claude Fable 5.1: Same Model as Mythos, 75% Cheaper Cache Reads Coffee Summary Anthropic launched Claude Fable 5.1 (GA) and Claude Mythos 5.1 (trusted access) in September 2026. Same underlying model; different safeguards. Model ID: `claude-fable-5-1`. Base price unchanged vs Fable 5: $10 / $50 per 1M input/output; cache reads now $0.25/1M (75% less). Anthropic…
-

Gemini 3.8 Flash: Intro Pricing Window Through Dec 31, 2026
Gemini 3.8 Flash: Intro Pricing Window Through Dec 31, 2026 Coffee Summary Google announced Gemini 3.8 Flash and 3.8 Flash Cyber on September 2, 2026. 3.8 Flash intro price matches 3.7 Flash: $0.75 / $3.75 per 1M input/output tokens through December 31, 2026. From January 1, 2027: $1.50 / $7.50 per 1M input/output. Positioned as…
-

GPT-Live-1 in the API: Full-Duplex Voice at $0.05/Min
GPT-Live-1 in the API: Full-Duplex Voice at $0.05/Min Coffee Summary GPT-Live-1 became generally available in the OpenAI API on September 10, 2026. Full-duplex: listen and speak at once; handles interruptions, tone/pace/style via system prompt, noise/silence, telephony. Voice layer priced at $0.05 per minute, billed per second (not rounded up); backend models/tools bill separately. Can delegate…
-

Meta Muse Explained: Secure VM, Sentinel, and What It Can Actually Do
Meta Muse Explained: Secure VM, Sentinel, and What It Can Actually Do Coffee Summary Meta announced Muse on September 8, 2026 — a personal agent powered by Muse Spark. It runs in Muse Secure VM: a dedicated VM with its own browser; chat via Muse app or WhatsApp. Sentinel gates internet actions; Muse asks before…
-

Gemini App for Windows: Alt+Space Desktop AI in 5 Minutes
Gemini App for Windows: Alt+Space Desktop AI in 5 Minutes Coffee Summary Google launched the Gemini app for Windows on September 10, 2026, globally for Windows 10 and 11. Download at gemini.google/desktop; press Alt+Space for an overlay without leaving your current screen. Inside: Gemini Spark for multi-step tasks, pull from Gmail/Drive, Nano Banana images, Gemini…
-

OpenAI Agents API Public Beta: Managed Codex Harness Explained
OpenAI Agents API Public Beta: Managed Codex Harness Explained Coffee Summary OpenAI launched the Agents API in public beta on September 10, 2026 — a managed Codex-style harness over a simple API. OpenAI hosts sessions, orchestration, context compaction, and recovery; you choose the compute environment. Environments: OpenAI-hosted sandbox, self-hosted, or sandbox partners; hosted = Linux…