2026-10-11 16:38 UTC

OpenAI claims its GPT-6 caching update preserves eligible prefixes for 30 minutes and adds explicit breakpoints, diagnostics, and cache-preserving reasoning changes, reducing latency and input costs for persistent agents.

state: corroboratedheat: highuncertainty: lowconvergesscott: highprompt-caching inference-economics agent-harnessesOpenAIGitHubManusWordsmith
Surfaced 2026-09-24T03:31:03Z — Better prompt caching for GPT-6 — Promoted to corroborated: beyond OpenAI's first-party announcement there are live developer docs making the controls independently verifiable and named third-party adopters (GitHub Copilot at billions-of-requests scale, Manus, Wordsmith) with quantified hit-rate improvements. The Reddit wave (317 pts, 98.6th percentile, 3 platforms) is amplification of the already-grounded announcement rather than new substance, but per the magnitude valve it justifies delivery heat — this is attention pricing, not a belief revision.

What is this?

On September 22, 2026 OpenAI announced improved prompt caching for the GPT-6 family: eligible shared prefixes reused within a 30-minute window get cache discounts (a guaranteed retention floor, up from 'as short as 5 minutes'), with 90% discounts on cached input-token reads. The update adds developer controls and observability — explicit cache breakpoints to pin what gets cached, a diagnostics payload for diagnosing misses (e.g. 'tools_changed'), and the ability to change reasoning effort or enable/disable tools without invalidating the cached prefix. OpenAI's GPT-6 announcement cites GitHub data that these changes cut the share of prompt tokens needing fresh processing by more than 50% across billions of Copilot requests. Third-party coverage (AIHubMix) adds that this comes with new cache-write fees on the GPT-5.6/GPT-6 generation, making it a deliberate trade of guaranteed retention and controllability for write costs. The case's key people list Manus and Wordsmith, but the supplied snippets don't mention them; only GitHub appears as a named beneficiary.

Why it matters to Scott

OpenAI has productized exactly what Scott's Prefix-Caching Economics and Scout-and-Senior work argues — guaranteed 30-minute retention makes stable append-only agent loops the cost-optimal shape — and, more interestingly, ships explicit cache breakpoints and cache-preserving tool/reasoning changes that directly bear on his open harness-design questions about tool churn breaking caches (radar:tool-schema-prompt-cache-invalidation) and mid-session reconfiguration (radar:claude-mid-conversation-reconfiguration). It also revalues adjacent open threads: guaranteed retention plus cache-write fees partially obsoletes the cache-warming hack (radar:cache-tax-idle-session-warming) and rewrites the five-minute-vs-one-hour tradeoff his routing stack (LiteLLM tiers, Codex default) sits on top of — a dated-receipts publishing window on 'the platform just ratified prefix-caching economics, and here's what the write fees change.'
ip:concept.prefix-caching-economicsip:source.the-scout-and-the-senior-ebookip:concept.model-barbellip:framework.context-engineeringdev:technology.litellmwork:technology.openai-apiradar:concept.prompt-cachingradar:concept.inference-economicsradar:cache-tax-idle-session-warmingradar:claude-code-cache-ttl-analyzerradar:tool-schema-prompt-cache-invalidationradar:claude-mid-conversation-reconfigurationradar:replay-prompt-cache-miss-audit
queries asked of Scott's wikis
  • prompt caching strategy in agent harness design — how do my harnesses structure system prompts and tool blocks for cache reuse
  • inference cost economics for persistent/long-running agents — input token costs, context reuse, cached vs fresh token pricing
  • agent memory and persistent context — how session state, wikis, and memory files interact with prompt prefix stability
  • breakpoints and prompt layout as explicit engineering control — any prior writing on prompt ordering, dynamic sections, tool-definition churn
  • local/open model inference vs frontier API costs — does 90% cached-read discount change the local-inference economics argument
  • cache-preserving reasoning effort and dynamic tool availability — patterns for varying effort/tools mid-session without breaking context

Measured heat

now 0 pts/hpeak 0 pts/hcomments 0/hpeers p14momentum: steady3 platformsage 452h
points/hour across evidence · reading as of 2026-10-12 02:59:37.977291+11:00 · deterministic, not a model opinion

How the heat travelled

09-22 21:00⭐ origin directly observedBetter prompt caching for GPT-6
OpenAI on openai
—
09-22 20:09first on hacker news · published · +-0.8hBetter prompt caching for GPT‑6
mehrdadrad
—
09-23 16:21first on r/OpenAI · published · +19.4hOpenAI rolls out upgraded prompt caching and reduced cached input rates for GPT-6
rhiever
—
09-22 20:09amplified on hacker newshn.story.49807383
mehrdadrad
peak 3 · 2 comments · 2% of case engagement
09-23 10:35amplified on hacker newshn.story.49814007
theanonymousone
peak 1 · 0 comments · 0% of case engagement
09-23 16:21amplified on r/OpenAI 👑reddit.post.1woal22
rhiever
peak 429 · 27 comments · 98% of case engagement
09-22 20:20our radar first saw it · +-0.7hdiscovery anchor: hn.story.49807383—
09-24 02:38reached heat=high · +29.6h · via queue+ledger——
pace: p83 vs 1032 stories at the 336h mark (now 452h old) — ahead of openai-chatgpt-ads-global-rollout (1.0x), behind microsoft-titan-jwt-bypass (1.0x)

Evidence (4) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟧 hnBetter prompt caching for GPT‑6
Retrieved article excerpt

Open article · Retrieved 2026-09-22T20:24:16.384144+00:00

September 22, 2026

[Product](https://openai.com/news/product-releases/)

# Better prompt caching for GPT‑6

Higher cache hit rates and new tools to help persistent agents run faster and cost less.

Loading…

Share

GPT‑6 enables persistent agents to work for hours on complex tasks, from refactoring codebases to producing well-researched documents and presentations. The applications behind these agents make a series of API requests that build on one another, often carrying forward the same instructions, tool definitions, and context from earlier turns. OpenAI caches that shared context to reuse computation across requests, reducing response times and giving developers discounts of up to 90% on cached input tokens.

With the GPT‑6 family, we launched an improved prompt caching system that delivers higher cache hit rates by default. We now give cache discounts for eligible shared prefixes reused within a 30-minute window. We’re also introducing new tools to help developers monitor cache performance, diagnose misses, and choose how much of a prompt to cache.

> “OpenAI’s prompt caching plays a critical role in helping GitHub Copilot deliver fast, efficient experiences at scale. Over the past several months, we’ve reduced by more than 50% the share of prompt tokens requiring fresh processing across billions of requests to OpenAI models, relative to our previous baseline. The result is a more efficient inference stack and faster time to first response for developers.”

—Mario Rodriguez, Chief Product Officer

## Monitor caching and diagnose cache misses

The new [Prompt Caching Dashboard⁠(opens in a new window)](https://platform.openai.com/usage?usage_section=prompt-caching) shows how much of your application’s input is served from cache. Track hit rates over time and use the input composition chart to compare cached and uncached tokens. These views help you spot drops in cache hits and evaluate how changes to your application impact caching performance.

Prompt caching dashboard showing cache hit rate, cache performance over time, and input token composition.

When you see an unexpected cache miss, use the [prompt caching diagnostics tool⁠(opens in a new window)](https://developers.openai.com/api/docs/guides/prompt-caching/diagnostics) to understand what happened. Compare a request with a recent response to identify changes to the model, tools, settings, or input that prevented reuse. The estimated number of affected tokens helps you assess the size of the impact and decide how you can optimize your integration to maximize cache hit rates.

```
{
  "prompt_cache_diagnostics": {
    "type": "cache_miss",
    "reason": "tools_changed",
    "comparison_reusable_tokens": 5629,
    "cache_missed_tokens": 5629
  }
}
```

## Optimize caching for your application

**Choose what to cache.** Explicit cache breakpoints let you choose which prompt prefixes to reuse. The refreshed [prompt caching guide⁠(opens in a new window)](https://developers.openai.com/api/docs/guides/prompt-caching) explains how to use them, how long cached prefixes remain eligible, and how changes to tools and inputs affect reuse.

**Adjust reasoning effort without breaking cache.** On GPT‑6 models, you can now [change reasoning effort⁠(opens in a new window)](https://developers.openai.com/api/docs/guides/reasoning?api-mode=responses#change-reasoning-mid-conversation) between responses without breaking cache. Raise effort for a harder task or lower it for a routine follow-up by appending a `configuration_update` while leaving request-level reasoning effort unchanged. This lets you adjust how much reasoning a task needs while preserving reusable context.

**Preserve cache as tools and instructions change.** As your agent’s tool use needs change, keep tool definitions, schemas, and ordering stable so earlier context stays reusable. Use `allowed_tools` to make only the relevant tools callable, or set tool\_choice to none when no tools are needed, instead of removing definitions. Use new developer messages to append new instructions towards the end of the context to override older ones. See our [guidance on managing tool changes⁠(opens in a new window)](https://developers.openai.com/api/docs/guides/prompt-caching#manage-tools-with-append-only-updates).

**Prewarm the cache to reduce latency.** [Prewarming⁠(opens in a new window)](https://developers.openai.com/api/docs/guides/prompt-caching#prewarm-the-cache) prepares known context ahead of time so the model can start responding sooner when a request arrives. For example, an application can prewarm shared instructions, tool definitions, or reference material during startup, before the user asks their first question. This moves processing out of the user’s wait time.

These optional controls build on the engine’s default performance, helping you tailor caching to your workload.

1 of 3

> “OpenAI’s prompt caching diagnostics and dashboard helped us improve cache hit rates by a few percentage points, reducing costs by 20%. We now get alerts when caching breaks unexpectedly and use Codex agents to diagnose the root cause. Explicit breakpoints also let us cache stable context while keeping frequently changing content at the end of the prompt. That’s made it economically viable to fork conversations for background tasks while reusing nearly all of the shared context.”

—Arian Hanifi, Chief Technology Officer

> “For long-running agents like Manus, reliable caching is fundamental to the economics. Working with OpenAI's engineering team, we refined cache breakpoint placement, combined explicit and automatic caching, and used real requests to pinpoint unexpected cache misses. In less than a week, our OpenAI model cache hit rate went from roughly 85% to consistently above 90%, further lowering inference costs in production. The progress came through a series of targeted improvements, with both teams validating the results along the way.”

—Bin Fan, Agent Team Lead

> “Working with OpenAI, we moved our session agents to explicit cache breakpoints. In under a week, cache hit rates on our evaluations rose from 83% to 91%. This meant fewer cache writes and lower inference costs with the same workload; the cache writes fell by roughly two-thirds, and inference costs by 36%.”

—Eugene Mikhantyev, AI Engineer

- Strawberry Browser
- Manus
- Wordsmith

- Strawberry Browser
- Manus
- Wordsmith

## Get started

- Monitor cache hit rates in the [Prompt Caching Dashboard⁠(opens in a new window)](https://platform.openai.com/usage?usage_section=prompt-caching).
- Investigate unexpected misses with the [diagnostics tool⁠(opens in a new window)](https://developers.openai.com/api/docs/guides/prompt-caching/diagnostics).
- Follow the [prompt caching guide⁠(opens in a new window)](https://developers.openai.com/api/docs/guides/prompt-caching) to improve your setup, or [use Codex⁠(opens in a new window)](https://developers.openai.com/api/docs/guides/prompt-caching#how-to-optimize-prompt-caching) to review your code, apply improvements, and measure results.

- [API](https://openai.com/news/?tags=api)
- [2026](https://openai.com/news/?tags=2026)

## Author

OpenAI

## Keep reading

[View all](https://openai.com/news/)

Introducing GPT-6 Sol and Luna — Art card

[Introducing GPT-6 Sol and Luna

ProductSep 22, 2026](https://openai.com/index/introducing-gpt-6-sol-and-luna/)

Reimagining advertising with AI — Card cover

[Reimagining advertising with AI

ProductSep 16, 2026](https://openai.com/index/reimagining-advertising-with-ai/)

How to connect AI usage to business value — art card

[How to connect AI usage to business value

ProductSep 16, 2026](https://openai.com/index/how-to-connect-ai-usage-to-business-value/)
mehrdadrad32
🟧 openai ⭐Better prompt caching for GPT-6
Retrieved article excerpt

Open article · Retrieved 2026-09-22T20:24:16.384144+00:00

September 22, 2026

[Product](https://openai.com/news/product-releases/)

# Better prompt caching for GPT‑6

Higher cache hit rates and new tools to help persistent agents run faster and cost less.

Loading…

Share

GPT‑6 enables persistent agents to work for hours on complex tasks, from refactoring codebases to producing well-researched documents and presentations. The applications behind these agents make a series of API requests that build on one another, often carrying forward the same instructions, tool definitions, and context from earlier turns. OpenAI caches that shared context to reuse computation across requests, reducing response times and giving developers discounts of up to 90% on cached input tokens.

With the GPT‑6 family, we launched an improved prompt caching system that delivers higher cache hit rates by default. We now give cache discounts for eligible shared prefixes reused within a 30-minute window. We’re also introducing new tools to help developers monitor cache performance, diagnose misses, and choose how much of a prompt to cache.

> “OpenAI’s prompt caching plays a critical role in helping GitHub Copilot deliver fast, efficient experiences at scale. Over the past several months, we’ve reduced by more than 50% the share of prompt tokens requiring fresh processing across billions of requests to OpenAI models, relative to our previous baseline. The result is a more efficient inference stack and faster time to first response for developers.”

—Mario Rodriguez, Chief Product Officer

## Monitor caching and diagnose cache misses

The new [Prompt Caching Dashboard⁠(opens in a new window)](https://platform.openai.com/usage?usage_section=prompt-caching) shows how much of your application’s input is served from cache. Track hit rates over time and use the input composition chart to compare cached and uncached tokens. These views help you spot drops in cache hits and evaluate how changes to your application impact caching performance.

Prompt caching dashboard showing cache hit rate, cache performance over time, and input token composition.

When you see an unexpected cache miss, use the [prompt caching diagnostics tool⁠(opens in a new window)](https://developers.openai.com/api/docs/guides/prompt-caching/diagnostics) to understand what happened. Compare a request with a recent response to identify changes to the model, tools, settings, or input that prevented reuse. The estimated number of affected tokens helps you assess the size of the impact and decide how you can optimize your integration to maximize cache hit rates.

```
{
  "prompt_cache_diagnostics": {
    "type": "cache_miss",
    "reason": "tools_changed",
    "comparison_reusable_tokens": 5629,
    "cache_missed_tokens": 5629
  }
}
```

## Optimize caching for your application

**Choose what to cache.** Explicit cache breakpoints let you choose which prompt prefixes to reuse. The refreshed [prompt caching guide⁠(opens in a new window)](https://developers.openai.com/api/docs/guides/prompt-caching) explains how to use them, how long cached prefixes remain eligible, and how changes to tools and inputs affect reuse.

**Adjust reasoning effort without breaking cache.** On GPT‑6 models, you can now [change reasoning effort⁠(opens in a new window)](https://developers.openai.com/api/docs/guides/reasoning?api-mode=responses#change-reasoning-mid-conversation) between responses without breaking cache. Raise effort for a harder task or lower it for a routine follow-up by appending a `configuration_update` while leaving request-level reasoning effort unchanged. This lets you adjust how much reasoning a task needs while preserving reusable context.

**Preserve cache as tools and instructions change.** As your agent’s tool use needs change, keep tool definitions, schemas, and ordering stable so earlier context stays reusable. Use `allowed_tools` to make only the relevant tools callable, or set tool\_choice to none when no tools are needed, instead of removing definitions. Use new developer messages to append new instructions towards the end of the context to override older ones. See our [guidance on managing tool changes⁠(opens in a new window)](https://developers.openai.com/api/docs/guides/prompt-caching#manage-tools-with-append-only-updates).

**Prewarm the cache to reduce latency.** [Prewarming⁠(opens in a new window)](https://developers.openai.com/api/docs/guides/prompt-caching#prewarm-the-cache) prepares known context ahead of time so the model can start responding sooner when a request arrives. For example, an application can prewarm shared instructions, tool definitions, or reference material during startup, before the user asks their first question. This moves processing out of the user’s wait time.

These optional controls build on the engine’s default performance, helping you tailor caching to your workload.

1 of 3

> “OpenAI’s prompt caching diagnostics and dashboard helped us improve cache hit rates by a few percentage points, reducing costs by 20%. We now get alerts when caching breaks unexpectedly and use Codex agents to diagnose the root cause. Explicit breakpoints also let us cache stable context while keeping frequently changing content at the end of the prompt. That’s made it economically viable to fork conversations for background tasks while reusing nearly all of the shared context.”

—Arian Hanifi, Chief Technology Officer

> “For long-running agents like Manus, reliable caching is fundamental to the economics. Working with OpenAI's engineering team, we refined cache breakpoint placement, combined explicit and automatic caching, and used real requests to pinpoint unexpected cache misses. In less than a week, our OpenAI model cache hit rate went from roughly 85% to consistently above 90%, further lowering inference costs in production. The progress came through a series of targeted improvements, with both teams validating the results along the way.”

—Bin Fan, Agent Team Lead

> “Working with OpenAI, we moved our session agents to explicit cache breakpoints. In under a week, cache hit rates on our evaluations rose from 83% to 91%. This meant fewer cache writes and lower inference costs with the same workload; the cache writes fell by roughly two-thirds, and inference costs by 36%.”

—Eugene Mikhantyev, AI Engineer

- Strawberry Browser
- Manus
- Wordsmith

- Strawberry Browser
- Manus
- Wordsmith

## Get started

- Monitor cache hit rates in the [Prompt Caching Dashboard⁠(opens in a new window)](https://platform.openai.com/usage?usage_section=prompt-caching).
- Investigate unexpected misses with the [diagnostics tool⁠(opens in a new window)](https://developers.openai.com/api/docs/guides/prompt-caching/diagnostics).
- Follow the [prompt caching guide⁠(opens in a new window)](https://developers.openai.com/api/docs/guides/prompt-caching) to improve your setup, or [use Codex⁠(opens in a new window)](https://developers.openai.com/api/docs/guides/prompt-caching#how-to-optimize-prompt-caching) to review your code, apply improvements, and measure results.

- [API](https://openai.com/news/?tags=api)
- [2026](https://openai.com/news/?tags=2026)

## Author

OpenAI

## Keep reading

[View all](https://openai.com/news/)

Introducing GPT-6 Sol and Luna — Art card

[Introducing GPT-6 Sol and Luna

ProductSep 22, 2026](https://openai.com/index/introducing-gpt-6-sol-and-luna/)

Reimagining advertising with AI — Card cover

[Reimagining advertising with AI

ProductSep 16, 2026](https://openai.com/index/reimagining-advertising-with-ai/)

How to connect AI usage to business value — art card

[How to connect AI usage to business value

ProductSep 16, 2026](https://openai.com/index/how-to-connect-ai-usage-to-business-value/)
OpenAI——
🟧 hnBetter prompt caching for GPT-6theanonymousone10
🟠 redditOpenAI rolls out upgraded prompt caching and reduced cached input rates for GPT-6
OpenAI
rhiever42427

Interpretation history

Decision trace