2026-10-11 16:38 UTC

Anthropic's official Opus 5.5 prompting guide documents harness-breaking mechanics — thinking can no longer be disabled, progress updates arrive as often-empty thinking blocks, and some turns end with plain text instead of tool calls that unattended agent loops mistake for task completion — forcing coding-agent harnesses to adopt the guide's mitigation patterns to keep Opus 5.5 agents running.

state: corroboratedheat: lowuncertainty: mediumconvergesscott: highagent-harnesses prompt-engineering claude-opus-5-5Anthropic
Surfaced 2026-09-28T08:33:41Z — Anthropic's official guide documents Opus 5.5 behavior and API changes per the retrieved page: unlike Opus 5, 'Claude Opus 5.5 doesn't' acce — Seed → corroborated: the harness-breaking mechanics are now multiply attested (retrieved first-party docs, third-party migration writeups tallying the same four breakers, substantive two-platform discussion), so the open question shifts from what the guide documents to whether harnesses actually adopt its mitigations — none observed yet. Heat to high on a 99th-percentile, accelerating 3-platform spread; the HN attach and Reddit velocity are repeated coverage of the same doc, so this is attention, not a new fact — no material change.

What is this?

Claude Opus 5.5 is Anthropic's flagship agentic-coding model, released September 22, 2026 at $4/$20 per million tokens with a 1M-token context window and adaptive thinking that is always on. Anthropic's official model docs and prompting guide document four breaking API changes for code built on Opus 5 — thinking can no longer be disabled, forced tool_choice (any/tool) is rejected, thinking blocks are tied to model and conversation (append-only history), and the old computer-use tool is replaced — plus a non-erroring response-shape change: narration between tool calls now arrives in thinking blocks whose text is empty at the default display setting, so clients that render only text blocks go silent mid-run. The guide specifically warns that on long unattended tasks Opus 5.5 may end a turn with plain text (stop_reason 'end_turn') instead of a tool call, and agent loops that treat any text-only turn as task completion will stop early; prescribed mitigations include aligning prompt and harness on what counts as finished, the display:'updates' beta to surface progress notes, and appending turn-scoped reminder messages after consecutive silent tool-calling steps. The supplied sources are consistent — Anthropic's own docs are corroborated by third-party migration writeups tallying the same four breakers — and nothing in them contradicts the case hypothesis.

Why it matters to Scott

Anthropic's own docs now codify Scott's closer doctrine vendor-side: Opus 5.5 ends unattended turns with plain text that loops misread as done, so the guide prescribes harness-agreed finish criteria and turn-scoped reminder messages — exactly his Five-Surface Loop Anatomy claim that a stop condition internal to the acting agent isn't a true closer and the premature-'done' discipline of Ask Yourself If You're Finished (a dated-receipts opportunity: his ebooks argued this before the vendor documented it). The breakers also hard-code his provider-bound-reasoning-continuity position as API law and land directly on his live stack — dev:ask's recurse-until-answer loop, proposal's Claude log-salvage/resume, and the LiteLLM gateway all need the guide's mitigations.
ip:concept.adversarial-closerip:framework.five-surface-loop-anatomyip:source.ask-yourself-if-you-re-finished-cron-as-the-poor-man-s-orchestrator-ebookip:framework.heartbeat-supervisory-programdev:concept.provider-bound-reasoning-continuitydev:project.askdev:project.proposaldev:technology.litellmradar:concept.agent-harnessesradar:concept.long-running-agentsradar:concept.tool-callingradar:concept.model-migrationradar:concept.reasoning-tokensradar:coding-agent-self-report-failure-blindnessradar:frontier-api-zero-output-voidsradar:ballast-goal-completion-harness
queries asked of Scott's wikis
  • agent loop completion detection text-only turn stop condition
  • harness migration model upgrade breaking API changes
  • thinking block parsing display streaming client rendering
  • agent progress updates narration long-running runs
  • forced tool_choice structured outputs fallback patterns
  • unattended agent reliability continuation heuristics

Measured heat

now 0 pts/hpeak 198 pts/hcomments 0/hpeers p14momentum: steady3 platformsage 328h
points/hour across evidence · reading as of 2026-10-12 02:59:37.977291+11:00 · deterministic, not a model opinion

How the heat travelled

09-28 01:25 (minted)⭐ origin echo-reconstructedAnthropic's official guide documents Opus 5.5 behavior and API changes per the retrieved page: unlike Opus 5, 'Claude Opus 5.5 doesn't' acce
Anthropic on blog (echo) · attributed from reddit.post.1ws08mu · published time unknown
—
09-28 00:28first on r/ClaudeAI · published · lag ?Claude Opus 5.5 official prompting guide
BuffaloConscious7919
—
09-28 07:33first on hacker news · published · lag ?Prompting Claude Opus 5.5
Michelangelo11
—
09-28 00:28amplified on r/ClaudeAI 👑reddit.post.1ws08mu
BuffaloConscious7919
peak 1947 · 95 comments · 72% of case engagement
09-28 07:33amplified on hacker newshn.story.49874728
Michelangelo11
peak 207 · 227 comments · 28% of case engagement
10-01 18:20amplified on r/ClaudeAIreddit.post.1wv714p
Remarkable_Rock5845
peak 18 · 6 comments · 1% of case engagement
09-28 01:20our radar first saw it · lag ?discovery anchor: reddit.post.1ws08mu—
09-28 08:28reached heat=high · lag ? · via ledger——
pace: p96 vs 1188 stories at the 168h mark (now 328h old) — ahead of opus55-video-as-code (1.0x), behind sonnet-55-release-economics (1.0x)

Evidence (4) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟠 redditClaude Opus 5.5 official prompting guide
ClaudeAI
Retrieved article excerpt

Open article · Retrieved 2026-09-28T01:24:21.417913+00:00

[Best practices](https://platform.claude.com/docs/en/about-claude/use-case-guides/overview)Prompt engineering

# Prompting Claude Opus 5.5

Copy page



Behavioral differences from Claude Opus 5 and the prompting and harness patterns that address them: effort calibration, thinking behavior in API integrations and chat, progress updates, unattended and multiagent tasks, safeguard refusals, frontend design, complex visual inputs, multi-app workflows, and pasted text in user messages.

Copy page



This guide covers the prompting patterns specific to Claude Opus 5.5. For the model's capabilities and API changes, see [What's new in Claude Opus 5.5](https://platform.claude.com/docs/en/models/opus-5-5/whats-new-opus-5-5). For techniques that apply across all current Claude models, see [Prompting best practices](https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/claude-prompting-best-practices).

Claude Opus 5.5 generates output tokens more than 30 percent faster than Claude Opus 5 and tends to finish the same task with fewer tokens. Existing Claude Opus 5 prompts should perform well without changes, and the patterns in [Prompting Claude Opus 5](https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/prompting-claude-opus-5) remain a reasonable starting point. Start with the section that matches what you observe:

- Unsure which effort level to run, or turns run longer and cost more than they did on Claude Opus 5: [Calibrate effort](https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/prompting-claude-opus-5-5#calibrate-effort)
- Your Claude Opus 5 integration ran with thinking disabled: [Prompts written for thinking disabled](https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/prompting-claude-opus-5-5#prompts-written-for-thinking-disabled)
- An unattended agent stops partway through a long task after reporting progress: [Unattended agentic runs](https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/prompting-claude-opus-5-5#unattended-agentic-runs)
- Requests return `stop_reason: "refusal"`: [Safeguard refusals](https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/prompting-claude-opus-5-5#safeguard-refusals)
- Long agentic turns look silent, or you want updates at predictable points: [User-facing progress updates](https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/prompting-claude-opus-5-5#user-facing-progress-updates)
- An agent that works across several connected apps misses information the task didn't point to: [Explore context in multi-app workflows](https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/prompting-claude-opus-5-5#explore-context-in-multi-app-workflows)
- You run a team of agents and want it to finish sooner: [Time signals for multiagent harnesses](https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/prompting-claude-opus-5-5#time-signals-for-multi-agent-harnesses)
- Replies in a chat application start slowly because the model thinks at length first: [Thinking instructions in chat system prompts](https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/prompting-claude-opus-5-5#thinking-instructions-in-chat-system-prompts)
- The model follows instructions that arrived inside text a user pasted: [Mark pasted text in user messages](https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/prompting-claude-opus-5-5#mark-pasted-text-in-user-messages)
- Answers about dense charts, diagrams, or screenshots miss detail: [Tools for complex visual inputs](https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/prompting-claude-opus-5-5#tools-for-complex-visual-inputs)
- Frontend output looks generic: [Frontend design defaults](https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/prompting-claude-opus-5-5#frontend-design-defaults)



For the four breaking API changes when migrating from Claude Opus 5, see the [migration guide](https://platform.claude.com/docs/en/models/opus-5-5/migration-guide#migrating-from-claude-opus-5).

## Capabilities relevant to prompting

The capabilities that matter most for prompting are:

- **Agentic coding and code review:** The model is strongest on multistep work in a real repository, such as carrying a change through a large code base until its tests pass. In Anthropic's testing, at its default `medium` effort the model matched or beat Claude Opus 5 at `high` effort on such tasks, in fewer steps and with fewer tokens. It also sustains long-running autonomous work better than Claude Opus 5, such as multi-hour audits and migrations of large code bases run end to end with parallel subagents and little oversight. Early testers also reported stronger code review, with more bugs caught than on Claude Opus 5 and fewer false alarms, and it explains its changes in plain language.
- **Knowledge work:** The model is much less likely to state an incorrect figure or cite the wrong source. It's better at financial modeling tasks, such as building a financial model and one-page summary for a transaction or finding and fixing errors in a valuation workbook, and it catches details that are easy to miss in large inputs, such as a date in a long planning thread that falls on the wrong weekday or a chart in a slide deck that doesn't match the underlying figures. The spreadsheets, slides, and documents it produces need less editing before you share them.
- **Communication:** Its reports on agentic work, both the updates while it works and the summary when it finishes, say plainly what it did, what it found, and what it needs from you. See [User-facing progress updates](https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/prompting-claude-opus-5-5#user-facing-progress-updates).
- **Charts, diagrams, screenshots, and computer use:** The model reads visual material more accurately than Claude Opus 5 without extra tooling: in Anthropic's testing, even at its lowest effort setting it read values off dense charts more accurately than Claude Opus 5 did at its highest, using a small fraction of the output tokens. It is better, too, where meaning depends on position rather than text: which boxes an arrow connects in a flowchart, what changed between two versions of a diagram, or exactly when a meeting starts and ends in a calendar screenshot. It's also more reliable at computer use, where it operates applications from screenshots over many steps: at its default effort it matched the success rate that Claude Opus 5 reached only at a much higher effort setting. See [Tools for complex visual inputs](https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/prompting-claude-opus-5-5#tools-for-complex-visual-inputs).

## Calibrate effort

[Effort](https://platform.claude.com/docs/en/build-with-claude/effort) is the main control for how much Claude Opus 5.5 thinks, and because thinking is always on, it's the first setting to adjust when trading off intelligence, latency, and cost. Start at `medium`, the default on Claude Opus 5.5 (Claude Opus 5 defaults to `high`), set it explicitly, and test several levels against your own evals rather than carrying over the setting you used on Claude Opus 5. Effort level names don't correspond to the same amount of thinking across models: in Anthropic's testing, Claude Opus 5.5 at `medium` matches or exceeds Claude Opus 5 at `high` on coding and knowledge-work evaluations, and on several coding evaluations `low` comes close to it at much lower cost. See [Recommended effort levels for Claude Opus 5.5](https://platform.claude.com/docs/en/build-with-claude/effort#recommended-effort-levels-for-claude-opus-5-5).

At a given level, Claude Opus 5.5 tends to think more per turn than Claude Opus 5, especially at `xhigh` and `max`. If you keep the `effort` value you set for Claude Opus 5, expect longer turns and more output tokens. Three adjustments help:

- Set `max_tokens` high enough to leave room for the model's thinking tokens and the reply. Thinking counts toward `max_tokens` even when thinking content isn't returned to you, so a limit sized for Claude Opus 5 with thinking off can cut replies off. For the long turns that agentic coding can produce, a `max_tokens` of 128,000, the model's maximum, has worked well in Anthropic's testing.
- Reserve `xhigh` and `max` for work where you've measured a quality gain.
- To get less thinking, lower the effort level first. Lowering effort reduces thinking, and with it cost and latency, more reliably than prompt instructions do.

Changing the top-level `effort` value between requests invalidates the prompt cache. To run individual turns at a different level, use a [per-message effort change](https://platform.claude.com/docs/en/build-with-claude/effort#change-effort-mid-conversation-beta) (beta) instead, which keeps the cache.

## Prompts written for thinking disabled

Claude Opus 5 accepts `thinking: {"type": "disabled"}` at `high` effort or below; Claude Opus 5.5 doesn't, and the [migration guide](https://platform.claude.com/docs/en/models/opus-5-5/migration-guide#migrating-from-claude-opus-5) covers the request change. If your Claude Opus 5 integration ran with thinking disabled, four changes go with it:

- **Start at `low` effort and measure.** At `low` the model keeps its thinking short. How often it skips thinking altogether depends on your prompts, so measure latency and quality on your own traffic and move to `medium` if quality drops. If time to first token still matters after that, a system prompt line such as "Answer directly without deliberating." can reduce thinking further; measure quality when you add it, because less thinking can lower it.
- **Remove instructions that stood in for thinking.** If your prompt asked the model to write out its reasoning in the response as a substitute for thinking, remove that instruction and read the reasoning from [summarized thinking](https://platform.claude.com/docs/en/build-with-claude/thinking#summarized-thinking) blocks instead (`display: "summarized"`); a prompt that pushes the model to reproduce its reasoning in the response text can be declined with the `reasoning_extraction` [refusal category](https://platform.claude.com/docs/en/build-with-claude/refusals-and-fallback#refusal-response).
- **Re-test the thinking-disabled mitigations.** [Running with thinking disabled](https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/prompting-claude-opus-5#running-with-thinking-disabled) recommends a combined instruction (permission to speak before a tool call, what to do when no tool fits, no internal tags) and removing any rule that tells the model not to think. Both address artifacts that appear on Claude Opus 5 only when thinking is disabled. With thinking always on, check whether you still need the instruction, and remove the no-thinking rule either way.
- **Read the response by block type.** Check each block's type instead of assuming the first content block is text: a response may or may not begin with a `thinking` block, whose `thinking` field is empty under the default `display: "omitted"`.

## Unattended agentic runs

On long tasks with several parts, Claude Opus 5.5 keeps the user updated as it works, and some of those updates end the turn with text rather than a tool call ([`stop_reason: "end_turn"`](https://platform.claude.com/docs/en/build-with-claude/handling-stop-reasons#end-turn)). An unattended agent loop that treats such a turn as the end of the task stops running there. A few harness and prompt changes help it keep running.

Treat a text-only end of turn as a report rather than as proof the task is done. Keep the task's parts in a checklist the model updates, such as a to-do tool or a file. If a turn ends with items still open and no blocker stated, send a short user message naming them, like the following one. You can also state th
BuffaloConscious7919194795
🟧 echo.blog ⭐Anthropic's official guide documents Opus 5.5 behavior and API changes per the retrieved page: unlike Opus 5, 'Claude Opus 5.5 doesn't' acceAnthropic——
🟧 hnPrompting Claude Opus 5.5Michelangelo11207227
🟠 redditClaude answering-in-its-thoughts bug make new models unusable
ClaudeAI
Remarkable_Rock5845186

Interpretation history

Decision trace