2026-10-11 16:36 UTC

api-reliability

band: coolmomentum: stable score: 0.029
temperature history

Episodes (2)

AI Stupid Level founder ionutvi claims observations from 31,352 repeated benchmark measurements show that static LLM scores can miss performance changes over time, potentially requiring ongoing evaluation rather than treating API model names as stable reliability guarantees.
resolvedknownscott: low
Reddit user Temporary_Method6365 reports Sonnet 5.5 buffers all output after a tool_result until message completion โ€” reproduced across the Anthropic API, Bedrock, and OpenRouter with repro and data filed as anthropic-sdk-python issue #1960 โ€” and Anthropic's fix or acknowledgment, or refutation of the repro, settles whether this is a provider-side streaming regression that streaming agent harnesses must work around.
seednovelscott: high