2026-10-11 17:09 UTC

agent-distillation

band: coolmomentum: stable score: 0.13
temperature history

Episodes (2)

Ankit Sonthalia and coauthors introduce BOTTLED, a benchmark where LLM agents must convert general capabilities into cheap task-specific artifacts ('bottling'), finding that zero-shot performance doesn't predict bottling success but successful bottling can retain ~82% performance at 657x lower cost.
seedconvergesscott: high
Independent evaluations will determine whether World Model Optimizer can route repetitive agent tasks to trace-distilled smaller models at roughly half the cost of frontier-only serving without material quality loss.
expirednovelscott: none