2026-10-11 17:12 UTC

Independent deployments will determine whether Yeschef can reliably dispatch Claude Code tasks across pooled LAN-hosted Ollama workers with useful throughput, task quality, and operational simplicity.

state: expiredheat: lowuncertainty: highknownscott: mediumcoding-agents local-inference agent-harnesses distributed-inferenceYeschefClaude CodeOllamaLabs Community

What is this?

Yeschef appears to be an agent harness that dispatches Claude Code tasks to a pool of Ollama workers on three LAN-connected NUCs, with the evidence title claiming aggregate throughput of 627 tokens per second. The supplied snippets establish that Claude Code can use Ollama-compatible backends and that developers are testing strong-model planner/local-model worker patterns to reduce costs, but they do not independently document Yeschef’s architecture or verify its throughput, task quality, reliability, or operational simplicity. Those claims therefore remain a single reported deployment awaiting replication.

Why it matters to Scott

The radar already tracks the core frontier-orchestrator/cheaper-worker claim in radar:multi-model-orchestrator-worker-agents, while commodity-machine pooling is covered by the Cascadia distributed-inference cases. Yeschef still bears directly on Scott’s active Ollama/LiteLLM routing stack and could provide a useful trace-backed replication fixture for testing whether headline throughput survives task-quality, reliability, and operating-complexity measurement, but the single reported deployment adds no validated result yet.
ip:framework.micro-agents-architectureip:concept.model-barbellip:source.subagents-speed-accuracy-ebookdev:technology.litellmdev:concept.task-aware-model-routingdev:technology.ollamadev:project.gamepcdev:concept.trace-backed-agent-comparisonradar:multi-model-orchestrator-worker-agentsradar:cascadia-distributed-intel-inferenceradar:concept.distributed-inferenceradar:concept.model-routing
queries asked of Scott's wikis
  • planner-worker architectures for coding agents
  • pooled local inference orchestration
  • strong-model validation of local agent workers
  • coding-agent throughput versus task quality
  • LAN inference scheduling and fault tolerance
  • operational simplicity of hybrid local-cloud agents

Measured heat

no measured readings yet — the hourly heat pass fills this in

How the heat travelled

no chain yet — the hourly chain pass fills this in

Evidence (1) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟧 hn ⭐Yeschef: Claude Code dispatches work to Ollama on my LAN (627 tok/s on 3 NUCs)hxrace31

Interpretation history

Decision trace