2026-10-11 17:11 UTC

Independent benchmarks will determine whether Krasis can serve the 397B-parameter Ornith MoE interactively at Q4 on a single 96GB workstation GPU by dynamically streaming experts from system RAM while sustaining roughly 20–24 tokens per second.

state: expiredheat: lowuncertainty: highnovelscott: nonemoe-streaming local-inference expert-offloadKrasisOrnith

What is this?

Ornith-1.0 is an open-source model family from DeepReinforce aimed at agentic coding; its publisher claims the 397B model scores 77.5 on Terminal-Bench 2.1 and 82.4 on SWE-Bench Verified. The case’s evidence title reports running a Q4 quantization on one 96GB RTX PRO 6000 Blackwell at 2,354 tokens/s prefill and roughly 20–24 tokens/s decode, allegedly using Krasis to stream experts from system RAM. However, the supplied search snippets neither identify Krasis nor independently verify that hardware setup, streaming mechanism, or throughput; one comparison site lists Ornith as unranked and notes limited comparable benchmark coverage.

Why it matters to Scott

No intersection found in Scott’s wikis, and the radar does not already track Krasis, Ornith, or this claimed expert-streaming result. The case is broadly topical to local inference, but the supplied hits establish no Scott-specific claim, project impact, or publishing opportunity.
queries asked of Scott's wikis
  • dynamic MoE expert streaming from system RAM
  • single-GPU local inference and expert offload
  • Q4 quantization quality versus interactive throughput
  • local frontier coding models versus hosted coding agents
  • memory bandwidth economics of RAM-offloaded inference
  • model sovereignty through workstation-scale inference

Measured heat

no measured readings yet — the hourly heat pass fills this in

How the heat travelled

no chain yet — the hourly chain pass fills this in

Evidence (2) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟠 reddit ⭐Ornith-397B running at Q4 on a single RTX PRO 6000 Blackwell 96GB - 2,354 tok/s prefill, ~20–24 tok/s decode
LocalLLaMA
mrstoatey1923
🟧 hnRunning Kimi K3 on MI355X at Better Performance per Dollar Than B300ilreb19092

Interpretation history

Decision trace