Ornith-1.0 is an open-source model family from DeepReinforce aimed at agentic coding; its publisher claims the 397B model scores 77.5 on Terminal-Bench 2.1 and 82.4 on SWE-Bench Verified. The case’s evidence title reports running a Q4 quantization on one 96GB RTX PRO 6000 Blackwell at 2,354 tokens/s prefill and roughly 20–24 tokens/s decode, allegedly using Krasis to stream experts from system RAM. However, the supplied search snippets neither identify Krasis nor independently verify that hardware setup, streaming mechanism, or throughput; one comparison site lists Ornith as unranked and notes limited comparable benchmark coverage.
No intersection found in Scott’s wikis, and the radar does not already track Krasis, Ornith, or this claimed expert-streaming result. The case is broadly topical to local inference, but the supplied hits establish no Scott-specific claim, project impact, or publishing opportunity.
queries asked of Scott's wikis
- dynamic MoE expert streaming from system RAM
- single-GPU local inference and expert offload
- Q4 quantization quality versus interactive throughput
- local frontier coding models versus hosted coding agents
- memory bandwidth economics of RAM-offloaded inference
- model sovereignty through workstation-scale inference
2026-08-02T13:24:17Z
Repeated reobservation has produced only adjacent MoE discussion, with no independent benchmark or implementation testing Krasis or its Ornith throughput claim. The signal has faded beyond its useful horizon and should return only if a direct reproduction appears.
2026-08-02T12:22:27Z
The attached material still offers no independent reproduction of Krasis, its RAM-streaming mechanism, or the claimed Ornith throughput; adjacent MoE economics evidence does not corroborate the specific result. The case is now a cold, isolated builder claim best revisited only if a direct benchmark or implementation appears.
2026-08-02T11:25:42Z
The newly attached activity still provides no independent test of Krasis, its RAM-streaming mechanism, or the claimed Ornith throughput. The case remains an isolated builder report, and repeated adjacent MoE evidence no longer warrants hourly review.
2026-08-02T10:23:17Z
The attached evidence remains adjacent evidence for large-MoE economics, not an independent test of Krasis, its RAM-streaming mechanism, or the claimed Ornith throughput. The case’s meaning is unchanged: an isolated builder result awaiting direct reproduction, with further engagement adding only repetitive amplification.
2026-08-02T09:22:52Z
No direct benchmark or independent reproduction of Krasis, its RAM-streamed expert offload, or the claimed Ornith throughput has appeared. The broader AMD MoE result remains adjacent rather than corroborating, so further engagement is repetitive amplification of an isolated builder claim.
2026-08-02T08:22:38Z
The attached evidence still concerns broader large-MoE economics rather than independently testing Krasis, its RAM-streaming mechanism, or the claimed Ornith throughput. This is repetitive amplification around an isolated builder result, so the case remains cold pending a direct reproduction or benchmark.
2026-08-02T07:21:36Z
The added material still does not benchmark Krasis, reproduce its RAM-streaming mechanism, or verify Ornith’s claimed single-GPU throughput. Broader MoE economics are increasingly plausible, but the core result remains an isolated builder claim.
2026-08-02T06:21:29Z
The attached AMD result supports only the broader feasibility of economical large-MoE inference and does not independently reproduce Krasis, its RAM-streaming mechanism, or the claimed Ornith throughput. With no substantive new verification, the case remains a low-heat builder claim awaiting benchmarks.
2026-08-02T05:27:18Z
The AMD MoE report modestly supports the broader plausibility of economical large-MoE inference, but it neither reproduces Krasis nor tests Ornith’s claimed single-GPU throughput, leaving the core hypothesis uncorroborated.
2026-08-02T05:21:17Z
evidence attached: hn.story.49141073 — The report is relevant independent evidence that large MoE inference can be made economical on AMD hardware, though it does not directly test the existing Ornith single-GPU hypothesis.
2026-07-27T23:25:51Z
The newly attached activity still supplies no independent reproduction, benchmark, or implementation evidence. The specific single-GPU throughput claim remains unverified builder testimony, and further engagement alone does not change its meaning.
2026-07-27T16:28:23Z
Additional activity still provides no independent benchmark, implementation, or technical verification; this remains a single builder-reported result despite the hot local-inference neighborhood.
2026-07-27T15:23:55Z
The lone additional comment adds no independent benchmark, implementation, or technical verification, so the claimed expert-streaming throughput remains a builder-reported result awaiting reproduction.
2026-07-27T14:24:47Z
grounded: novel/none — No intersection found in Scott’s wikis, and the radar does not already track Krasis, Ornith, or this claimed expert-streaming result. The case is broadly topica
2026-07-27T14:22:07Z
case created — The builder reports a concrete single-GPU result using a distinct RAM-based expert-streaming runtime that merits independent reproduction.