2026-10-11 17:11 UTC

Independent benchmarks will determine whether Ornith 1.5’s released 9B, 35B-A3B, and 397B models deliver competitive coding and reasoning quality with practical inference tradeoffs.

state: expiredheat: lowuncertainty: highknownscott: mediumopen-models local-inference coding-agentsOrnith AI

What is this?

Ornith AI released Ornith-1.5, an open model family comprising a 9B dense model and 35B-A3B and 397B mixture-of-experts models, reportedly trained through an end-to-end self-improvement process that generates tasks and task-specific scaffolds. Ornith claims state-of-the-art performance among similarly sized open models and results approaching proprietary frontier systems, particularly on software-engineering benchmarks. However, the supplied sources say the headline 397B results were largely self-reported and not yet independently reproduced, while early analysis indicates mixed performance outside coding, so practical quality and inference tradeoffs remain unsettled.

Why it matters to Scott

Scott already holds the core position in “Evaluation-Driven Development” and “Capability Audit”: vendor claims require repeatable, representative evaluation before adoption. Ornith’s sizes and quantizations could affect his local model-serving and coding-agent comparisons, while the radar already tracks the same 397B model’s practical inference question in “krasis-single-gpu-ornith-397b”; absent independent results, this is a model candidate rather than a changed conclusion.
ip:concept.evaluation-driven-developmentip:concept.capability-auditip:concept.usable-mass-over-unusable-powerdev:concept.hardware-aware-local-inferencedev:concept.trace-backed-agent-comparisondev:project.gamepcradar:krasis-single-gpu-ornith-397bradar:concept.model-evaluationradar:concept.open-modelsradar:concept.coding-modelsradar:concept.local-inferenceradar:concept.quantization
queries asked of Scott's wikis
  • independent evaluation of open coding models
  • local inference economics for dense versus MoE models
  • quantization tradeoffs for coding agents
  • self-improving models and generated training environments
  • SWE-bench versus real coding-agent performance
  • open-weight alternatives to proprietary coding models

Measured heat

no measured readings yet — the hourly heat pass fills this in

How the heat travelled

no chain yet — the hourly chain pass fills this in

Evidence (22) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟠 redditWe have Q3.8 35B at home: 3x new Ornith 1.5 released
LocalLLaMA
AppealSame4367471214
🟧 echo.blog ⭐Ornith's official blog introduces Ornith-1.5 as an open-source family of 397B, 35B, and 9B models using an end-to-end self-improvement loop.Ornith——
🟠 redditOrnith-1.5 (397B [DeepSWE 56], 35B-A3B, 9B)
LocalLLaMA
KokaOP27080
🟧 hnOrnith-1.5: From Self-Scaffolding to Self-ImprovementCommonGuy21373
🟠 redditOrnith-1.5 9B might not be bad at all
LocalLLaMA
zippydazoop4010
🟠 redditWe quantized the new Ornith 1.5 9B and 35B-A3B
LocalLLaMA
Fun-Meaning-64743910
🟠 redditornith9b on mac m3 pro 36gb
LocalLLaMA
hannibal2723
🟠 redditOrnith-1.5-35B-A3B Q4 running 60tk/s on 4070Ti.
LocalLLaMA
Seraphym871316
🟠 redditIf you are wondering why Ornith 1.5 35B A3B with MTP is so slow, this is why
LocalLLaMA
Max-_-Power17536
🟠 redditWhile Everyone Is Excited About Qwen 3.8 27B, Here’s the Reality for a 16GB AMD GPU User
LocalLLaMA
CrowKing63035
🟠 redditOrnith-1.5-35B-A3B-NInfer - 250 tok/s, 5-8k prefill, 5090
LocalLLaMA
koloved6142
🟠 redditIf you want to upgrade from Qwen3.6 27B, but dislike 3.8, give Ornith-1.5-35B-A3B a try.
LocalLLaMA
jinnyjuice037
🟠 redditFixed the MTP head on Ornith1.5 35B A3B. +3% TPS -33% wall clock
LocalLLaMA
frankentriple3522
🟠 redditTielCoder's 22 GB 4-bit quant matches Opus4.6 medium on recent real life coding issues, surpassing KAT-Coder and Nail as strongest and fastest MoE picks.
LocalLLaMA
peculiar-ragdoll287340
🟠 redditQwen-3.8-27B, Nemotron-3.5-Lightning-30B-A3B, Ornith-1.5-35B-A3B, Muse-Glimmer-30B oQ8e comparison
LocalLLaMA
DerTomsn7224
🟠 redditOrnith-1.5-35B-A3B on a Strix Halo iGPU lands 4 problems behind Qwen3.8-27B on 2x 3090s. Full LiveCodeBench v6 numbers.
LocalLLaMA
TrifleHopeful5418012
🟠 reddit35B-A3B tool calling benchmark: Original Qwen vs. KAT Coder, Ornith and Tiel-Coder
LocalLLaMA
OsmanthusBloom6638
🟠 redditOrnith-1.5-35B-A3B on 8 GB VRAM: I think I've found my sweet spot
LocalLLaMA
Elemental_Particle2410
🟠 redditOrnith 1.5 is actually pretty good
LocalLLaMA
deathcom654456
🟠 redditInfinite procedurally generated walking simulator coded entirely by Ornith-1.5-35B-Q4_K_M on an 8 GB RTX 4060
LocalLLaMA
Special_Condition6716725
🟠 redditDid anyone else notice the Ornith 1.5 35B GGUFs got a "silent" update?
LocalLLaMA
miki42422513
🟠 redditDoes anyone have real experience with Ornith-1.5-9B for coding
LocalLLaMA
Barni2751637

Interpretation history

Decision trace