2026-10-11 17:12 UTC

Sori-1B’s developer claims its 1B decoder, trained from scratch exclusively on audio-paired text, grounds responses in audio more strongly than text-pretrained audio-language models while remaining practical for local deployment.

state: expiredheat: lowuncertainty: highconvergesscott: mediumaudio-language-models multimodal local-inferenceSori-1BSeoul National University

What is this?

Sori-1B is presented in the supplied repository/model-card evidence as a 1B-parameter audio-language decoder trained from scratch using only audio-paired text, without language-only pretraining. Its developer claims this produces stronger audio grounding than text-pretrained audio-language models while keeping the model small enough for local deployment. The case associates it with Seoul National University, but the supplied search snippets do not independently establish the developer’s identity, affiliation, benchmark results, or deployment performance.

Why it matters to Scott

Sori-1B’s paired-audio-only training claim independently supports Scott’s training-distribution argument: modality grounding should follow the data distribution rather than remain inherently text-first. It is also a concrete candidate for his local speech laboratory, but the supplied evidence does not establish benchmark superiority or practical local performance, so it warrants testing rather than changing his position yet.
ip:concept.training-distribution-biasip:source.text-is-the-models-home-turfdev:project.audiodev:concept.hardware-aware-local-inferenceradar:concept.multimodal-modelsradar:concept.model-trainingradar:concept.local-audio-inferenceradar:audio-cpp-07-local-audio-arena
queries asked of Scott's wikis
  • audio-native models versus text-pretrained multimodal models
  • modality grounding through paired-data-only training
  • small multimodal models for local inference
  • local voice-agent latency privacy and economics
  • audio-language model evaluation and grounding benchmarks
  • from-scratch specialist models versus pretrained foundation models

Measured heat

no measured readings yet — the hourly heat pass fills this in

How the heat travelled

no chain yet — the hourly chain pass fills this in

Evidence (2) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟠 redditsnkii/Sori-1B: Audio-Grounded LM Trained From Scratch (No Text-Only Pretraining)
LocalLLaMA
Balance-323
🟧 echo.other ⭐The primary model card says Sori-1B was “trained from scratch, with no language-only pretraining,” and that every predicted text token was pSeonuk Kim——

Interpretation history

Decision trace