Ullis’s creator claims its RWKV-8 Heron and 1-bit ROSA design can train a 300M-parameter, 32-layer model in about 1.5GB of RAM on a base M1 Mac, making substantial local model training feasible on commodity Apple Silicon.
state: expiredheat: lowuncertainty: highconvergesscott: mediumlocal-training open-models memory-efficient-trainingVlad KalinkinRWKV
What is this?
Ullis is presented as a local model-training project whose creator claims that an RWKV-8 Heron architecture combined with a 1-bit ROSA design can train a 300M-parameter, 32-layer model using roughly 1.5GB of RAM on a base M1 Mac. If reproducible, this would move meaningful model training—not merely quantized inference—onto commodity Apple Silicon. The supplied search snippets discuss local inference, quantization, and Apple unified memory but do not independently verify Ullis’s benchmark, explain ROSA, or firmly establish Vlad Kalinkin’s role.
Why it matters to Scott
The claim converges with Scott’s sovereign, hardware-aware local-model direction by potentially extending commodity-device operation from inference into meaningful training while sharply lowering the cost of experimentation. It could affect what he builds or tests locally, but the benchmark and 1-bit ROSA design remain unverified, and the radar already follows several adjacent memory-efficient consumer-hardware training claims.
dev:concept.hardware-aware-local-inferenceip:framework.sovereign-software-assuranceip:concept.cost-of-cognitionradar:concept.memory-efficient-trainingradar:concept.open-model-trainingradar:concept.apple-siliconradar:gguf-lora-16gb-moe-training
queries asked of Scott's wikis
- low-bit training and optimizer-state compression
- local model training on commodity hardware
- Apple Silicon unified-memory training economics
- open-model sovereignty through local training
- RWKV and memory-efficient sequence architectures
- on-device fine-tuning versus cloud training
Measured heat
no measured readings yet — the hourly heat pass fills this in
How the heat travelled
no chain yet — the hourly chain pass fills this in
Evidence (1) — ⭐ canonical anchor
Interpretation history
2026-09-01T11:38:52Z
The claim has attracted no technical follow-up, reproduction, or evidence of completed, useful training within its observation window, so it no longer merits active tracking.
2026-08-30T11:33:26Z
No new evidence changes the case: it remains a first-party demonstration that training can start within the reported memory footprint, not proof that the model completes training, converges, or achieves useful quality.
2026-08-30T11:28:26Z
grounded: converges/medium — The claim converges with Scott’s sovereign, hardware-aware local-model direction by potentially extending commodity-device operation from inference into meaning
2026-08-30T11:24:49Z
case created — The first-party Show HN describes a concrete low-memory training result that is technically relevant but currently has no corroboration or demonstrated model quality.
Decision trace
- 09-01 21:38expireThe claim has attracted no technical follow-up, reproduction, or evidence of completed, useful training within its observation window, so it no longer merits active tracking.
- 09-01 21:38alert_silentThe only delta is a trivial engagement increase after 48 hours; it adds no evidence about convergence, model quality, or reproducibility and can be archived without briefing Scott.
- 09-01 21:38alert_routeThe only delta is a trivial engagement increase after 48 hours; it adds no evidence about convergence, model quality, or reproducibility and can be archived without briefing Scott.
- 08-30 21:33repriceNo new evidence changes the case: it remains a first-party demonstration that training can start within the reported memory footprint, not proof that the model completes training, converges, or achiev
- 08-30 21:33alert_silentThis is only a legacy-state reevaluation with unchanged engagement and no new technical evidence; the unresolved completion, quality, and reproducibility questions can wait for a normal briefing.
- 08-30 21:33alert_routeThis is only a legacy-state reevaluation with unchanged engagement and no new technical evidence; the unresolved completion, quality, and reproducibility questions can wait for a normal briefing.
- 08-30 21:31alert_silentThe creator provides a specific first-party implementation claim and logs, but the evidence only shows that training was started at a reported 1.5GB footprint—not that a useful 272M-parameter model co
- 08-30 21:31surface_candidateThe creator provides a specific first-party implementation claim and logs, but the evidence only shows that training was started at a reported 1.5GB footprint—not that a useful 272M-parameter model co
- 08-30 21:31alert_routeThe creator provides a specific first-party implementation claim and logs, but the evidence only shows that training was started at a reported 1.5GB footprint—not that a useful 272M-parameter model co
- 08-30 21:28groundThe claim converges with Scott’s sovereign, hardware-aware local-model direction by potentially extending commodity-device operation from inference into meaningful training while sharply lowering the
- 08-30 21:24createThe first-party Show HN describes a concrete low-memory training result that is technically relevant but currently has no corroboration or demonstrated model quality.