2026-10-11 16:38 UTC

Templar claims stage skipping with compressed cross-stage communication lets its Crucible pipeline-parallel pretraining platform keep healthy workers training through a failed pipeline stage instead of stalling for recovery, which if validated would remove a major availability constraint on distributed training runs.

state: seedheat: lowuncertainty: mediumconvergesscott: mediumai-infrastructure fault-tolerant-training distributed-trainingTemplar

What is this?

Templar (tplr.ai) is building Crucible, a platform for pretraining LLMs across globally distributed, heterogeneous machines rather than a single datacenter cluster. Its approach layers two techniques: SparseLoCo, which lets data-parallel replicas exchange compressed gradient fragments via object storage (Cloudflare R2) instead of direct collectives, so replica failures don't halt the run; and a compressed pipeline construction (based on Subspace Networks, from their heterogeneous low-bandwidth pre-training paper) that compresses activation-sized tensors crossing stage boundaries in pipeline parallelism, since those transfers are the critical path over slow internet links. The specific claim under test here โ€” stage skipping so healthy workers keep training through a failed pipeline *stage* โ€” is only partially covered by the snippets, which document replica-level fault tolerance and cross-stage compression but not the stage-skipping mechanism itself; that part rests on the case's own evidence item.

Why it matters to Scott

Templar's stage-skipping claim independently arrives at Scott's structural answer to failure โ€” design the system so failure doesn't stall it, rather than requesting careful behaviour or replaying checkpoints โ€” the 'Manners vs Physics' / 'Durability Beats Coordination' position, applied in a domain he doesn't own. It also sharpens a distinction the radar is already tracking: SpotWarp's two cases solve the same availability problem via checkpoint-recovery failover, while Crucible claims to train *through* the failed stage, which puts real tension on Scott's own 'checkpoint discipline as minimum shape' doctrine and is worth a dated-receipts note on protocol-level vs recovery-level fault tolerance. Stays medium rather than high because it's training infrastructure, not agent tooling โ€” it extends an argument, it doesn't change what he builds this month.
ip:framework.long-running-agentsip:concept.manners-vs-physicsip:concept.durability-beats-coordinationip:concept.checkpoint-disciplinework:project.chessbrain-netradar:concept.distributed-trainingradar:spotwarp-spot-gpu-failoverradar:spotwarp-vastai-spot-failoverradar:concept.training-efficiency
queries asked of Scott's wikis
  • decentralized collaborative training over the internet (Bittensor-style, consumer GPU swarms)
  • training fault tolerance: checkpoint/recovery vs designing failure out of the protocol
  • compression of activations/gradients as the enabler for low-bandwidth ML systems
  • heterogeneous compute economics: can distributed commodity hardware challenge datacenter training runs
  • object-storage / async message-passing patterns as a coordination substrate (vs direct collectives) โ€” same pattern in agent systems?
  • open frontier training runs and who can credibly claim to run them outside big labs

Measured heat

now 0 pts/hpeak 0 pts/hcomments 0/hpeers p50momentum: steady1 platformsage 456h
points/hour across evidence ยท reading as of 2026-10-12 02:59:37.977291+11:00 ยท deterministic, not a model opinion

How the heat travelled

09-22 15:47โญ origin directly observedSimulating fault tolerance with stage skipping in pipeline-parallel training [R]
covenant_ai on r/MachineLearning
โ€”
09-22 15:47amplified on r/MachineLearning ๐Ÿ‘‘reddit.post.1wnd5ys
covenant_ai
peak 2 ยท 1 comments ยท 100% of case engagement
09-22 16:20our radar first saw it ยท +0.5hdiscovery anchor: reddit.post.1wnd5ysโ€”
pace: p23 vs 1032 stories at the 336h mark (now 456h old) โ€” ahead of aafp-commons-signed-agent-notebook (2.0x), behind agentgate-signed-agent-receipts (0.7x)

Evidence (1) โ€” โญ canonical anchor

sourceobjectauthorscorecomments
๐ŸŸ  reddit โญSimulating fault tolerance with stage skipping in pipeline-parallel training [R]
MachineLearning
covenant_ai21

Interpretation history

Decision trace