2026-10-11 16:38 UTC

Redditor skeole reports that Qwen3.8-27B on one RTX 3090 sustained a roughly 21-day CUDA-engine development run with about 12 human messages, producing working kernels but no llama.cpp performance win and spending roughly 83 hours on compaction, suggesting local long-running agents are feasible but context maintenance is a major bottleneck.

state: seedheat: lowuncertainty: highlong-running-agents agent-harnesses local-inference context-compactionskeole

Measured heat

now 0 pts/hpeak 0 pts/hcomments 0/hpeers p0momentum: steady1 platformsage 502h
points/hour across evidence ยท reading as of 2026-10-12 02:59:37.977291+11:00 ยท deterministic, not a model opinion

How the heat travelled

09-20 18:26โญ origin directly observedThe bear can dance: Qwen 3.8 27B on one 3090 for 3 weeks
skeole on r/LocalLLaMA
โ€”
09-20 18:26amplified on r/LocalLLaMA ๐Ÿ‘‘reddit.post.1wloora
skeole
peak 182 ยท 39 comments ยท 100% of case engagement
09-20 19:20our radar first saw it ยท +0.9hdiscovery anchor: reddit.post.1wlooraโ€”
pace: p76 vs 1032 stories at the 336h mark (now 502h old) โ€” ahead of epoch-price-of-thought (1.0x), behind minimax-code-open-source (1.0x)

Evidence (1) โ€” โญ canonical anchor

sourceobjectauthorscorecomments
๐ŸŸ  reddit โญThe bear can dance: Qwen 3.8 27B on one 3090 for 3 weeks
LocalLLaMA
skeole18039

Interpretation history

Decision trace