2026-10-11 17:09 UTC

sparse-models

band: coolmomentum: stable score: 0.016
temperature history

Episodes (2)

Independent benchmarks will determine whether AFM3โ€™s prompt-conditioned expert and layer activation can substantially reduce local-inference memory bandwidth while preserving model quality.
expiredconvergesscott: medium
Extension-Bid-639 claims a build combining quantization, expert caching, host-RAM offload, and multi-token prediction raises full-261K-context Qwen3.8-Flash-Next decode throughput from 25โ€“29 to 37โ€“41 tokens per second on two RTX 3090 GPUs, potentially making long-context local coding inference practical on commodity multi-GPU systems.
resolvedknownscott: medium

Trajectory notes