2026-10-11 16:37 UTC

moe

band: coolmomentum: stable score: 0.001
temperature history

Episodes (2)

Independent benchmarks will determine whether dsv4-streaming can run the 284B DeepSeek V4 Flash checkpoint on 64GB Apple Silicon with near-lossless quality and practically useful expert-streaming throughput.
expiredknownscott: medium
The fork author claims selectively placing frequently used MoE experts in VRAM raises llama.cpp generation throughput from 20 to 30 tokens per second on partially offloaded coding workloads, potentially improving local inference on memory-constrained GPUs.
resolvedknownscott: medium

Trajectory notes