2026-10-11 17:11 UTC

Independent benchmarks will determine whether NVIDIA Nemotron 3.5 Lightning 30B-A3B delivers a practically useful quality-throughput tradeoff for local sparse-MoE inference.

state: expiredheat: lowuncertainty: mediumknownscott: lowopen-models local-inferenceNVIDIA

What is this?

The case concerns a purported NVIDIA open model, Nemotron 3.5 Lightning 30B-A3B, and the claim that independent benchmarking is needed to establish its real quality-versus-throughput tradeoff for local sparse-MoE inference. The supplied search results are unrelated results about the newspaper The Independent; they do not verify the model’s architecture, release details, benchmark performance, licensing, hardware requirements, or practical local-inference utility. Beyond the evidence titles naming NVIDIA and a Hugging Face repository, the event is therefore not substantively grounded by these snippets.

Why it matters to Scott

Scott already holds the relevant position in Capability Audit and Hardware-aware local inference: model utility must be established through representative, hardware-specific evaluation rather than vendor claims. The model could eventually matter to his gamepc local-model stack, but the supplied evidence does not verify its architecture, requirements, performance, or even substantive release details, so it currently adds only an ungrounded benchmark candidate.
ip:concept.capability-auditdev:concept.hardware-aware-local-inferencedev:project.gamepcradar:person.nvidiaradar:concept.mixture-of-expertsradar:concept.local-inferenceradar:concept.model-evaluation
queries asked of Scott's wikis
  • sparse MoE local inference economics
  • quality-throughput benchmarks for local models
  • active-parameter count versus deployment cost
  • open-weight model evaluation framework
  • NVIDIA models in local inference stacks
  • hardware constraints for mixture-of-experts inference

Measured heat

no measured readings yet — the hourly heat pass fills this in

How the heat travelled

no chain yet — the hourly chain pass fills this in

Evidence (8) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟠 reddit ⭐nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16 · Hugging Face
LocalLLaMA
coder543572170
🟧 hnNvidia Nemotron 3.5 Lightningbeklein12132
🟠 redditnvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16
LocalLLaMA
TheRealMasonMac9533
🟧 hnNvidia Nemotron 3.5 Lightning and NeMo Switchyarddroidjj261136
🟠 redditMoE task time comparison
LocalLLaMA
parepeg26
🟠 redditTested Nemotron 3.5 Lightning locally on coding, Hermes Agent and agentic work
LocalLLaMA
curiousily_1516
🟧 hnNemotron-3.5-lightning-30B-a3B Model by Nvidia for use on 1 GPUjanandonly10
🟠 redditNemotron 3.5 Lightning 30B-A3B: W4A16 vs IQ4_XS on the same RTX 3090: near-parity at B1, ~4.5× throughput by B16
LocalLLaMA
mitchins-au33

Interpretation history

Decision trace