2026-10-11 17:12 UTC

Independent reproduction will determine whether the published GGUF-based workflow can LoRA-train Qwen3.6-35B-A3B within 16GB of VRAM using APEX quantization and fused kernels without prohibitive performance or quality tradeoffs.

state: expiredheat: lowuncertainty: highknownscott: mediumgguf-training lora quantization low-vram-training mixture-of-expertswoct0rdho

What is this?

A repository attributed to woct0rdho presents a GGUF-native LoRA workflow that claims to fine-tune Qwen3.6-35B-A3B within 16GB of VRAM using APEX quantization and fused dequantization kernels. The supplied snippets independently show that quantized GGUF variants of this model family can run on 16GB GPUs and that APEX quantizations target usable quality and performance, but they discuss inference rather than LoRA training. They therefore do not independently establish the workflow’s training memory use, speed, stability, or quality tradeoffs; those claims still require reproduction and benchmarks.

Why it matters to Scott

Scott already treats precision, memory pressure, accelerator placement, and compilation as explicit policy in dev:concept.hardware-aware-local-inference, with dev:project.gamepc providing an active consumer-GPU substrate. The claimed 16GB GGUF-native training path is therefore not a new position, but successful reproduction could materially extend his local stack from inference into low-VRAM LoRA training; the radar already tracks the surrounding local-inference and extreme-quantization territory, including this same Qwen model’s optimized execution.
dev:concept.hardware-aware-local-inferencedev:project.gamepcradar:concept.local-inferenceradar:concept.extreme-quantizationradar:ninfer-qwen-5090-throughput
queries asked of Scott's wikis
  • GGUF-native training and fine-tuning
  • low-VRAM LoRA and QLoRA workflows
  • quantized training quality versus memory tradeoffs
  • consumer-GPU fine-tuning economics
  • mixture-of-experts local training
  • fused dequantization kernels

Measured heat

no measured readings yet — the hourly heat pass fills this in

How the heat travelled

no chain yet — the hourly chain pass fills this in

Evidence (3) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟠 redditLoRA over GGUF: Train Qwen3.6-35B-A3B in 16G VRAM
LocalLLaMA
woct0rdho2113
🟧 echo.github ⭐The repository presents a GGUF-native LoRA workflow claimed to fit Qwen3.6-35B-A3B into 16GB VRAM using APEX quantization and fused dequantiwoct0rdho——
🟠 redditLoRA over GGUF: Train DeepSeek-V4-Flash in 90G VRAM
LocalLLaMA
woct0rdho216

Interpretation history

Decision trace