A repository attributed to woct0rdho presents a GGUF-native LoRA workflow that claims to fine-tune Qwen3.6-35B-A3B within 16GB of VRAM using APEX quantization and fused dequantization kernels. The supplied snippets independently show that quantized GGUF variants of this model family can run on 16GB GPUs and that APEX quantizations target usable quality and performance, but they discuss inference rather than LoRA training. They therefore do not independently establish the workflow’s training memory use, speed, stability, or quality tradeoffs; those claims still require reproduction and benchmarks.
Scott already treats precision, memory pressure, accelerator placement, and compilation as explicit policy in dev:concept.hardware-aware-local-inference, with dev:project.gamepc providing an active consumer-GPU substrate. The claimed 16GB GGUF-native training path is therefore not a new position, but successful reproduction could materially extend his local stack from inference into low-VRAM LoRA training; the radar already tracks the surrounding local-inference and extreme-quantization territory, including this same Qwen model’s optimized execution.
dev:concept.hardware-aware-local-inferencedev:project.gamepcradar:concept.local-inferenceradar:concept.extreme-quantizationradar:ninfer-qwen-5090-throughput
queries asked of Scott's wikis
- GGUF-native training and fine-tuning
- low-VRAM LoRA and QLoRA workflows
- quantized training quality versus memory tradeoffs
- consumer-GPU fine-tuning economics
- mixture-of-experts local training
- fused dequantization kernels
2026-08-07T19:35:00Z
After repeated checks, no third-party reproduction or benchmark has emerged, and the latest trigger is staleness rather than evidence. The episode has faded without resolving the technical claim and can be rediscovered if an external implementation appears.
2026-08-05T06:29:17Z
The nominal attachment contains no identifiable new evidence, leaving the workflow dependent on same-author claims without independent memory, throughput, stability, or quality measurements. Repeated empty triggers are noise; keep the case cold until a third-party reproduction appears.
2026-08-05T05:21:54Z
The refreshed discussion is modest amplification rather than independent validation; no external reproduction or measurements change the specific 16GB workflow’s credibility. Keep the case cold and shift to a weekly cadence unless concrete third-party results appear.
2026-07-29T04:23:54Z
The nominal evidence trigger contains no identifiable new material or independent reproduction; the case still rests on one author’s related implementations. Repeated empty reobservations are noise, so revisit only when external memory, throughput, stability, or quality results appear.
2026-07-29T00:22:25Z
No identifiable new evidence accompanies the trigger; the case still rests on same-author implementations rather than an independent reproduction of the 16GB workflow. Repeated attachment and engagement events are noise until external memory, throughput, stability, and quality measurements appear.
2026-07-28T22:25:05Z
The trigger adds no identifiable evidence beyond the same-author implementations already incorporated. Independent reproduction and measurements of memory, throughput, stability, and training quality remain absent, so the case stays cold pending external validation.
2026-07-28T21:26:16Z
The attachment adds nothing beyond the already incorporated same-author follow-up, so the specific 16GB Qwen workflow still lacks independent reproduction or measurements of speed, stability, and quality. Repeated reobservation is now noise; revisit only on external implementation evidence.
2026-07-28T18:26:59Z
No new independent evidence has appeared since the same-author 284B follow-up was already incorporated. The specific 16GB Qwen workflow remains unvalidated on memory, speed, stability, and output quality, so further engagement or reobservation does not change the case.
2026-07-28T17:28:04Z
The same author’s 284B MoE follow-up indicates an actively expanding implementation and makes the underlying GGUF-training approach more plausible, warranting watching rather than seed status. It still provides no independent reproduction or quality/performance validation of the specific 16GB Qwen workflow.
2026-07-28T17:21:33Z
evidence attached: reddit.post.1v941oh — A concrete follow-up implementation shows GGUF-based LoRA training scaling to a 284B MoE in 90 GiB VRAM, materially contextualizing the broader low-VRAM training technique.
2026-07-25T23:21:39Z
The small engagement increase is repetitive amplification, not validation; the workflow remains a single-author claim with no reproduction, benchmark, or quality evidence. Keep it cold and revisit only if an external implementation report appears.
2026-07-22T23:21:11Z
The latest trigger adds no substantive evidence beyond the author’s original claim; independent reproduction, training benchmarks, and quality measurements remain absent. Repeated engagement updates are not changing the case’s meaning, so it should stay cold pending an actual implementation report.
2026-07-21T12:23:47Z
The newly attached material adds no independent reproduction, benchmark, or quality evidence beyond the author’s original claim. The case remains a technically relevant but unvalidated implementation awaiting external testing.
2026-07-21T10:21:22Z
No independent reproduction or benchmark has appeared; the only change is flat-to-declining engagement around the original author claim. The case remains technically relevant but uncorroborated, with no reason to raise attention.
2026-07-21T06:26:05Z
grounded: known/medium — Scott already treats precision, memory pressure, accelerator placement, and compilation as explicit policy in dev:concept.hardware-aware-local-inference, with d
2026-07-21T06:23:52Z
case created — The implementation makes a bounded and technically significant low-VRAM training claim, but currently has only a single author report and no independent validation.