2026-10-11 16:37 UTC

GLQ’s maintainer claims its released trellis-quantization kernels serve SmolLM3-3B at near-bf16 single-stream speed in one-third the memory through vLLM, potentially making compressed local inference practical without a substantial decode penalty.

state: watchingheat: mediumuncertainty: highconvergesscott: mediumquantization vllm local-inference inference-economicscnygaardGLQ

What is this?

The case describes GLQ as a trellis-based 2–8-bit weight-quantization project with fused serving kernels and a vLLM plugin, associating it with cnygaard. Its maintainer reportedly claims released kernels run 4-bit SmolLM3-3B at 176 tokens per second versus 180 for bf16 in single-stream serving, using one-third the memory. None of the supplied web-result snippets directly documents GLQ, its maintainer, or that benchmark: they concern separate Qwen serving results and TurboQuant KV-cache compression, which do not validate GLQ’s weight-quantization claim. Hardware, memory-accounting scope, quality impact, and independent reproduction remain unestablished here.

Why it matters to Scott

GLQ’s claimed precision–memory–speed tradeoff converges with Scott’s Hardware-aware local inference approach and offers a concrete candidate to evaluate for his gamepc serving substrate, rather than merely illustrating a general efficiency preference. This is a conditional testing opportunity, not a validated upgrade: the hits establish Ollama, not vLLM, in his stack; GLQ’s hardware compatibility, quality and benchmark remain unverified, and no supplied radar page tracks this same development.
dev:concept.hardware-aware-local-inferencedev:project.gamepcdev:technology.ollamaradar:concept.quantizationradar:concept.vllmradar:concept.local-inferenceradar:concept.inference-benchmarking
queries asked of Scott's wikis
  • local inference economics GPU memory decode latency
  • weight quantization quality speed tradeoffs
  • vLLM serving stack custom kernels plugins
  • self-hosted small models agent harness deployment
  • inference benchmarks single-stream throughput memory accounting

Measured heat

now 0 pts/hpeak 0 pts/hcomments 0/hpeers p14momentum: steady2 platformsage 746h
points/hour across evidence · reading as of 2026-10-12 02:59:37.977291+11:00 · deterministic, not a model opinion

How the heat travelled

09-10 15:40 (minted)⭐ origin echo-reconstructedGLQ provides 2–8-bit weight quantization and fused serving kernels, reporting 176 versus 180 tokens per second for 4-bit SmolLM3-3B versus b
cnygaard on github (echo) · attributed from hn.story.49644045 · published time unknown
—
09-10 14:12first on hacker news · published · lag ?Glq a port trellis quantization of large language models as a vLLM plugin
acd
—
09-10 14:12amplified on hacker news 👑hn.story.49644045
acd
peak 2 · 1 comments · 101% of case engagement
09-10 15:22our radar first saw it · lag ?discovery anchor: hn.story.49644045—
pace: p36 vs 519 stories at the 720h mark (now 746h old) — ahead of addom-local-coding-harness (1.5x), behind checkly-agentic-go-rewrite (0.8x)

Evidence (2) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟧 hnGlq a port trellis quantization of large language models as a vLLM plugin
Retrieved article excerpt

Open article · Retrieved 2026-09-10T15:27:45.022314+00:00

GLQ — fit larger LLMs on smaller GPUs Lattice and trellis-coded post-training quantization for LLM weights: 2–8 bits/weight ,
served on vLLM · HuggingFace Transformers , with deterministic fused CUDA kernels. Validated
from 24 GB 3090-class GPUs (A10G, sm_86) to a 96 GB RTX PRO 6000 Blackwell
(sm_120). The recommended codebook is trellis-coded quantization (QTIP-derived TCQ, --codebook trellis , since v0.7): it reaches an effective quantization dimension of 256 with a
lookup-free decode, which is why it wins where the bits are scarcest — 2 bpw on SmolLM3-3B
gives PPL 11.94 against 13.79 for the lattice path (bf16 9.12). It takes uniform
integer bit-rates, 2–8. The E8 lattice codebooks — each group of 8 weights as a 16-bit index into a
65,536-entry codebook — remain, and not only for the checkpoints published before v0.7:
they are what fractional and per-layer mixed bit-rates run on, which trellis refuses. Both share the rest of the pipeline. A Randomized Hadamard Transform makes the weights
incoherent
so Euclidean nearest-neighbour rounding is near-optimal under the Hessian-weighted
proxy loss, and a fused CUDA kernel matmuls directly against the compressed
indices — on the GPU serving path the dense weight is never materialized, so
GPU memory drops with the compression ratio (CPU inference and a few
architecture fallbacks dequantize instead). What you get NEW in v0.7 — trellis (TCQ) codebook : single-stream decode at bf16 speed (176 vs 180 tok/s, SmolLM3-3B 4 bpw vs bf16, RTX PRO 6000 / vLLM) in a third of the
memory , and good quality at 2–3 bpw. See Trellis codebook . 2–8 bpw , no group-size constraint , optional per-layer mixed precision . Serve anywhere — a vLLM plugin (weight + MoE + embedding) and an
HF Transformers integration. pip install glq , load, run. Small footprint — smallest of the ~4-bit quantizers we measured (vs AWQ /
NVFP4 on a 26B); a 31B fits ≈16.5 GiB at 5 bpw where bf16 needs ≈58 GiB, with
quality within noise of bf16 on our paired reasoning evals. Deterministic kernels — bit-identical logits across runs (reproducible
lm-eval scoring / on-policy RL rollouts). Pick your path → run a model · fit a bigger model on your card · quantize your own · how GLQ compares · serve with vLLM · how it works System requirements Linux x86_64. An NVIDIA GPU is the fast path; since 0.8.16 a CPU-only machine is
also supported — the installer detects the hardware and sets up the matching vLLM
backend automatically. GPU (recommended) : NVIDIA Ampere-class or newer ( sm_86+ ). Developed for 24–32 GB
cards (3090 / 4080 / 4090 / L4 / L40S); validated live on an Ada L4 ( sm_89 ) and an
RTX PRO 6000 Blackwell ( sm_120 ). A recent NVIDIA driver is required; a CUDA toolkit is not — pip install glq ships prebuilt kernels (CPython 3.12–3.14,
x86_64), and nvcc is only needed for the from-source fallback (and for FlashInfer's
sampler JIT on Blackwell — without it glq-chat falls back to vLLM's built-in sampler). CPU-only : no GPU → the installer sets up vLLM's CPU backend and GLQ's own fused
CPU kernels (AVX2 floor, AVX-512 tiers used when present); --cpu forces this on a
GPU machine. Trellis checkpoints only — dense or MoE since 0.8.17 (e8p/shell need
CUDA) — and expect single-digit tok/s. Measured serving under vLLM on an 8-vCPU
Sapphire Rapids (AVX-512-FP16), one request at a time, 64 decoded tokens, vLLM's
default core binding: gemma-4-26B-A4B trellis-4bpw at 3.0 tok/s , 0.4 s to first
token, 18.6 GiB resident (13.9 GiB weights + a 4 GiB KV pool) — against 2.7 tok/s for the dense gemma-4-E4B on the same box, because an MoE reads only its top-k
experts per token. glq-chat also gives the kernels the core vLLM holds back
( VLLM_CPU_NUM_OF_RESERVED_CPU=0 ), measured at 3.4 tok/s ; decode is
memory-bound, so more threads than physical cores add nothing. Usable for short
answers and background work, not fast chat. Model recommendations size against
system RAM, and glq-chat sizes the KV pool from the checkpoint and from the RAM
that is actually free — on CPU the weights, the pool, the runtime and the page cache
all come out of the same RAM, so a browser already holding 10 GiB is 10 GiB this cannot
have, and a machine with no swap (the cloud default) thrashes rather than failing
cleanly when they do not fit. It prints the arithmetic before starting and says how
much it is over by if it does not fit; VLLM_CPU_KVCACHE_SPACE overrides it. Distros : GLQ has only been tested end-to-end on Ubuntu . The installer's
pre-flight is additionally exercised (Docker + GPU) on Ubuntu 24.04/26.04, Debian 13,
Fedora 43/44, AlmaLinux 9, Arch and openSUSE Tumbleweed
( tests/test_installer_distros.py ), so other distros are expected to install and run. Windows : no native support (no Windows wheels, and vLLM is Linux-only). WSL2 with
the NVIDIA CUDA driver should work in principle — untested. macOS : not supported. There is no CUDA on Apple hardware, and Docker Desktop on
macOS cannot pass through an NVIDIA GPU, so a container does not help. Python : 3.12–3.14 (the installer creates its own venv). Quickstart Installer command (venv, glq, vLLM, chat UI, optional pi agent) curl -fsSL https://raw.githubusercontent.com/cnygaard/glq/main/install.sh | bash Or preselect components — bash -s -- is how arguments reach a piped script. With
the quantize deps, for making your own checkpoints: curl -fsSL https://raw.githubusercontent.com/cnygaard/glq/main/install.sh | bash -s -- --components core,vllm,chat,quantize Or with the pi coding agent (installs node via nvm; afterwards glq-code serves a
tool-calling vLLM and runs pi against it): curl -fsSL https://raw.githubusercontent.com/cnygaard/glq/main/install.sh | bash -s -- --components core,vllm,chat,picode CPU-only serving on a machine that has a GPU you don't want used (no flag needed on a
machine without one — the installer detects it): curl -fsSL https://raw.githubusercontent.com/cnygaard/glq/main/install.sh | bash -s -- --cpu Creates a venv at ~/.glq/venv , then discovers the published checkpoints, sizes
them against your GPU and offers the ones that fit. When it finishes it offers to
start GLQ and open the chat; answer no and it just prints the steps. It refuses to
run as root and never calls sudo . In a terminal the installer prompts for components and checkpoint (the prompts
read /dev/tty , so they appear even though the script itself arrives on stdin).
Scripted — --yes , --dry-run , or no terminal at all (CI, ssh host 'cmd' , docker build ) — it prompts for nothing and installs the defaults, so use --components there to get anything non-default. The components : Component What it installs Default core the venv + glq itself always vllm vLLM — the OpenAI-compatible server behind the chat UI and picode ✓ chat the Gradio chat UI ✓ picode the pi coding agent (installs node via nvm) — run with glq-code , which starts a tool-calling vLLM for it and frees the GPU when pi exits opt-in quantize the deps for quantizing your own models ( glq[quantize] ) opt-in Other flags: --dry-run prints every command without running it; --list shows the
checkpoints and exits; --start / --no-start decide the handoff without being
asked; --no-modify-path leaves your shell rc file alone (by default the venv's bin/ is appended to PATH in ~/.bashrc / ~/.zshrc so vllm , glq-chat and ninja resolve by name in new shells). ~/.glq/venv/bin/glq-chat is the one command afterwards: it starts vLLM, waits for
it, serves the Gradio UI on http://localhost:7860 , and stops the server again when
you press Ctrl-C — vLLM has no idle unload, so a server left running keeps its share
of the card. It sizes the VRAM reservation from the checkpoint — weights, runtime
overhead and a usable cache — and sizes the context window from the card's KV
headroom, in tiers from 8192 up to 65536, clamped to the model's own maximum (a
24 GB card serving a 26B stays at 8192; a 96 GB card reaches 65536). --gpu-memory-utilization / --max-model-len pin either; --no-serve attaches to
a server you started yourself. The first start takes minutes — weights download, model load, CUDA-graph capture —
so it reports progress in vLLM's own words while it waits and writes the full server
log to ~/.glq/vllm.log . --verbose streams that log instead of summarising it. It also publishes a public https://….gradio.live link by default, so the chat can be
opened from a phone or another machine with no port forwarding. That link is unauthenticated for as long as the chat runs — anyone holding it can use your GPU ; --no-share keeps everything on localhost. The rest of this section document the manual path. Run a pre-quantized model pip install ' glq[hf] ' # glq + transformers + accelerate; requires PyTorch ≥ 2.0 import glq . hf_integration # registers GLQ with transformers from transformers import AutoModelForCausalLM , AutoTokenizer model = AutoModelForCausalLM . from_pretrained ( "xv0y5ncu/SmolLM2-360M-Instruct-GLQ-block-diagonal-4bpw" , device_map = "auto" ,
) tok = AutoTokenizer . from_pretrained ( "xv0y5ncu/SmolLM2-360M-Instruct-GLQ-block-diagonal-4bpw" ) print ( tok . decode ( model . generate ( ** tok ( "The capital of France is" , return_tensors = "pt" ). to ( model . device ), max_new_tokens = 20 ,
)[ 0 ], skip_special_tokens = True )) import glq.hf_integration registers quant_method="glq" with HF
Transformers; from_pretrained then swaps nn.Linear for E8RHTLinear and uses the fused CUDA C kernel on inference. CPU falls back to a
naive dequantize-then-matmul. Or serve the fastest GLQ checkpoint on vLLM (the trellis-3INST decode —
single-stream speed at bf16 parity, 1.9 GiB of weights): pip install glq vllm # glq ≥ 0.7.0 (trellis kernel storage layout) vllm serve xv0y5ncu/SmolLM3-3B-trellis-3inst-4bpw-kernel --quantization glq Blackwell (sm_120 — RTX 5090, RTX PRO 6000): vLLM's FlashInfer sampler ships no
prebuilt kernel for this architecture and compiles one at startup. Without the NVIDIA
CUDA Toolkit that build fails and takes the engine down before the first token — GLQ's
own kernels are fine and load normally. Either install the toolkit, or run with VLLM_USE_FLASHINFER_SAMPLER=0 . glq-chat detects this and falls back on its own; vllm serve and the LLM(...) API do not. Available pre-quantized checkpoints The table mirrors the live start-here collection — the same curated list glq-setup --list and the installer picker read — in its order,
plus one downloads-earned extra at the end. Everything is on the xv0y5ncu HF org . Repo Base model bpw License Footprint¹ Best for SmolLM3-3B-trellis-3inst-4bpw-kernel SmolLM3-3B 4.0 trellis Apache 2.0 1.9 GiB flagship / fastest GLQ decode — single-stream at bf16 parity gemma-4-26B-A4B-it-GLQ-trellis-3inst-4bpw Gemma-4-26B-A4B (MoE) 4.0 trellis Apache 2.0 14.4 GiB best quality-per-GB (MoE) gemma-4-26B-A4B-it-GLQ-trellis-3inst-3bpw Gemma-4-26B-A4B (MoE) 3.0 trellis Apache 2.0 11.9 GiB the 26B for 12–16 GB cards — AIME-2026 avg@8 82.1% vs 86.25% at 4 bpw Qwen3.8-27B-GLQ-trellis-3inst-4bpw Qwen3.8-27B (hybrid GDN) 4.0 trellis Apache 2.0 16.7 GiB the biggest model on a 24 GB card — AIME-2026 avg@8 90.4% (bf16 87.9%, n=30) Gemma-4-31B-it-GLQ-5.0bpw-mix3-8 Gemma-4-31B 5.0 mix Apache 2.0 16.5 GiB a 31B on one 24–32 GB card (vs 57.9 GiB bf16) Gemma-4-12B-it-GLQ-5.0bpw Gemma-4-12B 5.0 mix Apache 2.0 6.9 GiB 12B, mixed 3–8 bpw allocation gemma-4-E4B-it-GLQ-trellis-3inst-4bpw Gemma-4-E4B (8B, multimodal) 4.0 trellis Apache 2.0 6.58 GiB a capable model on an 8–12 GB card SmolLM3-3B-GLQ-block-diagonal-3.5bpw SmolLM3-3B 3.5 mix Apache 2.0 1.8 GiB fractional-bpw example; small + fast SmolLM2-360M-Instruct-GLQ-trellis-3inst-6bpw SmolLM2-360M 6.0 trellis Apache 2.0 0.31 GiB near-lossless tiny — wikitext-2 PPL within +0.2% of bf16 SmolLM2-135M-Instruct-GLQ-block-diagonal-4bpw SmolLM2-135M 4.0 Apache 2.0 0.1 GiB smallest checkpoint; CI smoke tests Devstral-Small-2-24B-Instruct-GLQ-4bpw Devstral-Small 24B 4.0² Apache 2.0 ~20.5 GiB coding / agentic (top-15 by downloads; not in start-here) 44 public checkpoints total — the H
acd21
🟧 echo.github ⭐GLQ provides 2–8-bit weight quantization and fused serving kernels, reporting 176 versus 180 tokens per second for 4-bit SmolLM3-3B versus bcnygaard——

Interpretation history

Decision trace