Independent benchmarks will determine whether llama.cpp's SYCL TILE-kernel dispatch materially accelerates long-context quantized-KV decoding on Intel Battlemage GPUs across representative models and contexts.
state: expiredheat: lowuncertainty: highknownscott: lowlocal-inference llama-cpp intel-gpullama.cppIntel
What is this?
A llama.cpp pull request changes Intel SYCL dispatch for quantized key-value-cache decoding from a VEC kernel to a TILE kernel, with its author reporting up to 169% faster decoding at 118K-token context on Intel Battlemage hardware. The supplied results establish that llama.cpp’s SYCL backend targets Intel GPUs and can outperform its OpenCL backend, but they also document uneven SYCL performance and call for reproduction across GPUs, models, quantizations, drivers, and context lengths. The claimed gain is therefore a promising single-source optimization result, not yet evidence of a broadly reliable acceleration.
Why it matters to Scott
This is another hardware-specific llama.cpp optimization awaiting representative validation, a pattern already tracked in “llama.cpp,” “BeeLlama.cpp KVarN and low-bit KV-cache validation,” and several backend-kernel cases. It touches Scott’s hardware-aware local-inference and capability-audit work, but the supplied hits show his active substrate is NVIDIA/CUDA rather than Intel SYCL, so this result is unlikely to change what he builds unless independent tests establish broader portability or compelling Intel economics.
dev:concept.hardware-aware-local-inferenceip:concept.capability-auditradar:concept.llama-cppradar:person.llama-cppradar:concept.kv-cacheradar:concept.long-context-inferenceradar:beellama-kvarn-kv-cache-validation
queries asked of Scott's wikis
- local inference benchmark methodology
- long-context KV-cache quantization tradeoffs
- llama.cpp backend optimization strategy
- Intel GPU local inference economics
- hardware-specific kernel dispatch portability
- representative benchmarking for local models
Measured heat
no measured readings yet — the hourly heat pass fills this in
How the heat travelled
no chain yet — the hourly chain pass fills this in
Evidence (2) — ⭐ canonical anchor
Interpretation history
2026-08-13T14:39:05Z
The expected controlled VEC-versus-TILE replication has not emerged, and repeated checks have added only anecdotal broader-SYCL reports. This narrow validation episode has faded without resolving the kernel-specific claim.
2026-08-11T13:52:47Z
No controlled VEC-versus-TILE benchmark has appeared; the independent reports remain anecdotal evidence of broader SYCL improvement rather than validation of this kernel-specific claim. With only stale engagement and no substantive delta, the case should cool while awaiting reproducible comparisons.
2026-08-09T13:29:28Z
A second independent user now reports exercising recent SYCL changes on a B580 with a Gemma workload at roughly 49K context, broadening anecdotal evidence beyond the B70. The visible comment still lacks a controlled VEC-versus-TILE comparison, so the specific acceleration claim remains uncorroborated.
2026-08-07T20:29:27Z
An independent B70 user now reports running the change in a daily SYCL build and seeing continued performance improvement, making this more than author-only evidence. However, the comment lacks a controlled before/after benchmark, so it does not yet validate the claimed long-context gains or justify corroborated status.
2026-08-07T17:33:08Z
grounded: known/low — This is another hardware-specific llama.cpp optimization awaiting representative validation, a pattern already tracked in “llama.cpp,” “BeeLlama.cpp KVarN and l
2026-08-07T17:30:45Z
origin walked (codex/luna, conf 0.99): anchor reddit.post.1vi6hmw -> echo.github.e74fcf74a8 by johnkarlhill
2026-08-07T17:29:41Z
case created — The reported upstream kernel change has large and technically specific long-context gains that merit replication.
Decision trace
- 08-14 00:39expireThe expected controlled VEC-versus-TILE replication has not emerged, and repeated checks have added only anecdotal broader-SYCL reports. This narrow validation episode has faded without resolving the
- 08-14 00:39alert_silentThere is no new evidence or consequential event to surface; staleness alone does not justify an alert.
- 08-14 00:39alert_routeThere is no new evidence or consequential event to surface; staleness alone does not justify an alert.
- 08-11 23:52repriceNo controlled VEC-versus-TILE benchmark has appeared; the independent reports remain anecdotal evidence of broader SYCL improvement rather than validation of this kernel-specific claim. With only stal
- 08-11 23:52alert_silentNothing new establishes the claimed kernel-specific acceleration or changes Scott's decisions; engagement and elapsed time alone do not warrant an alert.
- 08-11 23:52alert_routeNothing new establishes the claimed kernel-specific acceleration or changes Scott's decisions; engagement and elapsed time alone do not warrant an alert.
- 08-10 01:21sensor_dirtyengagement_update
- 08-10 00:21sensor_dirtyengagement_update
- 08-09 23:29repriceA second independent user now reports exercising recent SYCL changes on a B580 with a Gemma workload at roughly 49K context, broadening anecdotal evidence beyond the B70. The visible comment still lac
- 08-09 23:29alert_silentThe refreshed discussion adds another real-world configuration but no controlled benchmark establishing the claimed kernel-specific gain; it can wait for a normal briefing or reproducible before/after
- 08-09 23:29alert_routeThe refreshed discussion adds another real-world configuration but no controlled benchmark establishing the claimed kernel-specific gain; it can wait for a normal briefing or reproducible before/after
- 08-09 23:21sensor_dirtycomment_update
- 08-09 05:21sensor_dirtyengagement_update
- 08-09 02:21sensor_dirtyengagement_update
- 08-08 22:21sensor_dirtyengagement_update
- 08-08 19:21sensor_dirtyengagement_update
- 08-08 15:21sensor_dirtyengagement_update
- 08-08 13:21sensor_dirtyengagement_update
- 08-08 09:21sensor_dirtyengagement_update
- 08-08 07:21sensor_dirtyengagement_update
- 08-08 06:29repriceAn independent B70 user now reports running the change in a daily SYCL build and seeing continued performance improvement, making this more than author-only evidence. However, the comment lacks a cont
- 08-08 06:29alert_silentThe new adopter testimony modestly strengthens the case but provides no detailed comparative results and does not change access, pricing, or Scott's near-term build decisions; it can wait for nor
- 08-08 06:29alert_routeThe new adopter testimony modestly strengthens the case but provides no detailed comparative results and does not change access, pricing, or Scott's near-term build decisions; it can wait for nor
- 08-08 06:21sensor_dirtyengagement_update
- 08-08 05:21sensor_dirtyengagement_update
- 08-08 04:21sensor_dirtycomment_update
- 08-08 03:40alert_silentAn open llama.cpp PR establishes a testable SYCL dispatch change and reports large Battlemage gains, but the results are author-reported, hardware-specific, unmerged, and outside Scott’s active NVIDIA
- 08-08 03:40surface_candidateAn open llama.cpp PR establishes a testable SYCL dispatch change and reports large Battlemage gains, but the results are author-reported, hardware-specific, unmerged, and outside Scott’s active NVIDIA
- 08-08 03:40alert_routeAn open llama.cpp PR establishes a testable SYCL dispatch change and reports large Battlemage gains, but the results are author-reported, hardware-specific, unmerged, and outside Scott’s active NVIDIA
- 08-08 03:33groundThis is another hardware-specific llama.cpp optimization awaiting representative validation, a pattern already tracked in “llama.cpp,” “BeeLlama.cpp KVarN and low-bit KV-cache validation,” and several
- 08-08 03:30promote_anchororigin walk conf 0.99
- 08-08 03:29createThe reported upstream kernel change has large and technically specific long-context gains that merit replication.