An NVIDIA developer-forum experiment claims partially encrypted CKKS inference can approach one second per token on a DGX Spark while every-layer encrypted generation remains near six minutes per token, defining a sharply limited near-term usability frontier for homomorphic LLM inference.
state: expiredheat: lowuncertainty: highconvergesscott: mediumprivacy-preserving-inference inference-economics ai-securityvkaufmannNVIDIA
What is this?
An NVIDIA Developer Forums post by user vkaufmann reports a same-hardware experiment using CKKS homomorphic encryption for LLM inference on one DGX Spark. It distinguishes an interactive protocol—where activations may be decrypted and re-encrypted at layer boundaries—from an every-layer-encrypted run in which the server reportedly never holds plaintext, claiming roughly 1.05 seconds per token for the former and nearly six minutes per token for the latter. The supplied evidence is primarily the experimenter’s forum post; it does not establish independent reproduction, NVIDIA endorsement, or broader generality across models and hardware.
Why it matters to Scott
The reported gap supports Scott’s architectural position that practical privacy currently depends on trusted boundaries, tokenised representations, and privilege separation rather than giving general models raw sensitive data or relying on fully encrypted computation. The measurements could sharpen what he builds and advises by quantifying that trade-off, but the evidence is only a single, unreplicated forum experiment—not an NVIDIA-endorsed result—and the radar already tracks the broader homomorphic-inference viability question.
ip:framework.separation-of-powers-for-cognitionip:concept.proxy-mediated-tokenisationdev:concept.privacy-tokenized-agent-boundarydev:concept.hardware-aware-local-inferenceradar:concept.confidential-computingradar:google-homomorphic-private-airadar:concept.inference-economics
queries asked of Scott's wikis
- homomorphic encryption versus trusted inference boundaries
- privacy-preserving inference usability frontier
- encrypted inference threat-model taxonomy
- confidential AI latency and inference economics
- client-assisted versus server-blind inference
- local inference as a privacy architecture
Measured heat
no measured readings yet — the hourly heat pass fills this in
How the heat travelled
no chain yet — the hourly chain pass fills this in
Evidence (2) — ⭐ canonical anchor
Interpretation history
2026-09-02T15:44:24Z
After another 48 hours without model disclosure, code, discussion, or independent reproduction, this remains an isolated anecdotal benchmark rather than a developing signal. Preserve it as a provisional bound on the interactive-versus-fully-encrypted trade-off, but close active tracking until a reproducible artifact appears.
2026-08-31T14:50:11Z
No new evidence, engagement, or independent reproduction has appeared; the reported latency gap remains a useful but highly provisional single-experiment bound. The case should cool while awaiting disclosed model details, code, or replication.
2026-08-31T14:43:15Z
grounded: converges/medium — The reported gap supports Scott’s architectural position that practical privacy currently depends on trusted boundaries, tokenised representations, and privileg
2026-08-31T14:41:18Z
origin walked (codex/luna, conf 0.99): anchor hn.story.49509884 -> echo.other.0fd1113420 by Vincent Kaufmann
2026-08-31T14:40:19Z
case created — The first-party measurements provide concrete performance and economic bounds for encrypted inference rather than a general privacy claim.
Decision trace
- 09-03 01:44expireAfter another 48 hours without model disclosure, code, discussion, or independent reproduction, this remains an isolated anecdotal benchmark rather than a developing signal. Preserve it as a provision
- 09-03 01:44alert_silentThe only delta is elapsed time with unchanged evidence and engagement; nothing new affects Scott's architecture or purchasing decisions, so interruption would add no value.
- 09-03 01:44alert_routeThe only delta is elapsed time with unchanged evidence and engagement; nothing new affects Scott's architecture or purchasing decisions, so interruption would add no value.
- 09-01 00:50repriceNo new evidence, engagement, or independent reproduction has appeared; the reported latency gap remains a useful but highly provisional single-experiment bound. The case should cool while awaiting dis
- 09-01 00:50alert_silentThe re-observation adds no consequential delta, and the underlying measurements remain too underspecified and unreplicated to warrant interrupting Scott before a normal briefing.
- 09-01 00:50alert_routeThe re-observation adds no consequential delta, and the underlying measurements remain too underspecified and unreplicated to warrant interrupting Scott before a normal briefing.
- 09-01 00:47alert_silentThe original forum artifact establishes that its author reported this experiment, but the withheld model details, absent code or paper, and lack of replication prevent the measurements from reliably c
- 09-01 00:47surface_candidateThe original forum artifact establishes that its author reported this experiment, but the withheld model details, absent code or paper, and lack of replication prevent the measurements from reliably c
- 09-01 00:47alert_routeThe original forum artifact establishes that its author reported this experiment, but the withheld model details, absent code or paper, and lack of replication prevent the measurements from reliably c
- 09-01 00:43groundThe reported gap supports Scott’s architectural position that practical privacy currently depends on trusted boundaries, tokenised representations, and privilege separation rather than giving general
- 09-01 00:41promote_anchororigin walk conf 0.99
- 09-01 00:40createThe first-party measurements provide concrete performance and economic bounds for encrypted inference rather than a general privacy claim.