2026-10-11 17:13 UTC

An NVIDIA developer-forum experiment claims partially encrypted CKKS inference can approach one second per token on a DGX Spark while every-layer encrypted generation remains near six minutes per token, defining a sharply limited near-term usability frontier for homomorphic LLM inference.

state: expiredheat: lowuncertainty: highconvergesscott: mediumprivacy-preserving-inference inference-economics ai-securityvkaufmannNVIDIA

What is this?

An NVIDIA Developer Forums post by user vkaufmann reports a same-hardware experiment using CKKS homomorphic encryption for LLM inference on one DGX Spark. It distinguishes an interactive protocol—where activations may be decrypted and re-encrypted at layer boundaries—from an every-layer-encrypted run in which the server reportedly never holds plaintext, claiming roughly 1.05 seconds per token for the former and nearly six minutes per token for the latter. The supplied evidence is primarily the experimenter’s forum post; it does not establish independent reproduction, NVIDIA endorsement, or broader generality across models and hardware.

Why it matters to Scott

The reported gap supports Scott’s architectural position that practical privacy currently depends on trusted boundaries, tokenised representations, and privilege separation rather than giving general models raw sensitive data or relying on fully encrypted computation. The measurements could sharpen what he builds and advises by quantifying that trade-off, but the evidence is only a single, unreplicated forum experiment—not an NVIDIA-endorsed result—and the radar already tracks the broader homomorphic-inference viability question.
ip:framework.separation-of-powers-for-cognitionip:concept.proxy-mediated-tokenisationdev:concept.privacy-tokenized-agent-boundarydev:concept.hardware-aware-local-inferenceradar:concept.confidential-computingradar:google-homomorphic-private-airadar:concept.inference-economics
queries asked of Scott's wikis
  • homomorphic encryption versus trusted inference boundaries
  • privacy-preserving inference usability frontier
  • encrypted inference threat-model taxonomy
  • confidential AI latency and inference economics
  • client-assisted versus server-blind inference
  • local inference as a privacy architecture

Measured heat

no measured readings yet — the hourly heat pass fills this in

How the heat travelled

no chain yet — the hourly chain pass fills this in

Evidence (2) — ⭐ canonical anchor

sourceobjectauthorscorecomments
🟧 hnLLM inference FHE homomorphic encryption: 1s/token interactive, 6min/token fullvkaufmann10
🟧 echo.other ⭐The NVIDIA forum post is the original artifact. It reports measurements made that month on one DGX Spark: “Interactive protocol: 1.05 s/tokeVincent Kaufmann——

Interpretation history

Decision trace