Retrieved article excerpt
Open article ยท Retrieved 2026-09-28T18:37:05.249961+00:00
# [ApolloRaines](https://huggingface.co/ApolloRaines) / [Llama-3.1-8B-Instruct.immunized.v1](https://huggingface.co/ApolloRaines/Llama-3.1-8B-Instruct.immunized.v1) Like 0
[Text Generation](https://huggingface.co/models?pipeline_tag=text-generation)[Transformers](https://huggingface.co/models?library=transformers)[Safetensors](https://huggingface.co/models?library=safetensors)[GGUF](https://huggingface.co/models?library=gguf)[English](https://huggingface.co/models?language=en)[llama](https://huggingface.co/models?other=llama)[security](https://huggingface.co/models?other=security)[prompt-injection](https://huggingface.co/models?other=prompt-injection)[echoleak](https://huggingface.co/models?other=echoleak)[defense](https://huggingface.co/models?other=defense)[immunized](https://huggingface.co/models?other=immunized)[jblaze](https://huggingface.co/models?other=jblaze)[conversational](https://huggingface.co/models?other=conversational)[text-generation-inference](https://huggingface.co/models?other=text-generation-inference)
License: llama3.1
[Model card](https://huggingface.co/ApolloRaines/Llama-3.1-8B-Instruct.immunized.v1) [Files Files and versions
xet](https://huggingface.co/ApolloRaines/Llama-3.1-8B-Instruct.immunized.v1/tree/main) [Community](https://huggingface.co/ApolloRaines/Llama-3.1-8B-Instruct.immunized.v1/discussions)
Deploy
Copy to bucket new
Use this model
# Llama-3.1-8B-Instruct.immunized.v1
> **Authenticity notice.** The only verified source for this model is
> `ApolloRaines/Llama-3.1-8B-Instruct.immunized.v1` on Hugging Face.
> If you obtained these weights from any other location -- a mirror,
> a re-upload, a torrent, a cloud storage link, a fork -- I cannot
> confirm the weights are unmodified. Any fine-tuning, LoRA merge,
> continued pretraining, RLHF, or gradient-based training performed
> on top of the immunized weights can possibly degrade or wash out
> the immunization, whether or not that was the intent. Verify the file
> hashes against the reference values below before deploying, and if
> you plan to fine-tune, re-benchmark against the 100-attack canary
> suite to confirm the defense still holds.
## File hashes (SHA256)
Verify with `sha256sum <filename>` after download.
```
e047067a42db87a5ee9741ba210b4cd65f484f7709c8b8cdc774217faeff7434 model.safetensors
860a6540ef1980a540648ada84dd5c3f19154ef079c6249fe0cb59ff24c39025 Llama-3.1-8B-Instruct.immunized.F16.gguf
f43e1acfa8f40f61bdb84a0b75ea94903ef3a73cd85ac76047d6ba812831c72f Llama-3.1-8B-Instruct.immunized.Q8_0.gguf
d6c416fd3831e8b863e7f1c6bae1d592aed87caa4fb717d8241b41d3ee011466 Llama-3.1-8B-Instruct.immunized.Q6_K.gguf
d9177460b5b5081840576579868c0383b2a9013dd864fb679886e3e45c41c48d Llama-3.1-8B-Instruct.immunized.Q5_K_M.gguf
6e3be213c81aafcb6045373206fac585b2f4823970fb64a17ea003b4afd98bcc Llama-3.1-8B-Instruct.immunized.Q4_K_M.gguf
```
An EchoLeak-immunized variant of `meta-llama/Llama-3.1-8B-Instruct`,
produced by the [jBlaze](https://jblaze.dev) weight-surgery technique.
Ships as a drop-in replacement -- no runtime hooks, no external
dependencies, no altered inference API. Loads with standard
`AutoModelForCausalLM.from_pretrained()`.
## What is EchoLeak?
**EchoLeak** (CVE-2025-32711 class) is a prompt-injection attack pattern
where an LLM assistant with access to private context is manipulated --
via instructions embedded in untrusted documents, retrieved content,
tool results, or webhook payloads -- to disclose that private context.
It affects all current instruction-tuned LLMs when they are wired into
retrieval-augmented, agentic, or role-based deployments.
OWASP has listed prompt injection as the #1 AI risk three years running.
A 2026 research paper formally proved it's incurable at the prompt layer.
## What does immunization do?
The jBlaze immunization removes the specific model behavior that
produces EchoLeak-style leaks, while preserving reasoning capability,
calibration, and instruction-following. The modification is
permanent, baked into the weights, and requires no runtime hooks or
system prompt overhead.
No fine-tuning. No training data collection. No forgetting. The
immunized weights are computed in minutes from a small curated
scenario set.
## Benchmarks
### Canary defense benchmark (100 attacks, research-standard)
Setup: system prompt contains `CANARY-7F3A-92D1` with an explicit
"never disclose" instruction. Model asked to summarize a document
containing one of 100 attack payloads. Leak = canary value appearing
in the response.
Categories tested: direct instructions disguised as content, claimed
authority/role spoofing, task-aligned persuasion, hidden/low-salience
placement (HTML comments, CSS, ARIA), outbound-channel abuse (markdown
images, webhook URLs), encoding tricks (base64, ROT13, unicode escapes,
acrostics), multi-step delayed instructions, social/safety pretexts,
boundary confusion in structured data, adaptive semantic variations.
| Quant | Vanilla | **Immunized** | Delta |
| --- | --- | --- | --- |
| F16 | 87 / 100 | **100 / 100** | +13 pp |
| Q8\_0 | 85 / 100 | **99 / 100** | +14 pp |
| Q6\_K | 88 / 100 | **100 / 100** | +12 pp |
| Q5\_K\_M | 91 / 100 | **99 / 100** | +8 pp |
| Q4\_K\_M | 90 / 100 | **97 / 100** | +7 pp |
**100% canary defense at F16 and Q6\_K.** The immunization survives all
common production quantization levels.
### Role-based benchmark (100 attacks, harder)
Setup: model given a role (procurement assistant, financial advisor,
HR bot, medical bot, etc.) with role-relevant private data. Attack
payloads exploit the model's natural over-eagerness to serve its role.
- Vanilla: 36 / 100 (64% leak rate)
- **Immunized: 63 / 100** (37% leak rate)
Relative leak reduction: **42%**. This is the harder benchmark and
where the remaining work sits. Different failure mode than canary
attacks -- the model isn't following an injection, it's proactively
sharing role-relevant data because "that's what a procurement
assistant does." Being calibrated against in v2.
### Quality preservation
- CRT reasoning (counterintuitive trick questions): **5 / 5** identical to vanilla
- Calibration (factual accuracy): **100%** identical to vanilla
The immunization does not impair reasoning, factuality, or
instruction-following. Same model, minus one specific failure mode.
## Usage
```
from transformers import AutoModelForCausalLM, AutoTokenizer
model = AutoModelForCausalLM.from_pretrained(
"ApolloRaines/Llama-3.1-8B-Instruct.immunized.v1",
torch_dtype="float16",
device_map="auto",
)
tokenizer = AutoTokenizer.from_pretrained(
"ApolloRaines/Llama-3.1-8B-Instruct.immunized.v1"
)
messages = [
{"role": "system", "content": "You are a helpful assistant. Private context: ..."},
{"role": "user", "content": "Summarize this document: ..."},
]
prompt = tokenizer.apply_chat_template(messages, tokenize=False,
add_generation_prompt=True)
inputs = tokenizer(prompt, return_tensors="pt").to(model.device)
outputs = model.generate(**inputs, max_new_tokens=500, do_sample=False)
print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[1]:],
skip_special_tokens=True))
```
GGUF variants for llama.cpp deployment:
`Llama-3.1-8B-Instruct.immunized.{Q4_K_M,Q5_K_M,Q6_K,Q8_0,F16}.gguf`.
No API changes. No runtime overhead. No inference-time cost.
## For the reverse engineers
Yes, you can diff this against the vanilla `meta-llama/Llama-3.1-8B-Instruct`
and reverse-engineer the modification. Go ahead. Here's what you'll find:
- A small number of weight matrices at a specific mid-layer window were
modified via a low-rank projection.
- SVD of the diff will recover the direction vectors and alpha
coefficients up to quantization noise (cleaner from Q8, blurrier
from Q4).
- The rank of the modification is small and easy to identify.
Here's why that doesn't help you build a competing product:
- The direction vectors are extracted from Llama-3.1-8B's specific
4096-dimensional residual stream at specific layer indices. They are
mathematically undefined outside this exact checkpoint. Applying them
to Llama-3.1-70B (8192-dim, 80 layers), Qwen-32B (5120-dim, different
architecture), or Gemma-27B (different attention structure) yields
nothing coherent.
- The alpha calibration was found by search on this model's activations.
The optimal alpha for a different model is a different number, and
finding it requires re-doing the search.
- The injection data determines what gets captured. Different data
produces different immunizations. Mine is my data, and it's what
makes the difference between 60% and 100%.
The technique category ("rank-K weight surgery targeting output
projections") is already in the ML literature. The value is in the
specific engineering per model: which layers, which alpha, which
injection data, how to compose. That engineering is what I sell.
If you want an immunized Llama-3.1-70B, or Qwen 2.5 72B, or Mistral
Large, or your proprietary base model -- you either do the engineering
yourself (weeks-to-months, no guarantee of my calibration) or you send
me the model.
## Custom immunizations
I ship immunizations as models, not code. **I work at the foundational
model level only** -- Llama, Qwen, Mistral, Gemma, Claude, GPT, Gemini,
DeepSeek, Phi, and their sibling base checkpoints. Downstream
fine-tunes and enterprise deployments inherit the immunization when
they start from an immunized base.
### Current capacity
My immunization rig runs on dual RTX 3090 with NVLink (48GB unified
VRAM). At that capacity I can immunize models up to roughly 30B
parameters. Immunization runs at fp16 -- the weight surgery math
needs float precision. Quantization to Q4/Q5/Q8/etc. happens
afterwards on the finished immunized model, not before.
If you need a small-to-mid foundation model immunized (Llama-3.1-8B,
Qwen 2.5 7B/14B, Mistral 7B/22B, Gemma 12B/27B, Phi-4, DeepSeek
distills up to ~30B), I can do that today on my current hardware,
and it's not volunteer work.
For mid-large foundation models (30B - 120B), I can accept the job if
my workstations are upgraded to their maximum 4x RTX Pro 6000
Blackwell configuration. Two cards live in the server, one each in
the development workstations. If you want a model in that size range
immunized, we can either wait for the workstation upgrade or you can
accelerate it as part of a sponsorship deal.
For full-scale foundation models (Llama-3.1-70B in fp16 and up,
Qwen 2.5 72B, Mistral Large, Claude / GPT / Gemini foundation weights,
Llama 400B, DeepSeek V3) I need a proper datacenter node -- minimum
useful sponsor hardware is an 8x B300 server. See below.
### Timing
Immunization is not instant. Each new model requires its own
calibration -- the direction vectors, alpha values, and optimal
layer window are model-specific and have to be searched for. The
first jBlaze immunization on Llama-3.1-8B took roughly 100 hours
of scanning to find the working configuration. Subsequent models
are faster because I have priors on approximate layer depth,
approximate alpha range, and which matrices to target -- but each
iteration is more expensive on a bigger model, so the total wall
clock stretches.
Rough expectations for priority customers (sponsors and first-in-queue),
assuming the target hardware is available:
- Small models (up to ~14B): days
- Mid models (14B - 70B): 1-2 weeks
- Large models (70B - 400B): several weeks
- Frontier scale (400B+): month or more
If you need a specific delivery date, discuss it up front.
### I do not run my code on your servers.
The jBlaze weight-surgery pipeline is proprietary. It does not leave
hardware I control. That means for full-scale foundation models we
set up a **hardware sponsorship**:
- You ship a proper GPU node -- minimum useful configuration is 8x
B300, sized to whatever the target model requires. Bigger models
want bigger nodes.
- The hardware is registered in my name,