Retrieved article excerpt
Open article · Retrieved 2026-09-15T04:21:51.374809+00:00
[Back to Articles](https://huggingface.co/blog)
# What If the Adaptation Were a Model? ShadowPEFT in 🤗 PEFT library
[Community Article](https://huggingface.co/blog/community)
Published
September 15, 2026
[Upvote
9](https://huggingface.co/login?next=%2Fblog%2Fshadow-llm%2Fshadowpeft-peft)
- +3
[Zx Li's avatar](https://huggingface.co/csroyli)
[Zx Li
csroyli
Follow](https://huggingface.co/csroyli)
[ShadowLLM's avatar](https://huggingface.co/shadow-llm "ShadowLLM") [shadow-llm](https://huggingface.co/shadow-llm)
[SeanLee's avatar](https://huggingface.co/SeanLee97)
[SeanLee
SeanLee97
Follow](https://huggingface.co/SeanLee97)
[ShadowLLM's avatar](https://huggingface.co/shadow-llm "ShadowLLM") [shadow-llm](https://huggingface.co/shadow-llm)
[HSURA's avatar](https://huggingface.co/HSURA)
[HSURA
HSURA
Follow](https://huggingface.co/HSURA)
[ShadowLLM's avatar](https://huggingface.co/shadow-llm "ShadowLLM") [shadow-llm](https://huggingface.co/shadow-llm)
[Jing Li's avatar](https://huggingface.co/girlgunner)
[Jing Li
girlgunner
Follow](https://huggingface.co/girlgunner)
[ShadowLLM's avatar](https://huggingface.co/shadow-llm "ShadowLLM") [shadow-llm](https://huggingface.co/shadow-llm)
Authors: Zongxi Li, Xianming Li, Tsz-fung Andrew Lee, Jing Li, Haoran Xie, Qing Li
This article is also available in Chinese [简体中文](https://huggingface.co/blog/shadow-llm/shadowpeft-peft-cn)
---
ShadowPEFT is now a first-class method in the 🤗 [PEFT](https://github.com/huggingface/peft) library: ShadowPEFT has a fundamentally different adapter geometry with LoRA, but it shares the same `get_peft_model` entry point as LoRA, the same adapter save/load path, and is easy to work with. This post starts with a brief PEFT recap, explains what ShadowPEFT changes conceptually, walks through the library integration, and presents head-to-head comparisons with LoRA and DoRA.
> Paper: [ShadowPEFT: Shadow Network for Parameter-Efficient Fine-Tuning](https://huggingface.co/papers/2604.19254).
## The Familiar Parameter-Efficient Finetuning Recipe
Parameter-efficient fine-tuning (PEFT) adapts a large pretrained model to various downstream tasks without updating all of its weights [1].
Full fine-tuning updates every weight. This approach works, but the resulting checkpoint is as large as the model itself, the optimizer parameter scale becomes enormous, and storing many task-specific copies quickly becomes impractical. PEFT addresses these issues by freezing the pretrained model and training only a small set of additional parameters. During inference, these extra parameters steer the frozen backbone; after training, only the additional parameters are stored rather than a full copy of the base model.
The manner in which these extra parameters are attached varies across methods. Prompt-based methods [2] prepend learned tokens to the input. Adapter-based methods insert small modules between layers. LoRA-style methods [3,4] leave the architecture unchanged and add a low-rank update alongside selected weights: for a frozen linear layer WWW, train matrices AAA and BBB and compute y=Wx+α⋅(xA)By = Wx + \alpha \cdot (xA)By=Wx+α⋅(xA)B, where α\alphaα is a scaling factor. Despite these differences, the core idea remains the same—most of the model stays frozen, and a small trainable component handles the task adaptation.
PEFT has changed how we specialize large language models. Instead of updating billions of parameters, methods such as LoRA have made fine-tuning dramatically cheaper and have become a standard way to adapt LLMs. These methods are commonly used through the 🤗 PEFT library [5], released by Hugging Face. It has become a widely adopted toolkit for parameter-efficient fine-tuning on top of Transformers and Diffusers. ShadowPEFT is now one of the methods included in the library.
## Introduction of ShadowPEFT
However, LoRA-style approaches encode a task as a collection of independent low-rank updates scattered across selected weights. The backbone carries representations from one layer to the next, so that these modifications interact through computation. However, the adaptation mechanism itself never maintains an explicit task-specific state that is updated and reused across depth, yielding a less globally coordinated finetuning process. This leads to a natural question: *Could adaptation itself have a state?*
**ShadowPEFT** starts from that question. Instead of representing adaptation as a set of weight perturbations, it introduces a compact and stateful **shadow model** with a persistent hidden state:
s(0)→s(1)→s(2)→⋯→s(L)s^{(0)} \rightarrow s^{(1)} \rightarrow s^{(2)} \rightarrow \cdots \rightarrow s^{(L)}s(0)→s(1)→s(2)→⋯→s(L)
The state s(ℓ)s^{(\ell)}s(ℓ) carries task-specific information at depth ℓ\ellℓ. At every Transformer layer, three operations occur:
1. **Shadow Injection (shadow → base):** The discrepancy h−sh - sh−s is projected through a low-rank bottleneck and added to the block input, where hhh **is the hidden state produced by the frozen base model** at the current layer, and sss **is the shadow state**—a persistent task-specific vector maintained by the trainable shadow network (as shown in Figure 2a).
2. **Base Encoding:** The frozen Transformer layer processes the refined representation (as shown in Figure 2b).
3. **Shadow Update (base → shadow):** Two small MLPs produce a candidate and a gate; the shadow state advances via a gated residual mixture, carrying the accumulated task signal forward to the next layer (as shown in Figure 2c).
This creates a **coupled, bidirectional** information flow—not just a side network, but a process where the shadow refines the backbone and the backbone simultaneously updates the shadow. The core operation is **carry `(h, s)` through the decoder**: at each layer, the frozen block processes `h` while the shadow network injects information from `s` into `h`, then updates `s` using the new `h`. Later layers thus see both the accumulated shadow state and the latest backbone representation.
[Architecture of ShadowPEFT](https://cdn-uploads.huggingface.co/production/uploads/64c76a70720d494bb2acc47b/T08XieCT5QoBza7w_0kBk.jpeg)
### Why this matters
Treating adaptation as a stateful computation unlocks capabilities that are difficult to express as weight deltas:
**1. Cross-layer coordination.** With distributed updates, each adapted layer owns its own parameters, indepedently initializing and separately updating. With ShadowPEFT, the shadow state provides continuity across the entire depth. A later layer can work with information accumulated by the shadow earlier in the computation—it learns *"given everything the task-specific pathway has accumulated so far, what should layer 12 do now?"*
The gated residual update explicitly preserves part of the previous state while incorporating information from the current backbone representation, giving adaptation a mechanism for depth-wise coordination.
**2. Adaptation capacity as a model-scaling.** LoRA's capacity is controlled by rank and the number of adapted matrices. ShadowPEFT introduces another dimension: how large should the task-specific model be, and how large can it be? In our experiments, increasing ShadowPEFT's parameter budget improved performance over a wide range, while LoRA remained flat and DoRA degraded—showing qualitatively different scaling behavior. Instead of asking only *how many parameters should be patched*, we can ask *how much capacity the task-specific computation should have*.
**3. The task-specific component becomes a model in its own right.** A LoRA adapter is meaningful only because it modifies a particular backbone; remove the backbone, and the LoRA parameters do not constitute a complete predictor. ShadowPEFT is different: the shadow network has its own prediction head and receives direct task supervision during training. The resulting checkpoint supports both attached inference and **detached shadow-only inference**. This enables edge-cloud deployments—the compact shadow handles simple requests locally and sends harder ones to the attached cloud model.
**3. Adapter as a pretrained model.** Since the Shadow module is a self-contained functional model, it can be initiated from another pretrained LLM. For example, a smaller model such as Qwen-0.5B can serve as the Shadow model for a larger backbone such as Qwen-8B, enabling the larger model to be adapted by training its smaller counterpart. In this configuration, the smaller model learns to refine the representations of the larger one, allowing adaptation capacity to be reused across model scales. This perspective expands PEFT beyond lightweight parameter injection toward reusable, cross-scale adaptation dynamics.
| | Existing PEFT (LoRA-style) | ShadowPEFT |
| --- | --- | --- |
| Trainable params | Distributed: separate low-rank update on each targeted linear | Centralized: one shadow network plus small per-block inject/update nets |
| Where the update happens | Selected `nn.Linear`s (or prompts / inserted adapters) | Each decoder layer |
| Information flow | One-way local delta: y=Wx+(xA)By = Wx + (xA)By=Wx+(xA)B | Bidirectional: inject shadow→base, then update base→shadow |
| After training | Merge into WWW, or keep the adapter on the same backbone | Keep it attached, or detach the shadow and run it alone |
[Differences between LoRAs and ShadowPEFT](https://cdn-uploads.huggingface.co/production/uploads/64c76a70720d494bb2acc47b/ZUNa7NIxr3u6BUY6NJAYQ.jpeg)
On generation and understanding benchmarks, compact ShadowPEFT is competitive with LoRA and DoRA at similar trainable-parameter budgets. The rest of this post explains how this geometry is wired into 🤗 PEFT—because an adapter that *is* a network does not fit the usual merge-and-unload path.
## How the method is wired into the PEFT library
ShadowPEFT plugs into PEFT through the same three‑piece pattern as LoRA: a config `ShadowConfig`, a `BaseTuner` model, and a `BaseTunerLayer`. That makes `get_peft_model(model, ShadowConfig(...))` work almost the same as LoRA. Checkpoints are saved and loaded via the usual `save_pretrained` / `from_pretrained` APIs, with only minimal key remapping under the hood.
### Usage example
ShadowPEFT has been merged into the `main` branch of PEFT and will be included in the next PEFT release. To try it now, install the development version directly from GitHub:
```
pip install --upgrade git+https://github.com/huggingface/peft.git
```
For full details, see the official documentation: <https://huggingface.co/docs/peft/main/en/package_reference/shadow#shadowpeft>
Below shows an example:
```
from transformers import AutoModelForCausalLM
from peft import ShadowConfig, get_peft_model
# Load the frozen base model
model = AutoModelForCausalLM.from_pretrained("Qwen/Qwen3-8B")
# Configure ShadowPEFT
config = ShadowConfig(
r=8, # rank of the injection bottleneck
shadow_num_hidden_layers=1, # size of the shadow backbone
shadow_model="mirror", # build a scaled-down copy of the base;
# or pass a pretrained checkpoint, e.g.
# "shadow-llm/Qwen3-0.6B-H8B" for Qwen3-8B
task_type="CAUSAL_LM"
)
# Wrap the base model
model = get_peft_model(model, config)
model.print_trainable_parameters()
# Generate as usual
out = model.generate(input_ids, max_new_tokens=32)
# Evaluate the standalone shadow path (the detachable small model)
shadow = model.base_model.unload_shadow() # returns a DetachedShadowModel
shadow_out = shadow.generate(input_ids, max_new_tokens=32)
```
The key difference from LoRA is conceptual: **LoRA lives *inside* selected weights and merges when done; ShadowPEFT is a separate network that you detach when done** via `unload_shadow()`.
### LoRA vs. ShadowPEFT at a glance
| Aspect | LoRA | ShadowPEFT |
| --- | --- | --- |
| Target | Individual linears (`q_proj`, etc.) | Whole decod