Skip to content

Commit a9532bc

Browse files
authored
[docs] peft (huggingface#44804)
* peft * feedback
1 parent dda5468 commit a9532bc

3 files changed

Lines changed: 137 additions & 89 deletions

File tree

‎docs/source/en/_toctree.yml‎

Lines changed: 2 additions & 2 deletions
Original file line numberDiff line numberDiff line change
@@ -163,6 +163,8 @@
163163
- local: hpo_train
164164
title: Hyperparameter search
165165
title: Customization
166+
- local: peft
167+
title: Parameter-efficient fine-tuning
166168
- isExpanded: false
167169
sections:
168170
- local: accelerator_selection
@@ -195,8 +197,6 @@
195197
- local: model_memory_anatomy
196198
title: Model training anatomy
197199
title: Hardware
198-
- local: peft
199-
title: PEFT
200200
title: Training
201201
- isExpanded: false
202202
sections:

‎docs/source/en/main_classes/peft.md‎

Lines changed: 3 additions & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -11,13 +11,15 @@ rendered properly in your Markdown viewer.
1111

1212
# PEFT
1313

14-
The [`~integrations.PeftAdapterMixin`] provides functions from the [PEFT](https://huggingface.co/docs/peft/index) library for managing adapters with Transformers. This mixin currently supports LoRA, IA3, and AdaLora. Prefix tuning methods (prompt tuning, prompt learning) aren't supported because they can't be injected into a torch module.
14+
The [`~integrations.PeftAdapterMixin`] provides functions from the [PEFT](https://huggingface.co/docs/peft/index) library for managing adapters with Transformers. This mixin supports all non-prompt-learning PEFT methods (LoRA, IA3, AdaLoRA, and others). Prefix tuning methods (prompt tuning, prompt learning) aren't supported because they can't be injected into a torch module.
1515

1616
[[autodoc]] integrations.PeftAdapterMixin
1717
- load_adapter
1818
- add_adapter
1919
- set_adapter
2020
- disable_adapters
2121
- enable_adapters
22+
- enable_peft_hotswap
2223
- active_adapters
2324
- get_adapter_state_dict
25+
- delete_adapter

‎docs/source/en/peft.md‎

Lines changed: 132 additions & 86 deletions
Original file line numberDiff line numberDiff line change
@@ -9,98 +9,130 @@ specific language governing permissions and limitations under the License.
99
rendered properly in your Markdown viewer.
1010
-->
1111

12-
# PEFT
12+
# Parameter-efficient fine-tuning
1313

14-
[[open-in-colab]]
14+
[Parameter-efficient fine-tuning (PEFT)](https://huggingface.co/docs/peft/index) methods only fine-tune a small number of extra model parameters (adapters) on top of a pretrained model. Because only adapter parameters are updated, the optimizer tracks far fewer gradients and states, reducing memory usage significantly. Adapters are lightweight, making them convenient to share, store, and load.
1515

16-
[PEFT](https://huggingface.co/docs/peft/index), a library of parameter-efficient fine-tuning methods, enables training and storing large models on consumer GPUs. These methods only fine-tune a small number of extra model parameters, also known as adapters, on top of the pretrained model. A significant amount of memory is saved because the GPU doesn't need to store the optimizer states and gradients for the pretrained base model. Adapters are very lightweight, making it convenient to share, store, and load them.
16+
Transformers integrates directly with the PEFT library through [`~integrations.PeftAdapterMixin`], added to all [`PreTrainedModel`] classes. You can load, add, train, switch, and delete adapters without wrapping your model in a separate [`~peft.PeftModel`]. All non-prompt-learning PEFT methods are supported (LoRA, IA3, AdaLoRA). Prompt-based methods like prompt tuning and prefix tuning require using the [PEFT library](https://huggingface.co/docs/peft/index) directly.
1717

18-
This guide provides a short introduction to the PEFT library and how to use it for training with Transformers. For more details, refer to the PEFT [documentation](https://huggingface.co/docs/peft/index).
18+
Install PEFT to get started. The integration requires `peft >= 0.18.0`.
1919

20-
Install PEFT with the command below.
21-
22-
<hfoptions id="install">
23-
<hfoption id="pip">
24-
25-
```bash
20+
```shell
2621
pip install -U peft
2722
```
2823

29-
</hfoption>
30-
<hfoption id="source">
24+
## Add an adapter
3125

32-
```bash
33-
pip install git+https://github.com/huggingface/peft.git
34-
```
26+
Create a PEFT config, like [`~peft.LoraConfig`] for example, and attach it to a model with [`~integrations.PeftAdapterMixin.add_adapter`].
3527

36-
</hfoption>
37-
</hfoptions>
28+
```py
29+
from peft import LoraConfig, TaskType
30+
from transformers import AutoModelForCausalLM
3831

39-
> [!TIP]
40-
> PEFT currently supports the LoRA, IA3, and AdaLoRA methods for Transformers. To use another PEFT method, such as prompt learning or prompt tuning, use the PEFT library directly.
32+
model = AutoModelForCausalLM.from_pretrained("google/gemma-2-2b")
4133

42-
[Low-Rank Adaptation (LoRA)](https://huggingface.co/docs/peft/conceptual_guides/adapter#low-rank-adaptation-lora) is a very common PEFT method that decomposes the weight matrix into two smaller trainable matrices. Start by defining a [LoraConfig](https://huggingface.co/docs/peft/package_reference/lora#peft.LoraConfig) object with the parameters shown below.
34+
lora_config = LoraConfig(
35+
task_type=TaskType.CAUSAL_LM,
36+
inference_mode=False,
37+
r=8,
38+
lora_alpha=32,
39+
lora_dropout=0.1,
40+
)
4341

44-
```py
45-
from peft import LoraConfig, TaskType, get_peft_model
46-
from transformers import AutoModelForCausalLM
42+
model.add_adapter(lora_config, adapter_name="my_adapter")
43+
```
44+
45+
### Fully fine-tuning specific layers
4746

48-
# create LoRA configuration object
47+
To train additional modules alongside an adapter (for example, the language model head), specify them in `modules_to_save`. `modules_to_save` specifies layers that are fully fine-tuned alongside the adapter, so *all* of their parameters are updated. This is useful when certain layers need updates, for example the language model head (`lm_head`), when adapting a causal LM for sequence classification.
48+
49+
```py
4950
lora_config = LoraConfig(
50-
task_type=TaskType.CAUSAL_LM, # type of task to train on
51-
inference_mode=False, # set to False for training
52-
r=8, # dimension of the smaller matrices
53-
lora_alpha=32, # scaling factor
54-
lora_dropout=0.1 # dropout of LoRA layers
51+
modules_to_save=["lm_head"],
52+
...
5553
)
54+
model.add_adapter(lora_config)
5655
```
5756

58-
Add [LoraConfig](https://huggingface.co/docs/peft/package_reference/lora#peft.LoraConfig) to the model with [`~integrations.PeftAdapterMixin.add_adapter`]. The model is now ready to be passed to [`Trainer`] for training.
57+
### Choosing which layers to adapt
58+
59+
For common architectures (Llama, Gemma, Qwen, etc.), PEFT has predefined default targets (like `q_proj` and `v_proj`), so you don't need to specify `target_modules`. If you want to target different layers, or the model doesn't have predefined targets, pass `target_modules` explicitly as a list of module names or a regex pattern.
5960

6061
```py
61-
model.add_adapter(lora_config, adapter_name="lora_1")
62-
trainer = Trainer(model=model, ...)
63-
trainer.train()
62+
lora_config = LoraConfig(
63+
target_modules=["q_proj", "k_proj"],
64+
...
65+
)
66+
model.add_adapter(lora_config)
6467
```
6568

66-
To add an additional trainable adapter on top of a model with an existing adapter attached, specify the modules you want to train in [modules_to_save()](https://huggingface.co/docs/peft/package_reference/lora#peft.LoraConfig.modules_to_save).
69+
## Training
6770

68-
For example, to train the `lm_head` module on top of a causal language model with a LoRA adapter attached, set `modules_to_save=["lm_head"]`. Add the adapter to the model as shown below, and then pass it to [`Trainer`].
71+
Pass the model with an attached adapter to [`Trainer`] and call [`~Trainer.train`]. [`Trainer`] only updates the adapter parameters (those with `requires_grad=True`) because the base model is frozen.
6972

7073
```py
71-
from transformers import AutoModelForCausalLM
72-
from peft import LoraConfig
74+
from transformers import Trainer, TrainingArguments
7375

74-
model = AutoModelForCausalLM.from_pretrained("google/gemma-2-2b")
76+
training_args = TrainingArguments(
77+
output_dir="./output",
78+
num_train_epochs=3,
79+
per_device_train_batch_size=4,
80+
)
7581

76-
lora_config = LoraConfig(
77-
target_modules=["q_proj", "k_proj"],
78-
modules_to_save=["lm_head"],
82+
trainer = Trainer(
83+
model=model,
84+
args=training_args,
85+
train_dataset=dataset,
7986
)
8087

81-
model.add_adapter(lora_config)
82-
trainer = Trainer(model=model, ...)
8388
trainer.train()
8489
```
8590

86-
Save your adapter with [`~PreTrainedModel.save_pretrained`] to reuse it.
91+
During training, [`Trainer`] checkpoints contain only the adapter weights (`adapter_model.safetensors`) and configuration (`adapter_config.json`), keeping checkpoints small. The base model isn't included.
92+
93+
After training, save the final adapter with [`~PreTrainedModel.save_pretrained`].
94+
95+
```py
96+
model.save_pretrained("./my_adapter")
97+
```
98+
99+
### Resuming from a checkpoint
100+
101+
[`Trainer`] automatically detects adapter checkpoints when resuming. [`Trainer`] scans the checkpoint directory for subdirectories containing adapter weights and reloads each adapter with the correct trainable state.
102+
103+
```py
104+
trainer.train(resume_from_checkpoint="./output/checkpoint-1000")
105+
```
106+
107+
### Distributed training
87108

88-
## Load adapter
109+
PEFT adapters work with distributed training out of the box.
89110

90-
To load an adapter with Transformers, the Hub repository or local directory must contain an `adapter_config.json` file and the adapter weights. Load the adapter with [`~PreTrainedModel.from_pretrained`] or with [`~integrations.PeftAdapterMixin.load_adapter`].
111+
For ZeRO-3, [`Trainer`] passes `exclude_frozen_parameters=True` when saving checkpoints with a PEFT model. Frozen base model weights are skipped. Only the trainable adapter parameters are saved, reducing checkpoint size and save time.
112+
113+
For FSDP, [`Trainer`] updates the FSDP auto-wrap policy to correctly handle LoRA layers. For QLoRA (quantized base model + LoRA), [`Trainer`] also adjusts the mixed precision policy to match the quantization storage dtype.
114+
115+
## Loading an adapter
116+
117+
To load an adapter, the Hub repository or local directory must contain an `adapter_config.json` file and the adapter weights.
91118

92119
<hfoptions id="load">
93120
<hfoption id="from_pretrained">
94121

122+
[`~PreTrainedModel.from_pretrained`] automatically detects adapters. When it finds an `adapter_config.json`, it reads the `base_model_name_or_path` field to load the correct base model, then loads the adapter on top.
123+
95124
```py
96125
from transformers import AutoModelForCausalLM
97126

127+
# Automatically loads the base model and attaches the adapter
98128
model = AutoModelForCausalLM.from_pretrained("klcsp/gemma7b-lora-alpaca-11-v1")
99129
```
100130

101131
</hfoption>
102132
<hfoption id="load_adapter">
103133

134+
To load an adapter onto an existing model, use [`~integrations.PeftAdapterMixin.load_adapter`].
135+
104136
```py
105137
from transformers import AutoModelForCausalLM
106138

@@ -111,9 +143,7 @@ model.load_adapter("klcsp/gemma7b-lora-alpaca-11-v1")
111143
</hfoption>
112144
</hfoptions>
113145

114-
For very large models, it is helpful to load a quantized version of the model in 8 or 4-bit precision to save memory. Transformers supports quantization with its [bitsandbytes](https://huggingface.co/docs/bitsandbytes/index) integration. Specify in [`BitsAndBytesConfig`] whether you want to load a model in 8 or 4-bit precision.
115-
116-
For multiple devices, add `device_map="auto"` to automatically distribute the model across your hardware.
146+
For large models, load a quantized version in 8-bit or 4-bit precision with [bitsandbytes](./quantization/bitsandbytes) to save memory. Add `device_map="auto"` to distribute the model across available hardware.
117147

118148
```py
119149
from transformers import AutoModelForCausalLM, BitsAndBytesConfig
@@ -125,75 +155,91 @@ model = AutoModelForCausalLM.from_pretrained(
125155
)
126156
```
127157

128-
## Set adapter
158+
## Managing multiple adapters
129159

130-
[`~integrations.PeftAdapterMixin.add_adapter`] adds a new adapter to a model. To add a second adapter, the new adapter must be the same type as the first adapter. Use the `adapter_name` parameter to assign a name to the adapter.
160+
A model can hold multiple adapters at once. Add adapters with unique names, and switch between them as needed.
131161

132162
```py
133-
model.add_adapter(lora_config, adapter_name="lora_2")
163+
from peft import LoraConfig
164+
165+
model.add_adapter(LoraConfig(r=8, lora_alpha=32), adapter_name="adapter_1")
166+
model.add_adapter(LoraConfig(r=16, lora_alpha=64), adapter_name="adapter_2")
134167
```
135168

136-
Once added, use [`~integrations.PeftAdapterMixin.set_adapter`] to force a model to use the specified adapter and disable the other adapters.
169+
Use [`~integrations.PeftAdapterMixin.set_adapter`] to activate a specific adapter. The other adapters are disabled but remain in memory.
137170

138171
```py
139-
model.set_adapter("lora_2")
172+
model.set_adapter("adapter_2")
140173
```
141174

142-
## Enable and disable adapter
143-
144-
[`~integrations.PeftAdapterMixin.enable_adapters`] is a broader function that enables *all* adapters attached to a model, and [`~integrations.PeftAdapterMixin.disable_adapters`] disables *all* attached adapters.
175+
[`~integrations.PeftAdapterMixin.enable_adapters`] enables all attached adapters, and [`~integrations.PeftAdapterMixin.disable_adapters`] disables all of them.
145176

146177
```py
147-
model.add_adapter(lora_1)
148-
model.add_adapter(lora_2)
178+
# Disable all adapters for base model inference
179+
model.disable_adapters()
180+
181+
# Re-enable all adapters
149182
model.enable_adapters()
183+
```
150184

151-
# disable all adapters
152-
model.disable_adapters()
185+
Use [`~integrations.PeftAdapterMixin.active_adapters`] to see which adapters are currently active.
186+
187+
```py
188+
model.active_adapters()
189+
# ["adapter_1"]
153190
```
154191

155-
## Hotswapping adapters
192+
Remove adapters you no longer need with [`~integrations.PeftAdapterMixin.delete_adapter`] to free memory.
156193

157-
A common use case when serving multiple adapters is to load one adapter first, generate output, load another adapter, generate more outputs, load another adapter, etc. This can be inefficient, since each time a new adapter is loaded, new memory is reserved; moreover, if the model is compiled with `torch.compile`, it needs to be re-compiled each time a new adapter is used. When switching frequently, the compilation time may never be amortized.
194+
```py
195+
model.delete_adapter("adapter_1")
196+
```
158197

159-
To better support this common workflow, you can "hotswap" a LoRA adapter, to avoid accumulating memory and, in some cases, recompilation. It requires an adapter to already be loaded, and the new adapter weights are swapped in-place for the existing adapter. Note that other PEFT methods are not supported yet, only LoRA.
198+
## Hotswapping adapters
199+
200+
Loading a new adapter each time you serve a request allocates new memory. If the model is compiled with `torch.compile`, each new adapter triggers recompilation. Hotswapping replaces adapter weights in-place, avoiding both issues. Only LoRA adapters are supported.
160201

161-
Pass `hotswap=True` when loading a LoRA adapter to enable this feature. It is important to indicate the name of the existing adapter (`"default"` is the default adapter name) to be swapped.
202+
Pass `hotswap=True` when loading a LoRA adapter to swap its weights into an existing adapter slot. Set `adapter_name` to the name of the adapter to replace (`"default"` is the default adapter name).
162203

163-
```python
204+
```py
164205
model = AutoModel.from_pretrained(...)
165-
# load adapter 1 as normal
166-
model.load_adapter(file_name_adapter_1)
167-
# generate outputs with adapter 1
206+
# Load the first adapter normally
207+
model.load_adapter(adapter_path_1)
208+
# Generate outputs with adapter 1
168209
...
169-
# now hotswap the 2nd adapter
170-
model.load_adapter(file_name_adapter_2, hotswap=True, adapter_name="default")
171-
# generate outputs with adapter 2
210+
# Hotswap the second adapter in-place
211+
model.load_adapter(adapter_path_2, hotswap=True, adapter_name="default")
212+
# Generate outputs with adapter 2
172213
```
173214

174-
For compiled models, it is often necessary to call [`~integrations.peft.PeftAdapterMixin.enable_peft_hotswap`] to avoid recompilation. Call this method *before* loading the first adapter, while `torch.compile` should be called *after* loading the first adapter.
215+
### torch.compile
216+
217+
For compiled models, call [`~integrations.peft.PeftAdapterMixin.enable_peft_hotswap`] *before* loading the first adapter and before compiling.
175218

176-
```python
219+
```py
177220
model = AutoModel.from_pretrained(...)
178-
max_rank = ... # the highest rank among all LoRAs that you want to load
179-
# call *before* compiling and loading the LoRA adapter
221+
max_rank = ... # highest rank among all LoRAs you'll load
180222
model.enable_peft_hotswap(target_rank=max_rank)
181-
model.load_adapter(file_name_1, adapter_name="default")
182-
# optionally compile the model now
223+
model.load_adapter(adapter_path_1, adapter_name="default")
183224
model = torch.compile(model, ...)
184225
output_1 = model(...)
185-
# now you can hotswap the 2nd adapter, use the same name as for the 1st
186-
model.load_adapter(file_name_2, adapter_name="default")
226+
227+
# Hotswap without recompilation
228+
model.load_adapter(adapter_path_2, adapter_name="default")
187229
output_2 = model(...)
188230
```
189231

190-
The `target_rank=max_rank` argument is important for setting the maximum rank among all LoRA adapters that will be loaded. If you have one adapter with rank 8 and another with rank 16, pass `target_rank=16`. You should use a higher value if in doubt. By default, this value is 128.
232+
The `target_rank` argument sets the maximum rank among all LoRA adapters you'll load. If you have adapters with rank 8 and rank 16, pass `target_rank=16`. The default is 128.
191233

192-
By default, hotswapping is disabled and requires you to pass `hotswap=True` to `load_adapter`. However, if you called `enable_peft_hotswap` first, hotswapping will be enabled by default. If you want to avoid using it, you need to pass `hotswap=False`.
234+
After calling `enable_peft_hotswap`, all subsequent `load_adapter` calls hotswap by default. Pass `hotswap=False` explicitly to disable hotswapping.
193235

194-
However, there can be situations where recompilation is unavoidable. For example, if the hotswapped adapter targets more layers than the initial adapter, then recompilation is triggered. Try to load the adapter that targets the most layers first. Refer to the PEFT docs on [hotswapping](https://huggingface.co/docs/peft/main/en/package_reference/hotswap#peft.utils.hotswap.hotswap_adapter) for more details about the limitations of this feature.
236+
Recompilation may still occur if the hotswapped adapter targets more layers than the initial adapter. Load the adapter that targets the most layers first to avoid recompilation.
237+
238+
> [!TIP]
239+
> Wrap your code in `with torch._dynamo.config.patch(error_on_recompile=True)` to detect unexpected recompilation. If you detect recompilation despite following the steps above, open an issue with [PEFT](https://github.com/huggingface/peft/issues) with a reproducible example.
195240
196-
> [!Tip]
197-
> Move your code inside the `with torch._dynamo.config.patch(error_on_recompile=True)` context manager to detect if a model was recompiled. If you detect recompilation despite following all the steps above, please open an issue with [PEFT](https://github.com/huggingface/peft/issues) with a reproducible example.
241+
## Next steps
198242

199-
For an example of how the use of `torch.compile` in combination with hotswapping can improve runtime, check out [this blogpost](https://huggingface.co/blog/lora-fast). Although that example uses Diffusers, similar improvements can be expected here.
243+
- The PEFT [documentation](https://huggingface.co/docs/peft/index) covers the full range of PEFT methods and options.
244+
- The PEFT [hotswapping reference](https://huggingface.co/docs/peft/main/en/package_reference/hotswap#peft.utils.hotswap.hotswap_adapter) details limitations and edge cases.
245+
- A [blog post](https://huggingface.co/blog/lora-fast) benchmarks how `torch.compile` with hotswapping improves runtime.

0 commit comments

Comments
 (0)