You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Copy file name to clipboardExpand all lines: docs/source/en/main_classes/peft.md
+3-1Lines changed: 3 additions & 1 deletion
Display the source diff
Display the rich diff
Original file line number
Diff line number
Diff line change
@@ -11,13 +11,15 @@ rendered properly in your Markdown viewer.
11
11
12
12
# PEFT
13
13
14
-
The [`~integrations.PeftAdapterMixin`] provides functions from the [PEFT](https://huggingface.co/docs/peft/index) library for managing adapters with Transformers. This mixin currently supports LoRA, IA3, and AdaLora. Prefix tuning methods (prompt tuning, prompt learning) aren't supported because they can't be injected into a torch module.
14
+
The [`~integrations.PeftAdapterMixin`] provides functions from the [PEFT](https://huggingface.co/docs/peft/index) library for managing adapters with Transformers. This mixin supports all non-prompt-learning PEFT methods (LoRA, IA3, AdaLoRA, and others). Prefix tuning methods (prompt tuning, prompt learning) aren't supported because they can't be injected into a torch module.
@@ -9,98 +9,130 @@ specific language governing permissions and limitations under the License.
9
9
rendered properly in your Markdown viewer.
10
10
-->
11
11
12
-
# PEFT
12
+
# Parameter-efficient fine-tuning
13
13
14
-
[[open-in-colab]]
14
+
[Parameter-efficient fine-tuning (PEFT)](https://huggingface.co/docs/peft/index) methods only fine-tune a small number of extra model parameters (adapters) on top of a pretrained model. Because only adapter parameters are updated, the optimizer tracks far fewer gradients and states, reducing memory usage significantly. Adapters are lightweight, making them convenient to share, store, and load.
15
15
16
-
[PEFT](https://huggingface.co/docs/peft/index), a library of parameter-efficient fine-tuning methods, enables training and storing large models on consumer GPUs. These methods only fine-tune a small number of extra model parameters, also known as adapters, on top of the pretrained model. A significant amount of memory is saved because the GPU doesn't need to store the optimizer states and gradients for the pretrained base model. Adapters are very lightweight, making it convenient to share, store, and load them.
16
+
Transformers integrates directly with the PEFT library through [`~integrations.PeftAdapterMixin`], added to all [`PreTrainedModel`] classes. You can load, add, train, switch, and delete adapters without wrapping your model in a separate [`~peft.PeftModel`]. All non-prompt-learning PEFT methods are supported (LoRA, IA3, AdaLoRA). Prompt-based methods like prompt tuning and prefix tuning require using the [PEFT library](https://huggingface.co/docs/peft/index) directly.
17
17
18
-
This guide provides a short introduction to the PEFT library and how to use it for training with Transformers. For more details, refer to the PEFT [documentation](https://huggingface.co/docs/peft/index).
18
+
Install PEFT to get started. The integration requires `peft >= 0.18.0`.
Create a PEFT config, like [`~peft.LoraConfig`] for example, and attach it to a model with [`~integrations.PeftAdapterMixin.add_adapter`].
35
27
36
-
</hfoption>
37
-
</hfoptions>
28
+
```py
29
+
from peft import LoraConfig, TaskType
30
+
from transformers import AutoModelForCausalLM
38
31
39
-
> [!TIP]
40
-
> PEFT currently supports the LoRA, IA3, and AdaLoRA methods for Transformers. To use another PEFT method, such as prompt learning or prompt tuning, use the PEFT library directly.
32
+
model = AutoModelForCausalLM.from_pretrained("google/gemma-2-2b")
41
33
42
-
[Low-Rank Adaptation (LoRA)](https://huggingface.co/docs/peft/conceptual_guides/adapter#low-rank-adaptation-lora) is a very common PEFT method that decomposes the weight matrix into two smaller trainable matrices. Start by defining a [LoraConfig](https://huggingface.co/docs/peft/package_reference/lora#peft.LoraConfig) object with the parameters shown below.
34
+
lora_config = LoraConfig(
35
+
task_type=TaskType.CAUSAL_LM,
36
+
inference_mode=False,
37
+
r=8,
38
+
lora_alpha=32,
39
+
lora_dropout=0.1,
40
+
)
43
41
44
-
```py
45
-
from peft import LoraConfig, TaskType, get_peft_model
To train additional modules alongside an adapter (for example, the language model head), specify them in `modules_to_save`. `modules_to_save` specifies layers that are fully fine-tuned alongside the adapter, so *all* of their parameters are updated. This is useful when certain layers need updates, for example the language model head (`lm_head`), when adapting a causal LM for sequence classification.
48
+
49
+
```py
49
50
lora_config = LoraConfig(
50
-
task_type=TaskType.CAUSAL_LM, # type of task to train on
51
-
inference_mode=False, # set to False for training
52
-
r=8, # dimension of the smaller matrices
53
-
lora_alpha=32, # scaling factor
54
-
lora_dropout=0.1# dropout of LoRA layers
51
+
modules_to_save=["lm_head"],
52
+
...
55
53
)
54
+
model.add_adapter(lora_config)
56
55
```
57
56
58
-
Add [LoraConfig](https://huggingface.co/docs/peft/package_reference/lora#peft.LoraConfig) to the model with [`~integrations.PeftAdapterMixin.add_adapter`]. The model is now ready to be passed to [`Trainer`] for training.
57
+
### Choosing which layers to adapt
58
+
59
+
For common architectures (Llama, Gemma, Qwen, etc.), PEFT has predefined default targets (like `q_proj` and `v_proj`), so you don't need to specify `target_modules`. If you want to target different layers, or the model doesn't have predefined targets, pass `target_modules` explicitly as a list of module names or a regex pattern.
To add an additional trainable adapter on top of a model with an existing adapter attached, specify the modules you want to train in [modules_to_save()](https://huggingface.co/docs/peft/package_reference/lora#peft.LoraConfig.modules_to_save).
69
+
## Training
67
70
68
-
For example, to train the `lm_head` module on top of a causal language model with a LoRA adapter attached, set `modules_to_save=["lm_head"]`. Add the adapter to the model as shown below, and then pass it to [`Trainer`].
71
+
Pass the model with an attached adapter to [`Trainer`] and call [`~Trainer.train`]. [`Trainer`] only updates the adapter parameters (those with `requires_grad=True`) because the base model is frozen.
69
72
70
73
```py
71
-
from transformers import AutoModelForCausalLM
72
-
from peft import LoraConfig
74
+
from transformers import Trainer, TrainingArguments
73
75
74
-
model = AutoModelForCausalLM.from_pretrained("google/gemma-2-2b")
76
+
training_args = TrainingArguments(
77
+
output_dir="./output",
78
+
num_train_epochs=3,
79
+
per_device_train_batch_size=4,
80
+
)
75
81
76
-
lora_config = LoraConfig(
77
-
target_modules=["q_proj", "k_proj"],
78
-
modules_to_save=["lm_head"],
82
+
trainer = Trainer(
83
+
model=model,
84
+
args=training_args,
85
+
train_dataset=dataset,
79
86
)
80
87
81
-
model.add_adapter(lora_config)
82
-
trainer = Trainer(model=model, ...)
83
88
trainer.train()
84
89
```
85
90
86
-
Save your adapter with [`~PreTrainedModel.save_pretrained`] to reuse it.
91
+
During training, [`Trainer`] checkpoints contain only the adapter weights (`adapter_model.safetensors`) and configuration (`adapter_config.json`), keeping checkpoints small. The base model isn't included.
92
+
93
+
After training, save the final adapter with [`~PreTrainedModel.save_pretrained`].
94
+
95
+
```py
96
+
model.save_pretrained("./my_adapter")
97
+
```
98
+
99
+
### Resuming from a checkpoint
100
+
101
+
[`Trainer`] automatically detects adapter checkpoints when resuming. [`Trainer`] scans the checkpoint directory for subdirectories containing adapter weights and reloads each adapter with the correct trainable state.
PEFT adapters work with distributed training out of the box.
89
110
90
-
To load an adapter with Transformers, the Hub repository or local directory must contain an `adapter_config.json` file and the adapter weights. Load the adapter with [`~PreTrainedModel.from_pretrained`] or with [`~integrations.PeftAdapterMixin.load_adapter`].
111
+
For ZeRO-3, [`Trainer`] passes `exclude_frozen_parameters=True` when saving checkpoints with a PEFT model. Frozen base model weights are skipped. Only the trainable adapter parameters are saved, reducing checkpoint size and save time.
112
+
113
+
For FSDP, [`Trainer`] updates the FSDP auto-wrap policy to correctly handle LoRA layers. For QLoRA (quantized base model + LoRA), [`Trainer`] also adjusts the mixed precision policy to match the quantization storage dtype.
114
+
115
+
## Loading an adapter
116
+
117
+
To load an adapter, the Hub repository or local directory must contain an `adapter_config.json` file and the adapter weights.
91
118
92
119
<hfoptionsid="load">
93
120
<hfoptionid="from_pretrained">
94
121
122
+
[`~PreTrainedModel.from_pretrained`] automatically detects adapters. When it finds an `adapter_config.json`, it reads the `base_model_name_or_path` field to load the correct base model, then loads the adapter on top.
123
+
95
124
```py
96
125
from transformers import AutoModelForCausalLM
97
126
127
+
# Automatically loads the base model and attaches the adapter
98
128
model = AutoModelForCausalLM.from_pretrained("klcsp/gemma7b-lora-alpaca-11-v1")
99
129
```
100
130
101
131
</hfoption>
102
132
<hfoptionid="load_adapter">
103
133
134
+
To load an adapter onto an existing model, use [`~integrations.PeftAdapterMixin.load_adapter`].
For very large models, it is helpful to load a quantized version of the model in 8 or 4-bit precision to save memory. Transformers supports quantization with its [bitsandbytes](https://huggingface.co/docs/bitsandbytes/index) integration. Specify in [`BitsAndBytesConfig`] whether you want to load a model in 8 or 4-bit precision.
115
-
116
-
For multiple devices, add `device_map="auto"` to automatically distribute the model across your hardware.
146
+
For large models, load a quantized version in 8-bit or 4-bit precision with [bitsandbytes](./quantization/bitsandbytes) to save memory. Add `device_map="auto"` to distribute the model across available hardware.
117
147
118
148
```py
119
149
from transformers import AutoModelForCausalLM, BitsAndBytesConfig
@@ -125,75 +155,91 @@ model = AutoModelForCausalLM.from_pretrained(
125
155
)
126
156
```
127
157
128
-
## Set adapter
158
+
## Managing multiple adapters
129
159
130
-
[`~integrations.PeftAdapterMixin.add_adapter`] adds a new adapter to a model. To add a second adapter, the new adapter must be the same type as the first adapter. Use the `adapter_name` parameter to assign a name to the adapter.
160
+
A model can hold multiple adapters at once. Add adapters with unique names, and switch between them as needed.
Once added, use [`~integrations.PeftAdapterMixin.set_adapter`] to force a model to use the specified adapter and disable the other adapters.
169
+
Use [`~integrations.PeftAdapterMixin.set_adapter`] to activate a specific adapter. The other adapters are disabled but remain in memory.
137
170
138
171
```py
139
-
model.set_adapter("lora_2")
172
+
model.set_adapter("adapter_2")
140
173
```
141
174
142
-
## Enable and disable adapter
143
-
144
-
[`~integrations.PeftAdapterMixin.enable_adapters`] is a broader function that enables *all* adapters attached to a model, and [`~integrations.PeftAdapterMixin.disable_adapters`] disables *all* attached adapters.
175
+
[`~integrations.PeftAdapterMixin.enable_adapters`] enables all attached adapters, and [`~integrations.PeftAdapterMixin.disable_adapters`] disables all of them.
145
176
146
177
```py
147
-
model.add_adapter(lora_1)
148
-
model.add_adapter(lora_2)
178
+
# Disable all adapters for base model inference
179
+
model.disable_adapters()
180
+
181
+
# Re-enable all adapters
149
182
model.enable_adapters()
183
+
```
150
184
151
-
# disable all adapters
152
-
model.disable_adapters()
185
+
Use [`~integrations.PeftAdapterMixin.active_adapters`] to see which adapters are currently active.
186
+
187
+
```py
188
+
model.active_adapters()
189
+
# ["adapter_1"]
153
190
```
154
191
155
-
## Hotswapping adapters
192
+
Remove adapters you no longer need with [`~integrations.PeftAdapterMixin.delete_adapter`] to free memory.
156
193
157
-
A common use case when serving multiple adapters is to load one adapter first, generate output, load another adapter, generate more outputs, load another adapter, etc. This can be inefficient, since each time a new adapter is loaded, new memory is reserved; moreover, if the model is compiled with `torch.compile`, it needs to be re-compiled each time a new adapter is used. When switching frequently, the compilation time may never be amortized.
194
+
```py
195
+
model.delete_adapter("adapter_1")
196
+
```
158
197
159
-
To better support this common workflow, you can "hotswap" a LoRA adapter, to avoid accumulating memory and, in some cases, recompilation. It requires an adapter to already be loaded, and the new adapter weights are swapped in-place for the existing adapter. Note that other PEFT methods are not supported yet, only LoRA.
198
+
## Hotswapping adapters
199
+
200
+
Loading a new adapter each time you serve a request allocates new memory. If the model is compiled with `torch.compile`, each new adapter triggers recompilation. Hotswapping replaces adapter weights in-place, avoiding both issues. Only LoRA adapters are supported.
160
201
161
-
Pass `hotswap=True` when loading a LoRA adapter to enable this feature. It is important to indicate the name of the existing adapter (`"default"` is the default adapter name) to be swapped.
202
+
Pass `hotswap=True` when loading a LoRA adapter to swap its weights into an existing adapter slot. Set `adapter_name` to the name of the adapter to replace (`"default"` is the default adapter name).
For compiled models, it is often necessary to call [`~integrations.peft.PeftAdapterMixin.enable_peft_hotswap`] to avoid recompilation. Call this method *before* loading the first adapter, while `torch.compile` should be called *after* loading the first adapter.
215
+
### torch.compile
216
+
217
+
For compiled models, call [`~integrations.peft.PeftAdapterMixin.enable_peft_hotswap`]*before* loading the first adapter and before compiling.
175
218
176
-
```python
219
+
```py
177
220
model = AutoModel.from_pretrained(...)
178
-
max_rank =...# the highest rank among all LoRAs that you want to load
179
-
# call *before* compiling and loading the LoRA adapter
221
+
max_rank =...# highest rank among all LoRAs you'll load
The `target_rank=max_rank` argument is important for setting the maximum rank among all LoRA adapters that will be loaded. If you have one adapter with rank 8 and another with rank 16, pass `target_rank=16`. You should use a higher value if in doubt. By default, this value is 128.
232
+
The `target_rank` argument sets the maximum rank among all LoRA adapters you'll load. If you have adapters with rank 8 and rank 16, pass `target_rank=16`. The default is 128.
191
233
192
-
By default, hotswapping is disabled and requires you to pass `hotswap=True` to `load_adapter`. However, if you called `enable_peft_hotswap` first, hotswapping will be enabled by default. If you want to avoid using it, you need to pass `hotswap=False`.
234
+
After calling `enable_peft_hotswap`, all subsequent `load_adapter` calls hotswap by default. Pass `hotswap=False` explicitly to disable hotswapping.
193
235
194
-
However, there can be situations where recompilation is unavoidable. For example, if the hotswapped adapter targets more layers than the initial adapter, then recompilation is triggered. Try to load the adapter that targets the most layers first. Refer to the PEFT docs on [hotswapping](https://huggingface.co/docs/peft/main/en/package_reference/hotswap#peft.utils.hotswap.hotswap_adapter) for more details about the limitations of this feature.
236
+
Recompilation may still occur if the hotswapped adapter targets more layers than the initial adapter. Load the adapter that targets the most layers first to avoid recompilation.
237
+
238
+
> [!TIP]
239
+
> Wrap your code in `with torch._dynamo.config.patch(error_on_recompile=True)` to detect unexpected recompilation. If you detect recompilation despite following the steps above, open an issue with [PEFT](https://github.com/huggingface/peft/issues) with a reproducible example.
195
240
196
-
> [!Tip]
197
-
> Move your code inside the `with torch._dynamo.config.patch(error_on_recompile=True)` context manager to detect if a model was recompiled. If you detect recompilation despite following all the steps above, please open an issue with [PEFT](https://github.com/huggingface/peft/issues) with a reproducible example.
241
+
## Next steps
198
242
199
-
For an example of how the use of `torch.compile` in combination with hotswapping can improve runtime, check out [this blogpost](https://huggingface.co/blog/lora-fast). Although that example uses Diffusers, similar improvements can be expected here.
243
+
- The PEFT [documentation](https://huggingface.co/docs/peft/index) covers the full range of PEFT methods and options.
244
+
- The PEFT [hotswapping reference](https://huggingface.co/docs/peft/main/en/package_reference/hotswap#peft.utils.hotswap.hotswap_adapter) details limitations and edge cases.
245
+
- A [blog post](https://huggingface.co/blog/lora-fast) benchmarks how `torch.compile` with hotswapping improves runtime.
0 commit comments