Repository navigation
Does T5 actually utilize Encoder prefixes in PrefixTuning? past_key_values seems ignored by the Encoder. #2974
Replies: 2 comments 3 replies
|
Thanks for opening this discussion. Your observation is correct, in encoder-decoder models, prefix tuning does not affect the encoder. The kv cache (aka I have asked internally if there is a different way we could adapt the encoder, but possibly there is no solution. I'll update once I have an answer. |
|
Your suspicion is correct the T5 encoder does not actually use the prefix Why the encoder ignores past_key_values In HuggingFace's T5 implementation, This was explicitly confirmed in transformers issue #15591 where the original reporter had the exact same use case (Li & Liang 2021 prefix tuning) and a maintainer confirmed the encoder cannot use How this deviates from the original paper The Li & Liang (2021) prefix tuning paper (figure 2) shows both encoder and decoder receiving separate trainable prefixes. PEFT's implementation for Current state There have been active PRs to address this (transformers PR #34312, peft PR #2096) but the conclusion from the PEFT team was that encoder-decoder models work fundamentally differently for prefix tuning and the fix needs to live on the PEFT side rather than transformers. The transformers PR was closed with a note that a PEFT-side fix was found instead. Practical implication If you are training with Workarounds If you genuinely need encoder side prefixes, the options are:
So to directly answer your question: no, the encoder prefixes are not being utilized in PEFT's T5 prefix tuning. The |
Uh oh!
There was an error while loading. Please reload this page.
System Info
Description
I am exploring Prefix-Tuning on a Seq2Seq model (T5/Flan-T5) and I noticed a potential discrepancy between the paper's implementation and the current codebase regarding the Encoder's prefix.
According to the [Prefix-Tuning paper](https://arxiv.org/abs/2101.00190), prefixes should be prepended to both the Encoder and the Decoder for Encoder-Decoder architectures.
I configured my PEFT model with
num_transformer_submodules=2to generate prompts for both components.However, tracing the code:
peft/src/peft/peft_model.py(PeftModelForSeq2SeqLM):The code generates the prompt and passes it into
kwargs["past_key_values"]:transformers/models/t5/modeling_t5.py(T5Model):The
forwardmethod separates the execution flow. It callsself.encoderwithout passingpast_key_values.Reference Line: [modeling_t5.py#L991 (v5.0.0rc1)](https://github.com/huggingface/transformers/blob/v5.0.0rc1/src/transformers/models/t5/modeling_t5.py#L991)
Question
It seems that even if
num_transformer_submodules=2is set, thepast_key_valuesintended for the Encoder are effectively discarded before reaching the encoder's attention layers.PrefixTuningon T5 effectively only works on the Decoder?transformersthat automatically handles this injection that I might have missed?Any clarification would be appreciated. Thanks!
All reactions