Skip to content

fix(perplexity): use all_special_tokens for GPT-2 tokenizer compatibility - #789

Open
akaashsa wants to merge 1 commit into
huggingface:mainfrom
akaashsa:fix/perplexity-gpt2-tokenizer
Open

fix(perplexity): use all_special_tokens for GPT-2 tokenizer compatibility#789
akaashsa wants to merge 1 commit into
huggingface:mainfrom
akaashsa:fix/perplexity-gpt2-tokenizer

Conversation

@akaashsa

Copy link
Copy Markdown

GPT-2's slow tokenizer (GPT2Tokenizer) does not expose special_tokens_map_extended, causing an AttributeError when computing perplexity with batch_size > 1.

Additionally, the pad_token guard was conditioned on batch_size > 1, but the tokenizer is always called with padding=True, meaning pad_token is required even for batch_size=1 when multiple predictions are passed.

Fix:

  • Replace special_tokens_map_extended.values() with all_special_tokens, which is defined on PreTrainedTokenizerBase and returns a flat list of strings for both slow and fast tokenizers.
  • Remove the batch_size > 1 condition so pad_token is set whenever it is missing, matching the actual padding requirement.

Fixes #766

…lity

GPT-2's slow tokenizer (GPT2Tokenizer) does not expose
special_tokens_map_extended, causing an AttributeError when computing
perplexity with batch_size > 1.

Additionally, the pad_token guard was conditioned on batch_size > 1, but
the tokenizer is always called with padding=True, meaning pad_token is
required even for batch_size=1 when multiple predictions are passed.

Fix:
- Replace special_tokens_map_extended.values() with all_special_tokens,
  which is defined on PreTrainedTokenizerBase and returns a flat list of
  strings for both slow and fast tokenizers.
- Remove the batch_size > 1 condition so pad_token is set whenever it
  is missing, matching the actual padding requirement.

Fixes huggingface#766
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Perplexity metric fails with GPT-2 tokenizer

1 participant