Skip to content

SentencePiece Tokenizer does not encode </s> as special token #383

Description

@le1nux

When tokenizing the </s> token, it does not get tokenized into a single special token, instead it gets split into multiple subtokens.

Activity

  1. le1nux commented on Jul 6, 2025

    @le1nux
    MemberAuthor
  2. le1nux commented on Jul 18, 2025

    @le1nux
    MemberAuthor

    This was a good hint here:
    google/sentencepiece#667 (comment)

    I wonder if this creates issues for our instruction tuning, where we do play around with the special tokens.
    What's your opinion on this @lllAlexanderlll ?

  3. lllAlexanderlll commented on Jul 18, 2025

    @lllAlexanderlll
    Contributor

    Yes, that is problematic, if we want to use the token - we could use any special token e.g. the "placeholder_tokens"
    However, adding new tokens is not possible yet due to #208. Currently, a given single token is marked as special and hence tokenized prior to non-special tokens, which is needed to reliably not resolve prefixed tokens and special tokens into a single token.

  4. added theissue type on Oct 8, 2025
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    bugSomething isn't working

    Type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions