tokenizers · Issues· 219 open
Open on GitHubLocally synced open issues (discussions stay on GitHub)
- #2236
[BUG] | The docs listing page for tokenizers is throwing 404
StaleUpdated Sep 18, 2026 - #2320
BpeTrainer emits a merge for a pair that occurs zero times (negative pair_counts sign-extended by `as u64`)
StaleUpdated Sep 18, 2026 - #2412
BertWordPieceTokenizer ignores wordpieces_prefix when loading a vocabulary
Updated Sep 16, 2026 - #2094
[Security] Load-time process abort when loading a malicious tokenizer.json (BPE merge buffer overrun)
Updated Sep 11, 2026 - #1663
Inconsistent behaviour of `PreTrainedTokenizerFast`s on diacritics marked texts
bugUpdated Sep 4, 2026 - #1784
Add cp314t tests, mark extension modules as compatible, and ship free-threaded wheels
Updated Sep 2, 2026 - #1975
[RFC] Korean Tokenization: Jamo Decomposition as Pre-tokenizer (A Blueprint for Compositional Scripts)
Feature RequestUpdated Aug 25, 2026 - #2334
`Precompiled` normalizer diverges from SentencePiece: shortest-prefix match + discarded remainder silently drops codepoints
Updated Aug 23, 2026 - #1581
[building on windows] onig_sys/oniguruma two or more data types in declaration specifiers
Updated Jul 30, 2026 - #1973
Make huggingface-hub an optional dependency (extras)
Feature RequestUpdated Jul 28, 2026 - #2225
[Feature/Security] Support trusted and untrusted special-token spans
Updated Jul 21, 2026 - #1791
SyntaxWarning: invalid escape sequence '\w'
Updated Jul 18, 2026 - #2210
EncodingVisualizer: data-* attributes built from token text are not HTML-escaped (XSS gap left by #1937)
Updated Jul 17, 2026 - #1407
How to add byte_fallback tokens?
bytefallbackFeature RequestUpdated Jul 16, 2026 - #1902
Guide: Compiling `tokenizers` on Android/Termux
Updated Jul 14, 2026