minbpe · Issues· 59 open
Open on GitHubLocally synced open issues (discussions stay on GitHub)
- #118
Multilingual Corpus collection utility tool
Updated Sep 3, 2026 - #114
train() raises ValueError: max() arg is an empty sequence on empty or whitespace-only input
Updated Aug 9, 2026 - #115
encode(..., allowed_special="none_raise") raises AssertionError instead of ValueError
Updated Aug 2, 2026 - #85
Python API with C extensions for faster training and encoding
Updated Jun 5, 2026 - #105
get_gpt4_merges()
Updated Nov 23, 2025 - #100
Simpler and faster encoding
Updated Nov 9, 2025 - #74
Notebook Issue In Google Colab
Updated Jul 13, 2025 - #77
What to support GPT-4O tokenizer?
Updated Jul 13, 2025 - #92
LLM is worse at non-English languages
Updated Jul 13, 2025 - #87
Question about Encoder Logic
Updated Dec 10, 2024 - #93
One problem in the annotations of `test_wikipedia_example` in the `tests/test_tokenizer` file
Updated Nov 18, 2024 - #51
counting pairs is inaccurate for repeating tokens?
Updated Jun 7, 2024 - #69
Instead of finding the one pair with the highest frequency and merging it at each step, do the highest N pairs
Updated Jun 7, 2024 - #81
LLM as calc
Updated Jun 6, 2024 - #80
OSS-Fuzz Integration
Updated May 30, 2024