tesseract · Issues· 484 open
Open on GitHubLocally synced open issues (discussions stay on GitHub)
- #4591
write runtime parameters to HOCR output
Updated Sep 16, 2026 - #4621
OCR does not recognize redaction overlay text after image redactions are applied
layout analysisUpdated Sep 13, 2026 - #4618
lstmtraining --convert_to_int leaves the ADAM training_flags_ bit set, unlike official int-optimized traineddata
bugUpdated Sep 3, 2026 - #4595
Regression: Tesseract 5.5.3 recognizes uppercase 'I' as '|' while 5.5.0 recognizes it correctly
Updated Aug 18, 2026 - #1714
[Feature Request] Table structure extraction at the API
feature requestaccuracytablesUpdated Aug 5, 2026 - #4589
Training unicharset_extractor broken in 5.5.3
Updated Jul 29, 2026 - #4504
Installation binaries are missing from releases page
feature requestwindowsUpdated Jul 25, 2026 - #4498
Plans for Tesseract in 2026
plansUpdated Jul 25, 2026 - #4582
textord noise filter silently drops entire text lines in dotted Arabic-script languages (verified: Persian, Arabic, Urdu); textord_noise_rejrows=0 recovers them
Updated Jul 15, 2026 - #4580
Request: enable Private Vulnerability Reporting / security contact
bugUpdated Jul 12, 2026 - #3446
Page numbers not detected in various cases
accuracylayout analysisUpdated Jul 8, 2026 - #2738
Duplicate Characters in Output Stream
accuracydiplopiaUpdated Jul 5, 2026 - #4544
Arabic words are sometimes reversed in OCR output
RTLUpdated Jun 19, 2026 - #4458
LSTM models recognize random characters instead of asterisk (*)
accuracytraineddatalegacyUpdated Jun 19, 2026 - #4430
Tesseract takes much time when processing image
binarizationUpdated Jun 19, 2026