opendataloader-pdf · Issues· 95 open
Open on GitHubLocally synced open issues (discussions stay on GitHub)
- #732
filterTinyText drops spaces, hyphens and em dashes at normal font sizes
Updated Sep 16, 2026 - #694
StringIndexOutOfBoundsException in TableStructureNormalizer aborts whole document (empty TextChunk substring)
bugfixed-in-devUpdated Sep 11, 2026 - #720
[Bug] Hybrid auto mode: TableFormer (hardcoded ACCURATE) collapses table rows on dense datasheets — orphan-cell fallback corrupts structure (cf. #627)
bugUpdated Sep 11, 2026 - #711
Python/Node wrappers silently drop empty-string hybrid_chunk_size
Updated Sep 11, 2026 - #708
test: add sync check for hybrid-chunk-size default across bindings
Updated Sep 11, 2026 - #709
test: assert hybrid CLI flags appear in the built JAR --help
Updated Sep 11, 2026 - #707
test: cover hybridChunkSize forwarding in Node bindings
Updated Sep 11, 2026 - #706
test: drive the production hybrid chunking loop with a configured chunk size
Updated Sep 11, 2026 - #725
[Bug] Hybrid docling-fast fails on Windows when the user profile path contains non-ASCII characters
bugUpdated Sep 11, 2026 - #695
Chart gridlines are detected as a table, producing a /Table with only empty cells
Updated Aug 22, 2026 - #568
[M1] Cap page-render DPI to bound single-page peak memory (the real lever for #458)
enhancementmemoryfollow-upfixed-in-devUpdated Aug 18, 2026 - #686
opendataloarder gets stuck
bugUpdated Aug 17, 2026 - #652
After upgrading to latest version 2.50 seeing issues in image extraction,extra grey scale image seen
bugUpdated Aug 17, 2026 - #676
hyperlinks and links are not handled properly in converted md
bugUpdated Aug 11, 2026 - #674
Add `py.typed` to Python package to denote that the package supports typing
enhancementUpdated Aug 10, 2026