anydoc · Issues· 94 open
Open on GitHubLocally synced open issues (discussions stay on GitHub)
- #170
anydoc-wasm still ships pdf-inspector 1.14.2 — the RTL extraction fix (pdf-inspector#440) isn't included
Updated Sep 15, 2026 - #173
Multi-column tables collapse into unreadable single-column blobs during PDF conversion
Updated Sep 12, 2026 - #172
Ligature characters (fi/fl/ffi) are dropped instead of expanded when extracting PDF text
Updated Sep 12, 2026 - #167
OOXML: recoverable allocation failure in Package::part is classified as malformed
Updated Sep 9, 2026 - #157
Feature request: local OCR provider hook (--ocr command) - delegate OCR to a user-supplied engine
Updated Sep 6, 2026 - #128
Support email file formats: `.eml` and `.msg`
Updated Sep 4, 2026 - #162
0.2.4's needsOcr gate refuses born-digital PDFs that 0.2.3 converts correctly
Updated Sep 3, 2026 - #41
docx to md puts a backslash between number and period in ordered lists
questionneeds-reproUpdated Sep 2, 2026 - #156
xlsx: resource limits hit on real workbooks with excess whitespace
Updated Sep 1, 2026 - #155
Node binding retains large XLSX RSS high-water per libuv worker
Updated Sep 1, 2026 - #146
Inbuilt Image Parser with Local OCR (Unlimited OCR, PaddleOCR etc.)
Updated Sep 1, 2026 - #144
0.2.4: one image-only page makes the whole PDF convert to nothing, including the pages that do have text
Updated Aug 31, 2026 - #139
XLSX embedded images are missing from Document.assets
Updated Aug 29, 2026 - #63
Expose embedded image assets in markdown output
enhancementP2Updated Aug 29, 2026 - #52
Feature request: Support for HTML / MHTML files (e.g. Jira exports)
enhancementP3Updated Aug 28, 2026