Enhancement / Discussion: Benchmark and Support for Arabic Language Parsing via arafix
Hi @baidu team,
First of all, congratulations on the release of Unlimited-OCR! The long-horizon document parsing capabilities and integration with SGLang/vLLM look very promising.
I am the maintainer of arafix, an open-source Python tool dedicated to Arabic text recovery, layout diagnostic, and evaluation for native PDFs and complex document streams.
Given the unique challenges of Arabic document parsing (right-to-left layout, complex ligatures, letter-joining rules, and encoding/stream issues in native PDFs), I would love to explore how Unlimited-OCR handles Arabic text extraction and offer my support to test/benchmark Arabic language performance.
Potential Areas of Collaboration / Testing:
- Arabic OCR & Parsing Benchmarking: Evaluating Unlimited-OCR on complex Arabic layouts, multi-column articles, and documents with mixed Arabic/English text.
- Post-Processing & Quality Verification: Leveraging
arafixto validate structural fidelity, character connectivity, and text directionality on extracted output. - Dataset / Edge-Case Sharing: Providing representative Arabic native PDF samples and edge cases to help improve model robustness for the Arabic language.
I would be happy to run some benchmark tests on a dataset of Arabic documents using Unlimited-OCR and share the detailed evaluation results here if the team is interested.
Looking forward to your thoughts!
Best regards,
Elias Sharar
Source: baidu/Unlimited-OCR