Enhancement / Discussion: Benchmark and Support for Arabic Language Parsing via arafix

Author: bio-colabCreated Aug 3, 2026Updated Aug 3, 2026

Hi @baidu team,

First of all, congratulations on the release of Unlimited-OCR! The long-horizon document parsing capabilities and integration with SGLang/vLLM look very promising.

I am the maintainer of arafix, an open-source Python tool dedicated to Arabic text recovery, layout diagnostic, and evaluation for native PDFs and complex document streams.

Given the unique challenges of Arabic document parsing (right-to-left layout, complex ligatures, letter-joining rules, and encoding/stream issues in native PDFs), I would love to explore how Unlimited-OCR handles Arabic text extraction and offer my support to test/benchmark Arabic language performance.

Potential Areas of Collaboration / Testing:

  1. Arabic OCR & Parsing Benchmarking: Evaluating Unlimited-OCR on complex Arabic layouts, multi-column articles, and documents with mixed Arabic/English text.
  2. Post-Processing & Quality Verification: Leveraging arafix to validate structural fidelity, character connectivity, and text directionality on extracted output.
  3. Dataset / Edge-Case Sharing: Providing representative Arabic native PDF samples and edge cases to help improve model robustness for the Arabic language.

I would be happy to run some benchmark tests on a dataset of Arabic documents using Unlimited-OCR and share the detailed evaluation results here if the team is interested.

Looking forward to your thoughts!

Best regards,
Elias Sharar