百科.dev
全部条目趋势榜开源项目技术资讯提交条目
登录
< 返回工具列表
O

OCRmyPDF

> 编程语言
开源

OCRmyPDF 在扫描的 PDF 文件上添加一个 OCR 文本层, 允许对其进行搜索

34.3K stars0 点赞0 次浏览
访问官网GitHub

工具介绍

OCRmyPDF 在扫描的 PDF 文件上添加一个 OCR 文本层, 允许对其进行搜索

[!][PyPI版本]pypi. [Homebrew版本][homebrew]. [Read TheDocs][docs]. [ZPython版本][pyversions][pypi]:https://img.shields.io/pypi/v/ocrmypdf.svg "PyPI版本"[homebrew]:https://img.shields.io/homebrew/v/ocrmypdf.svg "Homebrew版本"[docs]:https://readthedocs.org/project/ocrmypdf/badge/?version=late "pyvers]:https://img.shield.io/pypypyvers/o/ocyrmypdf/ "支持的ZZCRPD的版本" OCRP bash ocramypdf # 这是一个可脚写的命令行程序 - l eng+fra # 它支持多种语言 -- -- rotate-pages # 它可以修正被错误旋转的页面 -- -- Deskew # 它可以处理扭曲的PDF! -标题"我的PDF"# 它可以改变输出元数据 - jobs 4 # 它默认使用多个核心 - 输出型 pdfa # 它默认输入罐生成PDF/A..

核心特点

  • •Generates a searchable PDF/A file from a regular PDF
  • •Places OCR text accurately below the image to ease copy / paste
  • •Keeps the exact resolution of the original embedded images
  • •When possible, inserts OCR information as a "lossless" operation without disrupting any other content
  • •Optimizes PDF images, often producing files smaller than the input file
  • •If requested, deskews and/or cleans the image before performing OCR
  • •Validates input and output files
  • •Distributes work across all available CPU cores
  • •Uses Tesseract OCR engine to recognize more than 100 languages
  • •Keeps your private data private.

> 标签

Pythonimage-processingocrpdfpython

暂无评论,来聊聊你的看法吧

> 工具信息

发布日期2026年8月1日
最后更新2026年9月9日
分类编程语言
定价开源

> 相关工具

T
TypeScript
JavaScript 的超集,为前端与全栈提供静态类型
P
Python
通用编程语言,广泛用于 Web、数据与 AI
G
Go
Google 推出的简洁高效系统语言
百科.dev

开发者百科,帮助你快速发现语言、框架、数据库、DevOps 与云原生等优质开发者工具。

快捷入口

  • 首页
  • 全部条目
  • 趋势榜
  • 开源项目

关于我们

  • 关于我们
  • 社区公约
  • 技术资讯

参与我们

发现好用的开发者工具?欢迎提交分享。

提交条目
© 2026 开发者百科 baike.dev数据每日更新 · 发现优质开发者工具