百科.dev
全部条目AI 编程趋势榜开源项目技术资讯提交条目
登录
< 返回工具列表
P

paperless-gpt

> AI 编程
开源

使用 LLM 和 LLM Vision (OCR) 处理无纸化 - 由 AI 提供支持的文档数字化

2.6K stars0 点赞0 次浏览
访问官网GitHub

工具介绍

使用 LLM 和 LLM Vision (OCR) 处理无纸化 - 由 AI 提供支持的文档数字化

paperless-gpt

Maintained by Icereed. Proudly supported by BubbleTax.de – automated, BMF-compliant tax reports for Interactive Brokers traders in Germany.


paperless-gpt seamlessly pairs with [paperless-ngx][paperless-ngx] to generate AI-powered document titles and tags, saving you hours of manual sorting. While other tools may offer AI chat features, paperless-gpt stands out by supercharging OCR with LLMs-ensuring high accuracy, even with tricky scans. If you're craving next-level text extraction and effortless document organization, this is your solution.

https://github.com/user-attachments/assets/bd5d38b9-9309-40b9-93ca-918dfa4f3fd4

❤️ Support This Project
If paperless-gpt is helping you organize your documents and saving you time, please consider sponsoring its development. Your support helps ensure continued improvements and maintenance!


Key Highlights

  1. LLM-Enhanced OCR
    Harness Large Language Models (OpenAI or Ollama) for better-than-traditional OCR—turn messy or low-quality scans into context-aware, high-fidelity text.

  2. Use specialized AI OCR services

    • LLM OCR: Use OpenAI or Ollama to extract text from images.
    • Google Document AI: Leverage Google's powerful Document AI for OCR tasks.
    • Azure Document Intelligence: Use Microsoft's enterprise OCR solution.
    • Docling Server: Self-hosted OCR and document conversion service
  3. Automatic Title, Tag & Created Date Generation
    No more guesswork. Let the AI do the naming and categorizing. You can easily review suggestions and refine them if needed.

  4. Supports reasoning models in Ollama
    Greatly enhance accuracy by using a reasoning model like qwen3:8b. The perfect tradeoff between privacy and performance! Of course, if you got enough GPUs or NPUs, a bigger model will enhance the experience.

  5. Automatic Correspondent Generation
    Automatically identify and generate correspondents from your documents, making it easier to track and organize your communications.

  6. Automatic Custom Field Generation
    Extract and populate custom fields from your documents. Configure which fields to target and how they should be filled. This feature must be enabled in the settings, and you must select at least one custom field for it to function. Three write modes are available:

    • Append: This is the safest option: It only adds new fields that do not already exist on the document. It will never overwrite an existing field, even if it's empty.
    • Update: Adds new fields and overwrites existing fields with new suggestions. Fields on the document that don't have a new suggestion are left untouched.
    • Replace: Deletes all existing custom fields on the document and replaces them entirely with the suggested fields.
  7. Searchable & Selectable PDFs
    Generate PDFs with transparent text layers positioned accurately over each word, making your documents both searchable and selectable while preserving the original appearance.

  8. Extensive Customization

    • Customizable Prompts via Web UI: Tweak and manage all AI prompts for titles, tags, correspondents, and more directly within the web interface under the "Settings" menu. The application uses a safe default_prompts and prompts directory structure, ensuring your customizations are persistent.
    • Tagging: Decide how documents get tagged—manually, automatically, or via OCR-based flows.
    • PDF Processing: Configure how OCR-enhanced PDFs are handled, with options to save locally or upload to paperless-ngx.
  9. Simple Docker Deployment
    A few environment variables, and you're off! Compose it alongside paperless-ngx with minimal fuss.

  10. Unified Web UI

    • Manual Review: Approve or tweak AI's suggestions.
    • Auto Processing: Focus only on edge cases while the rest is sorted for you.
  11. Ad-hoc Document Analysis Perform ad-hoc analysis on a selection of documents using a custom prompt. Gain quick insights, summaries, or extract specific information from multiple documents at once.


Table of Contents

  • paperless-gpt
    • Key Highlights
    • Table of Contents
    • Getting Started
      • Prerequisites
      • Security
      • Installation
        • Docker Compose
        • Manual Setup
    • OCR Providers
      • 1. LLM-based OCR (Default)
      • 2. Azure Document Intelligence
      • 3. Google Document AI
      • 4. Docling Server
    • OCR Processing Modes
      • Image Mode (Default)
      • PDF Mode
      • Whole PDF Mode
      • Provider Compatibility
      • Existing OCR Detection
    • Enhanced OCR Features
      • PDF Text Layer Generation
      • Local File Saving
      • PDF Upload to paperless-ngx
      • Metadata Copying Limitations
      • Safety Features
      • Usage Recommendations
    • Configuration
      • Environment Variables
      • Using a Different AI Provider
      • Custom Prompt Templates
        • Template Variables
    • LLM-Based OCR: Compare for Yourself
      • Example 1
      • Example 2
      • How It Works
    • Usage
    • Troubleshooting
      • Working with Local LLMs
        • Token Management
      • PDF Processing Issues
      • Custom Field Generation Issues
    • Contributing
    • Support the Project
    • License
    • Star History
    • Disclaimer

Getting Started

Prerequisites

  • [Docker][docker-install] installed.
  • A running instance of [paperless-ngx][paperless-ngx] — tested against the 2.20.x release series and the 3.0.0 beta (v3.0.0-beta.rc1). paperless-gpt only uses the stable api/documents/, api/tags/, api/correspondents/, api/custom_fields/ and api/document_types/ endpoints, none of which have documented breaking changes in the v3 migration guide.
  • Access to an LLM provider:
    • OpenAI: An API key with models like gpt-4o or gpt-3.5-turbo.
    • Ollama: A running Ollama server with models like qwen3:8b.

Security

paperless-gpt has no built-in authentication. Its web UI and /api/* endpoints are open to anyone who can reach the port — by default it listens on all interfaces (LISTEN_INTERFACE defaults to :8080), so a plain -p 8080:8080 (as in the example below) exposes it to your whole LAN/VPN, not just localhost. Anyone who can reach it can rewrite documents in your connected paperless-ngx instance, trigger LLM/OCR jobs against your API keys, and change settings — with zero credentials required.

Do not expose it directly to the internet or an untrusted network. Put it behind a reverse proxy that adds authentication (e.g. Authelia, Authentik, a Basic Auth layer), restrict it to a VPN/Tailscale network, or otherwise limit who can reach the port.

Installation

Docker Compose

Here's an example docker-compose.yml to spin up paperless-gpt alongside paperless-ngx:

…

Pro Tip: Replace placeholders with real values and read the logs if something looks off.

Manual Setup

  1. Clone the Repository
    git clone https://github.com/icereed/paperless-gpt.git
    cd paperless-gpt
    
  2. Create a prompts Directory
    mkdir prompts
    
  3. Build the Docker Image
    docker build -t paperless-gpt .
    
  4. Run the Container
    docker run -d \
      -e PAPERLESS_BASE_URL='http://your_paperless_ngx_url' \
      -e PAPERLESS_API_TOKEN='your_paperless_api_token' \
      -e LLM_PROVIDER='openai' \
      -e LLM_MODEL='gpt-4o' \
      -e OPENAI_API_KEY='your_openai_api_key' \
      -e LLM_LANGUAGE='English' \
      -e VISION_LLM_PROVIDER='ollama' \
      -e VISION_LLM_MODEL='minicpm-v' \
      -e LOG_LEVEL='info' \
      -v $(pwd)/prompts:/app/prompts \
      -p 8080:8080 \
      paperless-gpt
    

OCR Providers

For detailed provider-specific documentation:

  • Mistral AI Integration

paperless-gpt supports four different OCR providers, each with unique strengths and capabilities:

1. LLM-based OCR (Default)

  • Key Features:
    • Uses vision-capable LLMs like gpt-4o or MiniCPM-V
    • High accuracy with complex layouts and difficult scans
    • Context-aware text recognition
    • Self-correcting capabilities for OCR errors
  • Best For:
    • Complex or unusual document layouts
    • Poor quality scans
    • Documents with mixed languages
  • Configuration:
    OCR_PROVIDER: "llm"
    VISION_LLM_PROVIDER: "openai" # or "ollama"
    VISION_LLM_MODEL: "gpt-4o" # or "minicpm-v"
    

2. Azure Document Intelligence

  • Key Features:
    • Enterprise-grade OCR solution
    • Prebuilt models for common document types
    • Layout preservation and table detection
    • Fast processing speeds
  • Best For:
    • Business documents and forms
    • High-volume processing
    • Documents requiring layout analysis
  • Configuration:
    OCR_PROVIDER: "azure"
    AZURE_DOCAI_ENDPOINT: "https://your-endpoint.cognitiveservices.azure.com/"
    AZURE_DOCAI_KEY: "your-key"
    AZURE_DOCAI_MODEL_ID: "prebuilt-read" # optional
    AZURE_DOCAI_TIMEOUT_SECONDS: "120" # optional
    AZURE_DOCAI_OUTPUT_CONTENT_FORMAT:
      "text" # optional, defaults to text, other valid option is 'markdown'
      # 'markdown' requires the 'prebuilt-layout' model
    

3. Google Document AI

  • Key Features:
    • Enterprise-grade OCR/HTR solution
    • Specialized document processors
    • Strong form field detection
    • Multi-language support
    • High accuracy on structured documents
    • Exclusive hOCR generation for creating searchable PDFs with text layers
    • Only provider that supports enhanced PDF generation features
  • Best For:
    • Forms and structured documents
    • Documents with tables
    • Multi-language documents
    • Handwritten text (HTR)
  • Configuration:
    OCR_PROVIDER: "google_docai"
    GOOGLE_PROJECT_ID: "your-project"
    GOOGLE_LOCATION: "us"
    GOOGLE_PROCESSOR_ID: "processor-id"
    CREATE_LOCAL_HOCR: "true" # Optional, for hOCR generation
    LOCAL_HOCR_PATH: "/app/hocr" # Optional, default path
    CREATE_LOCAL_PDF: "true" # Optional, for applying OCR to PDF
    LOCAL_PDF_PATH: "/app/pdf" # Optional, default path
    

4. Docling Server

  • Key Features:
    • Self-hosted OCR and document conversion service
    • Supports various input and output formats (including text)
    • Utilizes multiple OCR engines (EasyOCR, Tesseract, etc.)
    • Can be run locally or in a private network
  • Best For:
    • Users who prefer a self-hosted solution
    • Environments where data privacy is paramount
    • Processing a wide variety of document types
  • Configuration:
    OCR_PROVIDER: "docling"
    DOCLING_URL: "http://your-docling-server:port"
    DOCLING_IMAGE_EXPORT_MODE: "placeholder" # Optional, defaults to "embedd
    

GitHub Issues· 0 开放

在 GitHub 查看全部

暂无开放 Issues,或尚未同步最近议题。

核心特点

  • •LLM OCR: Use OpenAI or Ollama to extract text from images.
  • •Google Document AI: Leverage Google's powerful Document AI for OCR tasks.
  • •Azure Document Intelligence: Use Microsoft's enterprise OCR solution.
  • •Docling Server: Self-hosted OCR and document conversion service
  • •Append: This is the safest option: It only adds new fields that do not already exist on the document. It will never overwrite an existing field, even if it's empty.
  • •Update: Adds new fields and overwrites existing fields with new suggestions. Fields on the document that don't have a new suggestion are left untouched.
  • •Replace: Deletes all existing custom fields on the document and replaces them entirely with the suggested fields.
  • •Tagging: Decide how documents get tagged—manually, automatically, or via OCR-based flows.
  • •PDF Processing: Configure how OCR-enhanced PDFs are handled, with options to save locally or upload to paperless-ngx.
  • •Manual Review: Approve or tweak AI's suggestions.

> 标签

Goaichatgptllmmistral

暂无评论,来聊聊你的看法吧

> 工具信息

发布日期2026年8月1日
最后更新2026年9月17日
分类AI 编程
定价开源

> 相关工具

G
GitHub Copilot
GitHub 官方 AI 编程助手,覆盖补全、Chat 与 Agent 模式。
C
Cursor
AI 原生代码编辑器,对话改代码、多文件 Agent 与规则体系是其核心。
S
skills
Skills for Real Engineers. Straight from my .agents directory.