Elasticsearch 文件系统爬虫 (FS Crawler)
Welcome to FSCrawler for Elasticsearch
This crawler helps to index binary documents such as PDF, Open Office, MS Office.
Main features:
Current versions are:
Elasticsearch FSCrawler Released Docs 6.x, 7.x 2.9 2022-03-08 2.9 7.x, 8.x, 9.x 3.0 2026-08-26 3.0 8.x, 9.x 3.1-SNAPSHOT 3.1-SNAPSHOTRun Elasticsearch with start-local:
# Start Elasticsearch and Kibana
curl -fsSL https://elastic.co/start-local | sh
# Get the generated API key (you will need it for FSCrawler)
source elastic-start-local/.env
Run FSCrawler with Docker:
docker pull dadoonet/fscrawler
docker run -it --rm \
--add-host=host.docker.internal:host-gateway \
-v ~/.fscrawler:/root/.fscrawler \
-v $(pwd)/resumes:/tmp/es:ro \
-e FSCRAWLER_ELASTICSEARCH_URLS=http://host.docker.internal:9200 \
-e FSCRAWLER_ELASTICSEARCH_API_KEY="${ES_LOCAL_API_KEY}" \
-e FS_JAVA_OPTS="-DLOG_LEVEL=debug" \
dadoonet/fscrawler
Then open Kibana and watch for your documents coming to the fscrawler alias:
FROM fscrawler
| STATS numDocs = COUNT(*)
Or search for some text:
FROM fscrawler
| WHERE content : "David"
Or count by file.content_type:
FROM fscrawler
| STATS numDocs = COUNT(*) BY file.content_type
Note:
~/resumes contains the documents you want to index~/.fscrawler/fscrawler/_settings.yamlRead the documentation for more details and specifically the tutorial page.
Need help writing your job settings? Copy a ready-made prompt from the LLM assistant guide into ChatGPT, Claude, or your favorite AI assistant.
FSCrawler also publishes an llms.txt
index (mirrored in the repo) and clean Markdown versions of every docs page
(*.html.md) so agents can skip HTML chrome — see the llms.txt v2 proposal.
To test this SNAPSHOT with Docker, use the snapshot tag. Untagged dadoonet/fscrawler
(or :latest) is the last stable release:
docker pull dadoonet/fscrawler:snapshot
For Docker Compose, set FSCRAWLER_VERSION=snapshot in .env. Use snapshot-noocr if you
do not need OCR. The ZIP is published as a GitHub pre-release on every push to main.
Read more about the Apache2 License.
Thanks to JetBrains for the IntelliJ IDEA License! The best IDE out there!
Thanks to SonarCloud for the free analysis! You guys rock!
File not deleted from elasticsearch
remove_deleted only reconciles the first 10,000 files in a directory + Orphan files
support for child doc types
Add/document Support for OpenSearch
Use External API for OCR ie amazon textract or google vision
Add an ElasticsearchPasswordProvider
Generate a thumbnail version of the document
Submit job settings via REST api
Do not extract ALL raw metadata
fscrawler statistics in a monitoring stack