百科.dev
全部条目AI 编程趋势榜开源项目技术资讯提交条目
登录
< 返回工具列表
trawl

trawl

> 后端框架
免费

自主托管的爬虫引擎 — 绕过任何 JS 挑战和验证码: Cloudflare、转盘、reCAPTCHA

793 stars0 点赞1 次浏览
访问官网GitHub

工具介绍

自主托管的爬虫引擎 — 绕过任何 JS 挑战和验证码: Cloudflare、转盘、reCAPTCHA

TRAWL

Welcome to TRAWL!

Self-hosted web scraping engine with best-effort JS challenge and CAPTCHA solving.
Dedicated flows for Cloudflare, Akamai Bot Manager, and Imperva/Incapsula (best effort), plus Turnstile, reCAPTCHA, hCaptcha, GeeTest, ALTCHA, and Friendly Captcha.
Much faster and more reliable FlareSolverr & Byparr alternative and drop-in replacement for your *arr stack.

Features

  • 2-6x faster - compared to FlareSolverr or Byparr it returns much faster with higher success rate
  • 4-tier execution - plain HTTP fetch → cached browser session → fresh challenge solve → residential proxy
  • Challenge-aware HTTP/HTTPS proxy - direct forwarding for normal traffic, automatic tier escalation for detected walls, plus WebSockets, binary bodies, and Range/206 support
  • Multi-WAF handling - dedicated Cloudflare, Akamai Bot Manager, and Imperva/Incapsula detection and browser flows
  • Native captcha solving - CF Turnstile/Interstitial, reCAPTCHA v2 (free STT), hCaptcha, GeeTest v4 Slide, ALTCHA, and Friendly Captcha v1/v2
  • Camoufox Firefox - fingerprint-patched at the C++/Juggler level to reduce automation signals
  • Session cache - solved cookies and browser identity stored in Redis; accepted sessions can avoid a fresh solve
  • FlareSolverr compatible - works with Prowlarr, Jackett, Sonarr, and the full *arr ecosystem out of the box
  • No paid solver API required - reCAPTCHA audio can use Google's free STT endpoint or an optional local Whisper service

Sponsors

View/Collapse All

    Bright Data - The most powerful platform for Web Unlocker, SERP API and web scraping tools.

    Why Bright Data?

    • Web Unlocker - bypass any anti-bot protection

    • SERP API - real-time Google, Bing & more results

    • Scraping Browser & dedicated scrapers

    • Massive residential proxy network

    • Built for scale and reliability

    Get started for free with Bright Data!
  




  
    
      
    
  
  
    NodeMaven - The most efficient proxy provider for Web Scrapping and Automation with the Highest Quality IP on the market.

    Why NodeMaven?

    • ZIP targeting

    • 99.9% uptime

    • IP filtering: all proxies have fraud score 

Quick start

# Clone and configure
git clone https://github.com/germondai/trawl
cd trawl
cp .env.example .env

# Start scraper + Redis
docker compose up -d

# Verify
curl http://localhost:8191/health

First boot takes 15–30s while the browser pool warms up. Subsequent starts are fast.

NAS app catalogs

Prefer a one-click installation? TRAWL is available from the community app catalogs for both TrueNAS and Unraid:

  • TrueNAS Community Apps — open Apps → Discover Apps and search for TRAWL.
  • Unraid Community Apps — open the Apps tab and search for Trawl.

Thanks to the TrueNAS and Unraid community contributors who packaged and published these integrations.

API

FlareSolverr-compatible (/v1)

curl -X POST http://localhost:8191/v1 \
  -H 'Content-Type: application/json' \
  -d '{"cmd":"request.get","url":"https://nowsecure.nl","maxTimeout":60000}'

Native API (/scrape)

Returns richer metadata: tier, timings, sessionCached, full cookie list.

curl -X POST http://localhost:8191/scrape \
  -H 'Content-Type: application/json' \
  -d '{"url":"https://nowsecure.nl","maxTimeout":60000}'

MCP tools (/mcp)

Set MCP_ENABLED=true to expose TRAWL's client-independent Streamable HTTP tools for readable content, HTML, screenshots and browser diagnostics to any MCP-compatible AI application or agent. They load known public URLs; TRAWL does not provide web search or ranking. See the MCP integration guide.

Connect Prowlarr / Jackett

Set the FlareSolverr URL to:

http://localhost:8191        # running on the same host
http://trawl:8191            # running via Docker Compose on the same network

Challenge-bypassing HTTP/HTTPS proxy

Some sites bind their Cloudflare clearance to the solving browser's full connection fingerprint. The /v1 flow can't help there: Prowlarr keeps only the cookie + user-agent and re-fetches the page with its own HTTP client, which Cloudflare re-challenges — the cookie isn't portable. For those indexers, enable TRAWL's forward proxy and add it to Prowlarr as an HTTP proxy:

MITM_ENABLED=true
MITM_PORT=8192
MITM_CA_DIR=/data/proxy-ca   # persist the CA (mount a volume)
MITM_MAX_TIER=4              # cap escalation (e.g. 3 to stay off residential)
MITM_ALWAYS_SCRAPE=false     # opt in to bypass the proxy's direct Tier 0 probe

By default the listener binds 0.0.0.0 so clients on a Docker bridge network can reach it; set MITM_HOST=127.0.0.1 to restrict it to loopback on a bare-metal host.

  1. Install the proxy's CA into the client's trust store so it accepts the per-host certs: curl http://:8191/proxy-ca.crt → add to the Prowlarr container's CA store (e.g. a linuxserver /custom-cont-init.d script that copies it to /usr/local/share/ca-certificates/ and runs update-ca-certificates).
  2. Prowlarr → Settings → Indexer Proxies → HTTP, host ``, port 8192. Give it a tag if only selected indexers should use it.

Ordinary requests use a direct HTTP/TLS path. Small HTML, JSON, and text responses are buffered for challenge detection; detected challenges escalate through the same tier pipeline as POST /scrape. Videos and large binary responses stream directly. Range requests are forwarded end to end and can escalate when their response is a detected challenge; WebSocket upgrades use a direct relay without browser escalation.

See the complete proxy documentation for routing details, supported traffic, limitations, CA installation, and client examples.

⚠️ A MITM proxy can impersonate any host to a client that trusts its CA. Only expose it on a private interface (localhost / a private Docker network), never publicly.

Installing the proxy CA certificate

The proxy self-generates a root CA on first run. Its certificate and private key are persisted under MITM_CA_DIR (default /data/proxy-ca). Per-host certificates are minted and cached in memory while TRAWL runs; they do not need separate installation because they are signed by the persistent root. Every client that uses the proxy must trust that root. Without it, HTTPS fails with ERR_CERT_AUTHORITY_INVALID (browsers) or PKIX path building failed (Java).

Download the CA once per client:

curl http://:8191/proxy-ca.crt -o trawl-ca.crt
# or in a Docker setup where the API isn't reachable from outside:
docker cp trawl:/data/proxy-ca/ca.crt ./trawl-ca.crt

macOS (system keychain — affects most apps including Safari, curl, wget)

sudo security add-trusted-cert -d -r trustRoot \
  -k /Library/Keychains/System.keychain ./trawl-ca.crt
# Verify
security find-certificate -c "TRAWL MITM Proxy CA"
# Remove later
sudo security delete-certificate -c "TRAWL MITM Proxy CA" \
  /Library/Keychains/System.keychain

Linux (Debian/Ubuntu — system-wide for curl, wget, apt, etc.)

sudo cp trawl-ca.crt /usr/local/share/ca-certificates/trawl-ca.crt
sudo update-ca-certificates
# Verify
awk '/BEGIN/{c++} c==2' /etc/ssl/certs/ca-certificates.crt | grep -c "TRAWL MITM"

Linux (RHEL/Fedora/Amazon)

sudo cp trawl-ca.crt /etc/pki/ca-trust/source/anchors/trawl-ca.crt
sudo update-ca-trust

Firefox and NSS trust stores

Firefox installations that do not use operating-system roots need a per-profile NSS import:

# Firefox 115+ uses a file-backed NSS DB; older versions use the legacy libnssdb format.
# The certutil command is the same either way.
certutil -A -n "TRAWL MITM" -t "CT,C,C" -i trawl-ca.crt \
  -d sql:$HOME/.mozilla/firefox/
# Or via Firefox UI: Settings → Privacy & Security → Certificates → View Certificates →
# Authorities → Import… → check "Trust this CA to identify websites".
# Profile dir location: about:profiles in Firefox.

Chrome / Chromium (Linux: separate from system trust)

Chrome uses the system trust store on macOS and Windows but has its own on Linux:

# Option A: launch Chrome with --user-data-dir + NSS DB update (same as Firefox).
# Option B: use Chrome's --ignore-certificate-errors-spki-list= (per-session, less safe).
# Option C: add the cert to the system store (above) — Chrome picks it up automatically on
# most Linux distros via the nss-tool lookup.

Java (including JDownloader)

# Find the JRE cacerts file for your client.
#   JDownloader:    /jre/lib/security/cacerts
keytool -importcert -alias trawl -file trawl-ca.crt \
  -keystore "" -storepass changeit
# If `keytool` reports "Certificate already exists in keystore", use -delete first:
#   keytool -delete -alias trawl -keystore "" -storepass changeit

Prowlarr, Sonarr, and Radarr are .NET applications, not Java applications. For their Docker-based installations, add the CA to the container's Linux system trust store. A common LinuxServer pattern is a /custom-cont-init.d script:

# In the client's Compose service:
volumes:
  - ./trawl-ca.crt:/config/trawl-ca.crt:ro
  - ./install-trawl-ca.sh:/custom-cont-init.d/50-install-trawl-ca:ro
#!/usr/bin/with-contenv bash
cp /config/trawl-ca.crt /usr/local/share/ca-certificates/trawl-ca.crt
update-ca-certificates

LinuxServer runs scripts in /custom-cont-init.d/ when the container starts. Java clients such as JDownloader require the separate keytool import described above.

JDownloader 2 (Windows / macOS / Linux — manual install)

JDownloader bundles its own JRE; the CA must be imported into it.

  1. Find the JRE: Settings → Advanced → Java Path (in JDownloader) or look in the install dir:
    • Windows: C:\Program Files\JDownloader 2\jre\lib\security\cacerts
    • macOS: /Applications/JDownloader 2.app/Contents/app/jre/lib/security/cacerts
    • Linux: /jre/lib/security/cacerts
  2. Run the keytool -importcert command above against that file.
  3. Restart JDownloader.

Windows (system trust store)

# Run PowerShell as Administrator.
Import-Certificate -FilePath .\trawl-ca.crt `
  -CertStoreLocation Cert:\LocalMachine\Root
# Remove later
Get-ChildItem Cert:\LocalMachine\Root | Where-Object { $_.Subject -like "*TRAWL MITM*" } | Remove-Item

Removing the CA (cleanup)

Every installation method has a symmetric removal path. Search your trust store for TRAWL MITM Proxy CA (the CA's CN) and delete that entry. The CA certificate and key also live at /ca.crt and ca.key on the TRAWL host. Deleting either causes TRAWL to generate a new root on its next start, so existing clients must install the new certificate.

Tiers

…

bash git tag -a v1.0.1 -m "..." git push origin v1.0.1


## Configuration

TRAWL supports HTTP proxies, authenticated HTTP proxies, and SOCKS5 proxies. The standard Compose
files read proxy settings from the local `.env` file:

```ini
# Optional Tier 3 datacenter proxy
PROXY_URL=http://user:[email protected]:8080

# Optional Tier 4 residential proxy
RESIDENTIAL_PROXY_URL=socks5://user:[email protected]:1080
docker compose up -d

Leave either value empty to disable that proxy tier. Multiple endpoints can be separated with commas; larger pools can use the corresponding *_LIST_FILE variable. See Configuration → Proxies for pool and mounted-file examples.

| Variable | Default | Description

Issues· 3 开放

查看全部 Issues在 GitHub 打开

暂无开放 Issues,或尚未同步最近议题。

> 标签

bunbyparrflaresolverr

暂无评论,来聊聊你的看法吧

> 工具信息

发布日期2026年9月9日
最后更新2026年9月17日
分类后端框架
定价免费

> 相关工具

N
Node.js
基于 V8 的 JavaScript 运行时
D
Django
Python 高级 Web 框架
S
Spring Boot
Java 生态主流微服务框架