Autonomous media extraction and stealth crawl engine engineered for high-throughput asset discovery and resilient WAF bypass.
Getting Started • Key Features • WAF Pipeline • Architecture • Documentation • Threat Model (v0.30.0)
scrAPE is an autonomous media extraction & stealth crawl engine that runs locally on your machine. Built for domain crawling, high-throughput asset discovery, WAF bypass, and AI dataset curation, it handles complex single-page applications (SPAs), Cloudflare Turnstile protections, and high-concurrency downloads with real-time hardware telemetry.
-
Distributed Task Leasing & Autonomous Worker Daemons: Scalable multi-node cluster architecture using Redis Streams (
RedisStreamTaskBroker), atomic idempotency locks (SET NX EX 86400) preventing duplicate execution, dead-letter stream routing (scrape:dead_letter), background worker heartbeats, and autonomous CLI worker daemons (DistributedWorkerNode). -
Cloud Content-Addressable Storage (CAS) Synchronization: Asynchronous cloud block replication to Amazon S3, Cloudflare R2, and MinIO with bounded spooling (
maxsize=1000) and backpressure, 64-hex key validation against traversal, unconditional SSRF endpoint defense, and source-level credential redaction. Decoupled zero-cloud-dependency core. -
Multimodal Vision-Language (VLM) DOM Healing: Tier 4 DOM healing fallback using vision models (Gemini Flash, GPT-4o-mini, Ollama Vision) with strict prompt-injection immunity (
<untrusted_scraped_data>tags and 75-vector fuzzing rejection), structural default-deny allowlist for interactive elements, live DOM validation gate, and 7-day TTL cache expiration. - Pre-Warmed Anti-Bot Browser Pool: Eliminates 3–5s cold starts by pre-warming browser sessions (DrissionPage, Camoufox, Chromium) asynchronously in a background pool (<50ms lease time).
-
Hardware Device Manager & Multi-Provider LLM Gateway: Automatic CUDA, DirectML, MPS, and CPU routing with FP16/FP32 precision. Multi-provider LLM self-healing DOM parser (Ollama
qwen2.5-coder, Google Gemini 1.5 Flash, OpenAIgpt-4o-mini) with SQLite rule caching. -
Global Content-Addressable Storage (CAS): SHA-256 content deduplication with atomic NTFS/POSIX hardlinks (
os.link), consuming 0 additional disk bytes for identical media across runs and queries. -
Columnar Apache Parquet Dataset Exporter: Snappy-compressed Apache Parquet tables (
images.parquet,videos.parquet,run_summary.parquet) for high-performance ML analytics with JSONL fallback. - Dynamic Tactical WebUI: Decoupled FastAPI + HTMX dashboard featuring an interactive HTML5 Canvas crawl network tree, live Node Health telemetry (CPU, RAM, Disk, concurrency throttle indicator), process controls, and dual speed limiters.
- 8-Tier WAF Escalation & Stealth: Defeats Cloudflare Turnstile, reCAPTCHA v2/v3, and auth walls using an 8-tier fallback pipeline (Local Cookies → Crawl4AI → Crawlee Cheerio → DrissionPage → Crawlee Puppeteer → Helium → undetected-chromedriver → Camoufox → FlareSolverr).
-
Universal CAPTCHA Strategy: Commercial API auto-solving (
CapSolver,2Captcha,AntiCaptcha) with automatic fallback toFreeAudioCaptchaProvider(local Whisper speech-to-text solver). -
Asynchronous Inline ML Pipeline: Decoupled background worker (
AsyncMLPipelineWorker) running non-blocking aesthetic scoring/culling, smart face/body cropping, and WD14 dataset tagging. -
Pluggable Multi-Tier Storage Sinks: Atomic
LocalStorageSink(traversal-safe) and direct multipartS3StorageSink(Amazon S3 / MinIO) with automatic offline spillover. -
3-Tier Hierarchical Deduplication:
$O(1)$ SHA-256 Bloom filter$\to$ 64-bit DCT pHash indexed in a BK-Tree metric tree$\to$ vector cosine similarity index. -
Hybrid Worker Concurrency Pool: CPU/GPU process isolation via
ProcessPoolExecutor(spawncontext) andpsutilprocess-tree tracking guaranteeing zero zombie child processes. - HardwareLoadGovernor: Dynamic system RAM & CPU monitoring that overrides the pipeline's concurrency factor (scales 1x to 3x) and forces garbage collection under load.
-
Dual Speed Limiters: Token-bucket rate-limiting on outgoing page requests (
RPS) and network bandwidth throttling on media asset downloads (KBPS). - Resumable HTTP Range Downloads: Persistent SQLite queue paired with HTTP 206 Partial Content byte resumption and Pillow image sanitization.
- Pluggable Notification Architecture: Multi-channel webhook notifier supporting Telegram Bot alerts, Discord rich embeds, Slack, and generic webhooks.
- Language: Python 3.10+ (CLI & Core Engine), Node.js 18+ (Crawlee Bridge)
- Framework: FastAPI (Backend API)
- Frontend: HTMX with HTML5 Canvas (No heavy SPA frameworks)
- Database: Persistent SQLite (WAL mode)
- Extraction Tools: Crawl4AI, DrissionPage, undetected-chromedriver, Helium, Camoufox, yt-dlp, BeautifulSoup4
- Node.js Bridge: Express.js, Crawlee, Puppeteer, puppeteer-extra-plugin-stealth
- Deployment/Solving: Docker (FlareSolverr native binding on port 8191)
- Python 3.10+ (Python 3.12/3.13 fully supported)
- Node.js 18+ and
npm(Required for thecrawlee_bridgestealth tiers) - Git
- Docker (Optional, but highly recommended for FlareSolverr background auto-start)
git clone https://github.com/your-username/scraper.git
cd scraperCreate and activate a virtual environment (recommended):
# Windows
python -m venv .venv
.venv\Scripts\activate
# macOS / Linux
python3 -m venv .venv
source .venv/bin/activateInstall requirements:
pip install -r requirements.txtNavigate to the bridge directory and install the required Crawlee/Puppeteer stealth packages:
cd crawlee_bridge
npm install
cd ..Note: If you run the unified master launcher (run.bat / run.sh), it will attempt to detect and install missing Node.js dependencies automatically.
Copy the example environment file:
cp .env.example .envConfigure the essential variables in .env:
| Variable | Description | Example |
|---|---|---|
WEBUI_HOST |
Host address for the dashboard | 0.0.0.0 |
WEBUI_PORT |
Port for the dashboard | 10001 |
TELEGRAM_BOT_TOKEN |
(Optional) Alerts for completion/WAF issues | 123456:ABC-DEF1234... |
CAPSOLVER_API_KEY |
(Optional) Third-party Captcha solving | CAP-... |
DOWNLOAD_RATE_LIMIT_RPS |
Hard cap on download requests per second | 5.0 |
Using the Unified Master Launcher (Recommended):
# Windows
.\run.bat
# macOS / Linux
./run.shThis opens an interactive menu to start the WebUI, run the CLI Wizard, or execute the continuous Watchdog monitoring agent.
Running the Dashboard Manually:
# Starts the FastAPI/HTMX Command Center on http://localhost:10001
.\run_frontend.bat # Or python frontend/app.pyRunning a Direct CLI Scrape:
python src/cli/main.py --keyword "architecture" --seed-file seeds/architecture.txt \
--download-media --workers 8 --dl-workers 12 \
--enable-cas --export-parquet --enable-self-healingscrape-dashboard/
├── crawlee_bridge/ # Node.js Express Server for Puppeteer/Cheerio stealth
├── frontend/ # FastAPI + HTMX WebUI
│ ├── app.py # Core FastAPI application
│ ├── routers/ # Decoupled API routes (dashboard, dataset, seeds, watchdog)
│ └── templates/ # HTMX Brutalist templates
├── src/ # Python Source Core
│ ├── cli/ # CLI entrypoints, wizards, watchdog
│ ├── core/ # BFS crawl loop, filters, worker pool, self-healing DOM parser
│ ├── network/ # 8-tier WAF stealth pipeline, Domain Tier Memory, proxy rotation
│ ├── storage/ # SQLite state cache, CAS store, Parquet dataset exporter
│ └── ml/ # Hardware detection, LLM healing gateway, taggers
├── tests/ # Domain-structured test suite (533+ tests)
├── data/ # Configuration Registries (domain_config, normalisation)
├── docs/ # Technical documentation
└── seeds/ # Per-subject manifest target files
- URL is discovered and enqueued by the BFS Crawler (
src/core/orchestrator.py, coordinated withsrc/core/domain_rules.pyandsrc/core/media_processor.py). HttpClient(src/network/http_client.py) attempts to fetch the URL using standard TLS parameters and local harvested cookies.- If an auth wall (302 redirect), HTTP 403, or HTTP 429 is encountered, the request hits the Stealth Pipeline.
- The pipeline iterates through configured engines (Crawl4AI → Crawlee → DrissionPage → Camoufox → FlareSolverr) based on the domain's historical success memory.
- If a CAPTCHA is detected (Turnstile, reCAPTCHA), the
CaptchaStrategydelegates to CapSolver/2Captcha. - Once bypassed, the HTML is passed back to the
ScrapingEnginefor CSS selection and asset extraction.
| Variable | Description | Default |
|---|---|---|
WEBUI_HOST |
Bound host IP for FastAPI | 0.0.0.0 |
WEBUI_PORT |
Bound port for FastAPI | 10001 |
MAX_CONCURRENT_PER_HOST |
Max concurrent connections to a single domain | 4 |
PROXY_MAX_BANDWIDTH_MB |
Circuit breaker for proxy bandwidth | 500.0 |
DOWNLOAD_RATE_LIMIT_RPS |
Overall download requests per second throttle | 5.0 |
| Variable | Description |
|---|---|
TELEGRAM_BOT_TOKEN |
Token for Telegram Bot integration |
TELEGRAM_CHAT_ID |
Your personal or group Chat ID |
HARVEST_NOTIFY_THRESHOLD |
Send an alert every N items downloaded |
DISCORD_WEBHOOK_URL |
Discord webhook for rich embedded alerts |
CAPSOLVER_API_KEY |
Key for automated CAPTCHA bypassing |
| Command | Description |
|---|---|
.\run.bat |
Unified interactive master menu (Windows) |
./run.sh |
Unified interactive master menu (POSIX) |
scrape |
Global CLI execution (if installed via pip install -e .) |
python src/cli/main.py --help |
View all scraping arguments and thresholds |
python src/cli/cli_wizard.py |
Step-by-step interactive configuration wizard |
python src/cli/monitor_agent.py |
Launch the continuous background Watchdog daemon |
scrAPE uses pytest for unit and integration testing. Over 546 automated tests validate network fallback simulation, database transaction integrity, UI state verification, Content-Addressable Storage (CAS), Parquet exports, Domain Tier Memory, and anti-SSRF protections.
# Run the complete test suite (546 tests)
pytest tests/ -v
# Run specific functional areas
pytest tests/test_security_ssrf_and_tier_memory.py -v
pytest tests/storage/test_cas_and_parquet_exporter.py -v
pytest tests/core/test_self_healing_parser.py -vAll test scripts are maintained exclusively inside the tests/ directory. Temporary diagnostic or scratch scripts should be kept in scratch/.
scrAPE is primarily designed to run locally or on a dedicated VPS/bare-metal server, due to its heavy reliance on headful browsers (for Turnstile bypass) and extensive disk I/O.
The scraper seamlessly integrates with FlareSolverr to handle Cloudflare IUAM.
- The scraper attempts to bind to
http://127.0.0.1:8191/v1. - If unreachable, the internal
FlareSolverrMonitorexecutes a background Docker auto-start:docker start flaresolverr
- Ensure you have pulled the FlareSolverr image:
docker pull ghcr.io/flaresolverr/flaresolverr:latest
(See the docker-compose.yml file for multi-container orchestration of the entire stack).
If you deploy to a memory-constrained environment (e.g., a VPS with <4GB RAM), the HardwareLoadGovernor will automatically:
- Detect system load and throttle the
ThreadPoolExecutorworker count dynamically (down to 1x scaling). - Force Python garbage collection (
gc.collect()) when RAM utilization spikes. - Restrict heavy stealth tiers (like Puppeteer and Camoufox) if resources are exhausted.
Error: Executable doesn't exist at C:\Users\... or BrowserType.launch: Failed to launch
Solution: You need to install the Playwright browsers.
playwright install chromiumError: Failed to connect to Crawlee bridge at http://localhost:3000
Solution: The Node.js dependencies are missing or the bridge crashed.
- Navigate to
crawlee_bridge/. - Delete
node_modulesand runnpm install. - Check
crawlee_bridge.login the root directory for specific Express.js crash traces.
Error: Domains are instantly cutting off or failing downloads. Solution:
- Check
run_summary.jsonfor rejection reasons. - If a CDN is aggressively blocking you, lower
--dl-workers(e.g.,4), or define# Rate-limit: 0.5 req/sin your seed manifest for that specific domain. - Ensure your stealth tiers are active and you have a valid Captcha API key configured.
Distributed under the MIT License. See LICENSE for more information.
