Self-hosted web scraping toolkit combining Patchright (undetected Playwright) with the Bypass Paywalls Clean (BPC) extension. Extracts article content as clean markdown + structured JSON. CLI, REST API, and Docker deployment.
# Install dependencies
bun install
# Install Patchright Chromium
bunx patchright install chromium
# Scrape a page
bun run scrape -- https://example.com --json
# Map links on a page
bun run map -- https://example.com --json
# Crawl a site
bun run crawl -- https://example.com --depth 2 --limit 10 --json
# Take a screenshot
bun run screenshot -- https://example.com --json
# Render a PDF
bun run pdf -- https://example.com --jsondocker compose up -d
curl http://localhost:3777/health
# If SHUVCRAWL_API_TOKEN is set:
# curl -H "Authorization: Bearer <token>" http://localhost:3777/healthsc starts a stopped local container before API commands (sc health, sc scrape, and the rest). It does not rebuild. sc up still builds. sc down and sc logs do not start it. Pass --no-autostart to skip this, or point --api-url / SHUVCRAWL_API_URL at a non-local host.
The API exits after 2 hours with no request other than GET /health, and with no running crawl. Docker restart: on-failure leaves that exit stopped. Set SHUVCRAWL_API_IDLETIMEOUTMS=0 to disable.
shuvcrawl <command> [options]
Commands:
scrape <url> Scrape a URL and extract content
map <url> Discover URLs on a page (links + sitemap)
crawl <url> Crawl a site starting from URL
screenshot <url> Capture a screenshot
pdf <url> Render a page as PDF
config Show current configuration
version Show version info
cache <sub> Cache management (status, list, clear)
serve Start the REST API server
update-bpc Inspect BPC extension status
| Option | Description |
|---|---|
--config <path> |
Config file override |
--output <dir> |
Output directory |
--format <fmt> |
Output format (markdown or json) |
--json |
Output as JSON |
--no-cache |
Bypass response cache |
--no-robots |
Skip robots.txt checking |
--proxy <url> |
Proxy URL (HTTP or SOCKS5) |
--user-agent <ua> |
Custom user agent |
--verbose |
Debug-level logging |
--quiet |
Error-only logging |
| Option | Description |
|---|---|
--wait <strategy> |
Wait strategy: load, networkidle, selector, sleep |
--wait-for <sel> |
CSS selector to wait for |
--wait-timeout <ms> |
Wait timeout in ms (default: 30000) |
--sleep <ms> |
Fixed delay after load |
--headers <json> |
Custom HTTP headers as JSON |
--mobile |
Mobile viewport (390×844) |
--raw-html |
Include raw HTML in output |
--only-main-content |
Extract main content only (default) |
--no-only-main-content |
Extract full body |
| Option | Description |
|---|---|
--depth <n> |
Maximum link depth (default: 3) |
--limit <n> |
Maximum pages (default: 50) |
--include <glob> |
Allowed URL patterns |
--exclude <glob> |
Denied URL patterns |
--delay <ms> |
Per-domain delay (default: 1000ms) |
--source <type> |
Discovery source: links, sitemap, or both |
--resume |
Resume from saved crawl state |
bun run cache status # Show cache stats
bun run cache list # List cache entries
bun run cache clear # Clear all
bun run cache clear --older-than 86400 # Clear entries older than 24hStart the server:
bun run serve -- --port 3777| Method | Path | Description |
|---|---|---|
GET |
/health |
Service health + config summary |
GET |
/config |
Current configuration (redacted) |
POST |
/scrape |
Scrape a URL |
POST |
/map |
Discover URLs |
POST |
/crawl |
Start async crawl (returns jobId) |
GET |
/crawl/:jobId |
Get crawl job status |
DELETE |
/crawl/:jobId |
Cancel a crawl job |
POST |
/screenshot |
Capture a screenshot |
POST |
/pdf |
Render a PDF |
Set SHUVCRAWL_API_TOKEN to require bearer token auth:
SHUVCRAWL_API_TOKEN=secret123 bun run serve
curl -H "Authorization: Bearer secret123" http://localhost:3777/healthcurl -X POST http://localhost:3777/scrape \
-H "Content-Type: application/json" \
-d '{"url": "https://example.com/article"}'{
"success": true,
"data": {
"url": "https://example.com/article",
"content": "# Article Title\n\nArticle body...",
"html": "<article>...</article>",
"metadata": {
"requestId": "req_abc123",
"title": "Article Title",
"bypassMethod": "bpc-extension",
"status": "success",
"elapsed": 2341
}
}
}Config file: ~/.shuvcrawl/config.json (or $SHUVCRAWL_CONFIG, or --config <path>)
All values can be overridden via environment variables: SHUVCRAWL_{SECTION}_{KEY}.
| Variable | Description | Default |
|---|---|---|
SHUVCRAWL_API_PORT |
API server port | 3777 |
SHUVCRAWL_API_TOKEN |
Bearer auth token | none |
SHUVCRAWL_API_IDLETIMEOUTMS |
Exit after this many idle ms (0 disables) |
7200000 |
SHUVCRAWL_CACHE_TTL |
Cache TTL in seconds | 3600 |
SHUVCRAWL_CRAWL_DELAY |
Per-domain delay in ms | 1000 |
SHUVCRAWL_BPC_ENABLED |
Enable BPC extension | true |
SHUVCRAWL_BROWSER_HEADLESS |
Headless mode | true |
SHUVCRAWL_TLS_REJECT_UNAUTHORIZED |
TLS certificate validation | true |
SHUVCRAWL_TELEMETRY_OTLPHTTPENDPOINT |
OTLP collector URL | none |
output/
├── {domain}/
│ ├── {slug}.md # Extracted markdown
│ ├── {slug}.json # Structured data
│ ├── _meta.jsonl # Metadata log (append-only)
│ └── _crawl-state.json # Crawl resume state
└── _artifacts/
└── {requestId}/
├── page.png # Screenshot
├── page.pdf # PDF render
├── raw.html # Raw page HTML
├── clean.html # Cleaned HTML
└── console.json # Browser console logs
| Code | Meaning |
|---|---|
| 0 | Success |
| 1 | General error |
| 2 | Configuration or auth error |
| 3 | Network error |
| 4 | Validation error |
| 5 | Extraction failed |
| 6 | Robots denied |
| 7 | Rate limited |
| 8 | Browser/extension init failed |
# Type checking
bun run typecheck
# Unit tests (126 tests)
bun test
# Integration tests (requires Chromium — 17 tests)
bun run test:integration
# Dockerized black-box API suite + report
bun run test:api
# or
./scripts/test-api-docker.sh --verify-otlp --open-report
# All tests
bun run test:all
# Start dev server
bun run serveCLI (commander) REST API (Hono)
│ │
└──────────┬──────────────┘
▼
Engine ──► BrowserPool, DomainRateLimiter, JobRegistry (SQLite)
│
┌──────────┼──────────┐
▼ ▼ ▼
Scrape Map/Crawl Capture
Pipeline (BFS+ (screenshot,
(fast-path sitemap) PDF, console)
→ browser
→ extract
→ convert)
│ │ │
└──────────┴──────────┘
▼
Storage Layer
(output/, cache/, artifacts/, SQLite job store)
Run the full containerized API suite:
bun run test:api
# or
./scripts/test-api-docker.sh --verify-otlp --open-reportWhat it does:
- starts the Dockerized API and fixture server with isolated run-local state
- runs black-box API tests against
http://localhost:3777 - stores artifacts in
test-results/<run-id>/ - generates
report.htmlfor review
Result package layout:
test-results/<run-id>/
├── summary.json
├── suite.json
├── manifest.json
├── report.html
├── env.json
├── docker/
├── api/
└── runtime/
Each run gets isolated runtime state under test-results/<run-id>/runtime/, including output, artifacts, cache, browser profiles, telemetry captures, and the SQLite job DB.
This repo ships a bundled skill at skills/shuvcrawl/.
Pi discovers it automatically from the conventional skills/ directory when you run Pi from the repo root. The package also includes optional pi.skills metadata in package.json for shipping clarity.
Use it when an agent needs to:
- scrape a page
- map links
- run a bounded crawl
- capture a screenshot or PDF
- choose between CLI, API, and Docker workflows safely
Private — not yet published.

