Open source · Self-hosted · Built for agents
ScrapeQL loads a page in a real browser, digests it, and extracts only the relevant data your agent needs, as strict JSON. Plug it into your agent over MCP or plain HTTP. Run it on your own box with one docker compose up.
{ books[] { title price(with currency) rating(1 to 5) in_stock } }
{ "books": [ { "title": "It's Only the Himalayas", "price": "£45.17", "rating": "2", "in_stock": "yes" }, { "title": "Full Moon over Noah's…", … } ] }
One pull, four views
One headless-Chrome pull captures markdown, HTML, the ARIA tree and a full-page screenshot at the same moment. What you inspect in the workbench is what the model receives.
markdownNav, footers, cookie banners and link dumps removed. Trimmed to a token budget, 4000 by default.
htmlThe rendered DOM after JavaScript ran, for when structure matters.
aria_treeRoles and names from the accessibility tree. A compact way to describe an app screen.
screenshotA full-page capture of what the browser saw, returned with every pull.
From playground to production
Click an element or type a CSS selector to scope the page. A live readout shows the characters and tokens you are about to send, before you spend them.
Name the fields, mark lists with [], add hints in parentheses. The query compiles to a strict JSON schema the model must follow. Every value comes back as a string.
Ask a free-form question about a page, or leave it blank for a summary. Answers stay inside a token budget.
scrape, extract, ask, search and screenshot, served from /mcp. Point Claude, Cursor or any MCP client at your host and the agent fetches its own pages.
Spider through pagination, hop from list to detail pages, and keep every run's rows in Postgres. Trigger from the UI or POST /api/jobs/{id}/run.
Web search runs through a self-hosted SearXNG that ships in the stack. Your agent searches, picks a result, and extracts from it in one session.
Every experiment has a copy-paste POST /api/extract or /api/ask snippet. The OpenAPI reference and llms.txt are generated from the same schemas, so they never drift.
Some sites treat datacenter IPs differently. Send pulls out through DataImpulse residential IPs, the default provider, with per-request country targeting and sticky IPs. Or route them through your own home connection with one compose profile. Both work today.
Your OpenAI and Browserless keys stay server-side. Agents and scripts hold only a ScrapeQL token. Private and internal addresses are refused by default.
Built for the agent era
Output is context-safe by default. Cleaned markdown, capped at a token budget. Big pages come back as a summary and an outline, with a hint to use extract or ask for just the fields. Add the MCP server to your client and the agent calls these tools itself.
scrapeextractasksearchscreenshot// claude, cursor, any MCP client { "mcpServers": { "scrapeql": { "url": "https://your-host/mcp" } } }
Set SCRAPEQL_API_TOKEN or put the host behind Cloudflare Access first. /mcp stays off until one of them is configured. Private and internal addresses are refused by default.
curl -s https://your-host/api/extract \ -H "Authorization: Bearer $SCRAPEQL_API_TOKEN" \ -H "Content-Type: application/json" \ -d '{ "url": "https://example.com/pricing", "formats": ["json"], "extract": { "query": "{ plans[] { name price } }" } }'
Upstream keys (OpenAI, Browserless) never leave the server. Callers hold only your ScrapeQL token.
Self-host
App, Browserless, SearXNG and Postgres come up together. The only key you need is an OpenAI API key. GPT-6 Luna is the default model.
git clone https://github.com/columenlabs/scrapeql
cd scrapeql
echo "OPENAI_API_KEY=sk-..." > .env
docker compose up -d
# then open http://localhost:3000
appWorkbench, HTTP API and the MCP server at /mcpbrowserlessHeadless Chrome that renders each page and takes the screenshotsearxngSelf-hosted metasearch behind the search toolpostgresJobs, run history and extracted rowsresidential-egressRoutes pulls through your own home connectionTwo ways to run it
Free · Apache-2.0
Waitlist