Open source · Self-hosted · Built for agents

Ask the web for exactly the fields you need.

ScrapeQL loads a page in a real browser, digests it, and extracts only the relevant data your agent needs, as strict JSON. Plug it into your agent over MCP or plain HTTP. Run it on your own box with one docker compose up.

Open source, Apache-2.0LLM and browser keys stay on your serverMCP server at /mcp
books.toscrape.com/catalogue/category/books/travel_2
Query
{
  books[] {
    title
    price(with currency)
    rating(1 to 5)
    in_stock
  }
}
Strict JSON
{ "books": [
  { "title": "It's Only the Himalayas",
    "price": "£45.17",
    "rating": "2",
    "in_stock": "yes" },
  { "title": "Full Moon over Noah's…",
    … }
] }
scope ol.rowinput 3,904 chars · ~976 tokens

One pull, four views

Fetch once. See the page the way a model will.

One headless-Chrome pull captures markdown, HTML, the ARIA tree and a full-page screenshot at the same moment. What you inspect in the workbench is what the model receives.

Any URLhttps://…fetchBrowserlessreal Chrome, one pullmarkdownhtmlaria_treescreenshotScrapeQL querycompiles to a schemafill{ }strict JSONscope with a selector · chars and tokens counted before the model sees them
One fetch, four views. The query only ever sees the part of the page you scoped.
markdown

Clean text

Nav, footers, cookie banners and link dumps removed. Trimmed to a token budget, 4000 by default.

html

Raw DOM

The rendered DOM after JavaScript ran, for when structure matters.

aria_tree

Accessibility tree

Roles and names from the accessibility tree. A compact way to describe an app screen.

screenshot

Full page

A full-page capture of what the browser saw, returned with every pull.

From playground to production

Prove an extraction in the workbench. Ship it as an API call or an MCP tool.

Scope

Point at the part that matters

Click an element or type a CSS selector to scope the page. A live readout shows the characters and tokens you are about to send, before you spend them.

Extract

Typed queries, strict JSON

Name the fields, mark lists with [], add hints in parentheses. The query compiles to a strict JSON schema the model must follow. Every value comes back as a string.

Ask

Questions about a page

Ask a free-form question about a page, or leave it blank for a summary. Answers stay inside a token budget.

MCP

Five tools your agent can call

scrape, extract, ask, search and screenshot, served from /mcp. Point Claude, Cursor or any MCP client at your host and the agent fetches its own pages.

Jobs

Multi-page crawls

Spider through pagination, hop from list to detail pages, and keep every run's rows in Postgres. Trigger from the UI or POST /api/jobs/{id}/run.

Search

Find the URL first

Web search runs through a self-hosted SearXNG that ships in the stack. Your agent searches, picks a result, and extracts from it in one session.

Productionize

Copy the curl, done

Every experiment has a copy-paste POST /api/extract or /api/ask snippet. The OpenAPI reference and llms.txt are generated from the same schemas, so they never drift.

Egress

Residential exit, two ways

Some sites treat datacenter IPs differently. Send pulls out through DataImpulse residential IPs, the default provider, with per-request country targeting and sticky IPs. Or route them through your own home connection with one compose profile. Both work today.

Keys

One token for callers

Your OpenAI and Browserless keys stay server-side. Agents and scripts hold only a ScrapeQL token. Private and internal addresses are refused by default.

Built for the agent era

Your agent gets the digest. Over MCP or HTTP.

Output is context-safe by default. Cleaned markdown, capped at a token budget. Big pages come back as a summary and an outline, with a hint to use extract or ask for just the fields. Add the MCP server to your client and the agent calls these tools itself.

Your agentClaude, Cursor, MCPcall + tokenScrapeQL/mcp · /api/extractone pullBrowserlessheadless ChromeThe webdigest ≤ 4koptionalResidential exitDataImpulse or home lineprivate IPs refused by default
Your agent talks to ScrapeQL. ScrapeQL talks to the browser. The agent gets back a digest, not a page.
MCP server · /mcp
scrapeextractasksearchscreenshot
// claude, cursor, any MCP client
{
  "mcpServers": {
    "scrapeql": {
      "url": "https://your-host/mcp"
    }
  }
}

Set SCRAPEQL_API_TOKEN or put the host behind Cloudflare Access first. /mcp stays off until one of them is configured. Private and internal addresses are refused by default.

HTTP · POST /api/extract
curl -s https://your-host/api/extract \
  -H "Authorization: Bearer $SCRAPEQL_API_TOKEN" \
  -H "Content-Type: application/json" \
  -d '{
    "url": "https://example.com/pricing",
    "formats": ["json"],
    "extract": {
      "query": "{ plans[] { name price } }"
    }
  }'

Upstream keys (OpenAI, Browserless) never leave the server. Callers hold only your ScrapeQL token.

Self-host

One compose file. Every service included.

App, Browserless, SearXNG and Postgres come up together. The only key you need is an OpenAI API key. GPT-6 Luna is the default model.

~/scrapeql
git clone https://github.com/columenlabs/scrapeql
cd scrapeql
echo "OPENAI_API_KEY=sk-..." > .env
docker compose up -d
# then open http://localhost:3000
your server · docker composeappUI · API · /mcpChromebrowserlessSearchsearxngPostgresjobs, runsAgentsscripts, CItokenOpenAILLM APIkeyoptionalresidential-egress
Everything runs inside your compose stack. Callers only ever hold a ScrapeQL token.
appWorkbench, HTTP API and the MCP server at /mcp
browserlessHeadless Chrome that renders each page and takes the screenshot
searxngSelf-hosted metasearch behind the search tool
postgresJobs, run history and extracted rows
Optional profileresidential-egressRoutes pulls through your own home connection

Two ways to run it

Run it yourself today. Let us run it for you later.

Available at launch

Self-hosted

Free · Apache-2.0

  • Every feature, including the MCP server
  • Runs wherever docker compose runs
  • Bring your own OpenAI key and pay OpenAI directly
  • Fork it, embed it, ship it inside your own product
Planned

Hosted

Waitlist

  • The same open-source build, run by Columen Labs
  • No servers or keys to manage
  • Connect your agent over MCP
  • Not open yet