Skip to content
looot docs
Esc
↑↓navigate↵open⌘Jpreview
On this page

Context API / Crawl Website & Scrape Markdown

Context.dev tool context-dev-web-crawl on looot: input fields, $0.0025 per result, output shape, and code to run it with curl, JavaScript or Python.

Crawls from a start URL, extracts each page’s content as Markdown, and returns results for all crawled pages. maxPages caps the pages (up to 100) and maxDepth the link depth (up to 5). 1 Context.dev credit per page.

  • Tool id: context-dev-web-crawl
  • Provider: Context.dev
  • Job: Crawl web (web.crawl)
  • Price: $0.0025 per result. A call that fails at the provider costs $0.

Inputs

Name Type Required Description
country string no Fetch the target page through a residential proxy in this country (ISO 3166-1 alpha-2). Example: “de”
excludeSelectors array no CSS selectors to remove before each crawled page is converted to Markdown. Applied after includeSelectors. Exclusion takes precedence: an element matching both is removed. Examples: “nav”, “footer”, “.ad-banner”, “[aria-hidden=true]”.
followSubdomains boolean no When true, follow links on subdomains of the starting URL’s domain (e.g. docs.example.com when starting from example.com). www and apex are always treated as equivalent.
includeFrames boolean no When true, the contents of iframes are rendered to Markdown for each crawled page.
includeImages boolean no Include image references in the Markdown output
includeLinks boolean no Preserve hyperlinks in the Markdown output
includeSelectors array no CSS selectors. When provided, only matching HTML subtrees (and their descendants) are kept before each crawled page is converted to Markdown. When omitted, the entire document is kept. Examples: “article.main”, “#content”, “[role=main]”.
maxAgeMs integer no Return a cached result if a prior scrape for the same parameters exists and is younger than this many milliseconds. Defaults to 1 day (86400000 ms) when omitted. Max is 30 days (2592000000 ms). Set to 0 to always scrape fresh.
maxDepth integer no Maximum link depth from the starting URL (0 = only the starting page)
maxPages integer no Maximum number of pages to crawl. Hard cap: 500. Example: 10
pdf string no PDF parsing controls. Use start/end to limit text extraction and embedded-image detection/OCR to an inclusive 1-based page range.
settleAnimations boolean no When true, waits briefly for CSS and transition animations to settle before extracting each crawled page. Defaults to false. This adds a bit of latency in exchange for more stable output on animated pages.
shortenBase64Images boolean no Truncate base64-encoded image data in the Markdown output
stopAfterMs integer no Soft time budget for the crawl in milliseconds. After each scrape, the crawler checks the elapsed time and, if exceeded, returns the pages collected so far instead of continuing. Min: 10000 (10s). Max: 110000 (110s). Default: 80000 (80s).
tags array no Optional tags for tracking usage. Up to 20 tags, each 1 to 50 characters.
timeoutMS integer no Optional timeout in milliseconds for the request. If the request takes longer than this value, it will be aborted with a 408 status code. Maximum allowed value is 300000ms (5 minutes).
url string yes The starting URL for the crawl (must include http:// or https:// protocol). Example: “https://www.example.com”
urlRegex string no Regex pattern. Only URLs matching this pattern will be followed and scraped. An automatic prefix scope in the form ^<starting URL> follows a redirect of the starting page. Example: “https?://[/]+/blog/”
useMainContentOnly boolean no Extract only the main content, stripping headers, footers, sidebars, and navigation
waitForMs integer no Browser wait time in milliseconds after initial page load for each crawled page. Defaults to 3500 (3.5 seconds). Min: 0. Max: 30000 (30 seconds).
zdr string no Set to enabled to bypass shared caches and omit request and response content from retained usage logs. Requires zero data retention to be enabled for your organization (contact support@context.dev), otherwise the request fails with ZDR_NOT_ENABLED. Successful ZDR responses include X-Context-ZDR: true.

Output

The run’s result holds the provider’s answer. Signed in, looot inspect context-dev-web-crawl prints its fields.

Run it

Every call needs your API token in LOOOT_TOKEN; Sign in shows how to get one. Each run also needs a new idempotency key, so a retry never pays twice. With wait: 30 the answer comes back inline when the run ends within 30 seconds. Otherwise you get the running run back: poll GET /v1/runs/<runId>.

Example input with placeholder values:

curl -X POST "https://api.looot.ai/v1/runs" \
  -H "Authorization: Bearer $LOOOT_TOKEN" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: $(uuidgen)" \
  -d '{"endpointId":"context-dev-web-crawl","input":{"url":"https://www.example.com","tags":["production","team-alpha"],"country":"de","maxPages":10,"urlRegex":"^https?://[^/]+/blog/"},"wait":30}'
const response = await fetch("https://api.looot.ai/v1/runs", {
  method: "POST",
  headers: {
    Authorization: `Bearer ${process.env.LOOOT_TOKEN}`,
    "Content-Type": "application/json",
    "Idempotency-Key": crypto.randomUUID(),
  },
  body: JSON.stringify({
    endpointId: "context-dev-web-crawl",
    input: {
      url: "https://www.example.com",
      tags: ["production", "team-alpha"],
      country: "de",
      maxPages: 10,
      urlRegex: "^https?://[^/]+/blog/",
    },
    wait: 30,
  }),
});
const run = await response.json();
console.log(run.status, run.result);
import os
import uuid

import requests

response = requests.post(
    "https://api.looot.ai/v1/runs",
    headers={
        "Authorization": f"Bearer {os.environ['LOOOT_TOKEN']}",
        "Idempotency-Key": str(uuid.uuid4()),
    },
    json={
        "endpointId": "context-dev-web-crawl",
        "input": {
            "url": "https://www.example.com",
            "tags": ["production", "team-alpha"],
            "country": "de",
            "maxPages": 10,
            "urlRegex": "^https?://[^/]+/blog/",
        },
        "wait": 30,
    },
    timeout=90,
)
run = response.json()
print(run["status"], run.get("result"))

To let looot pick among every provider of this job instead, send job:web.crawl as endpointId; see the job page.

Was this page helpful?