---
title: "Context API / Crawl Website & Scrape Markdown"
description: "Context.dev tool context-dev-web-crawl on looot: input fields, $0.0025 per result, output shape, and code to run it with curl, JavaScript or Python."
sidebar:
  hidden: true
---

{/* Generated by scripts/generate-api-reference.mjs from data/api-reference.json. Do not edit. */}

Crawls from a start URL, extracts each page's content as Markdown, and returns results for all crawled pages. maxPages caps the pages (up to 100) and maxDepth the link depth (up to 5). 1 Context.dev credit per page.

- **Tool id:** `context-dev-web-crawl`
- **Provider:** [Context.dev](/providers/context-dev)
- **Job:** [Crawl web](/reference/jobs/web-crawl) (`web.crawl`)
- **Price:** $0.0025 per result. A call that fails at the provider costs $0.

## Inputs

| Name | Type | Required | Description |
| --- | --- | --- | --- |
| `country` | string | no | Fetch the target page through a residential proxy in this country (ISO 3166-1 alpha-2). Example: "de" |
| `excludeSelectors` | array | no | CSS selectors to remove before each crawled page is converted to Markdown. Applied after includeSelectors. Exclusion takes precedence: an element matching both is removed. Examples: "nav", "footer", ".ad-banner", "[aria-hidden=true]". |
| `followSubdomains` | boolean | no | When true, follow links on subdomains of the starting URL's domain (e.g. docs.example.com when starting from example.com). www and apex are always treated as equivalent. |
| `includeFrames` | boolean | no | When true, the contents of iframes are rendered to Markdown for each crawled page. |
| `includeImages` | boolean | no | Include image references in the Markdown output |
| `includeLinks` | boolean | no | Preserve hyperlinks in the Markdown output |
| `includeSelectors` | array | no | CSS selectors. When provided, only matching HTML subtrees (and their descendants) are kept before each crawled page is converted to Markdown. When omitted, the entire document is kept. Examples: "article.main", "#content", "[role=main]". |
| `maxAgeMs` | integer | no | Return a cached result if a prior scrape for the same parameters exists and is younger than this many milliseconds. Defaults to 1 day (86400000 ms) when omitted. Max is 30 days (2592000000 ms). Set to 0 to always scrape fresh. |
| `maxDepth` | integer | no | Maximum link depth from the starting URL (0 = only the starting page) |
| `maxPages` | integer | no | Maximum number of pages to crawl. Hard cap: 500. Example: 10 |
| `pdf` | string | no | PDF parsing controls. Use start/end to limit text extraction and embedded-image detection/OCR to an inclusive 1-based page range. |
| `settleAnimations` | boolean | no | When true, waits briefly for CSS and transition animations to settle before extracting each crawled page. Defaults to false. This adds a bit of latency in exchange for more stable output on animated pages. |
| `shortenBase64Images` | boolean | no | Truncate base64-encoded image data in the Markdown output |
| `stopAfterMs` | integer | no | Soft time budget for the crawl in milliseconds. After each scrape, the crawler checks the elapsed time and, if exceeded, returns the pages collected so far instead of continuing. Min: 10000 (10s). Max: 110000 (110s). Default: 80000 (80s). |
| `tags` | array | no | Optional tags for tracking usage. Up to 20 tags, each 1 to 50 characters. |
| `timeoutMS` | integer | no | Optional timeout in milliseconds for the request. If the request takes longer than this value, it will be aborted with a 408 status code. Maximum allowed value is 300000ms (5 minutes). |
| `url` | string | yes | The starting URL for the crawl (must include http:// or https:// protocol). Example: "https://www.example.com" |
| `urlRegex` | string | no | Regex pattern. Only URLs matching this pattern will be followed and scraped. An automatic prefix scope in the form ^&lt;starting URL&gt; follows a redirect of the starting page. Example: "^https?://[^/]+/blog/" |
| `useMainContentOnly` | boolean | no | Extract only the main content, stripping headers, footers, sidebars, and navigation |
| `waitForMs` | integer | no | Browser wait time in milliseconds after initial page load for each crawled page. Defaults to 3500 (3.5 seconds). Min: 0. Max: 30000 (30 seconds). |
| `zdr` | string | no | Set to enabled to bypass shared caches and omit request and response content from retained usage logs. Requires zero data retention to be enabled for your organization (contact support@context.dev), otherwise the request fails with ZDR_NOT_ENABLED. Successful ZDR responses include X-Context-ZDR: true. |

## Output

The run's `result` holds the provider's answer. Signed in, `looot inspect context-dev-web-crawl` prints its fields.

## Run it

Every call needs your API token in `LOOOT_TOKEN`; [Sign in](/get-started/sign-in#use-the-token-in-scripts-and-agents) shows how to get one. Each run also needs a new idempotency key, so a retry never pays twice. With `wait: 30` the answer comes back inline when the run ends within 30 seconds. Otherwise you get the running run back: poll `GET /v1/runs/<runId>`.

Example input with placeholder values:

<CodeGroup>

```bash curl
curl -X POST "https://api.looot.ai/v1/runs" \
  -H "Authorization: Bearer $LOOOT_TOKEN" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: $(uuidgen)" \
  -d '{"endpointId":"context-dev-web-crawl","input":{"url":"https://www.example.com","tags":["production","team-alpha"],"country":"de","maxPages":10,"urlRegex":"^https?://[^/]+/blog/"},"wait":30}'
```

```js JavaScript
const response = await fetch("https://api.looot.ai/v1/runs", {
  method: "POST",
  headers: {
    Authorization: `Bearer ${process.env.LOOOT_TOKEN}`,
    "Content-Type": "application/json",
    "Idempotency-Key": crypto.randomUUID(),
  },
  body: JSON.stringify({
    endpointId: "context-dev-web-crawl",
    input: {
      url: "https://www.example.com",
      tags: ["production", "team-alpha"],
      country: "de",
      maxPages: 10,
      urlRegex: "^https?://[^/]+/blog/",
    },
    wait: 30,
  }),
});
const run = await response.json();
console.log(run.status, run.result);
```

```python Python
import os
import uuid

import requests

response = requests.post(
    "https://api.looot.ai/v1/runs",
    headers={
        "Authorization": f"Bearer {os.environ['LOOOT_TOKEN']}",
        "Idempotency-Key": str(uuid.uuid4()),
    },
    json={
        "endpointId": "context-dev-web-crawl",
        "input": {
            "url": "https://www.example.com",
            "tags": ["production", "team-alpha"],
            "country": "de",
            "maxPages": 10,
            "urlRegex": "^https?://[^/]+/blog/",
        },
        "wait": 30,
    },
    timeout=90,
)
run = response.json()
print(run["status"], run.get("result"))
```

</CodeGroup>

To let looot pick among every provider of this job instead, send `job:web.crawl` as `endpointId`; see [the job page](/reference/jobs/web-crawl).
