---
title: "Context API / Extract Structured Website Data"
description: "Context.dev tool context-dev-web-extract on looot: input fields, $0.025 per call, output shape, and code to run it with curl, JavaScript or Python."
sidebar:
  hidden: true
---

{/* Generated by scripts/generate-api-reference.mjs from data/api-reference.json. Do not edit. */}

Crawl a website, use the provided JSON Schema and instructions to prioritize relevant internal links, and extract structured data from the selected pages.

- **Tool id:** `context-dev-web-extract`
- **Provider:** [Context.dev](/providers/context-dev)
- **Job:** [Extract web](/reference/jobs/web-extract) (`web.extract`)
- **Price:** $0.025 per call. A call that fails at the provider costs $0.

## Inputs

| Name | Type | Required | Description |
| --- | --- | --- | --- |
| `actions` | array | no | Optional browser actions executed in order on the requested page after it loads, before links are discovered or additional pages are crawled. Requires a paid plan. When actions are provided and stopAfterMs is omitted, the crawl budget defaults to 110000 ms. |
| `factCheck` | boolean | no | When true, every returned value must be grounded in facts stated on the page; fields that cannot be supported by the page are returned as null/empty. When false (default), the model may make reasonable inferences and derivations from the page content (e.g. ideal customer, competitor analysis, recommendations) while keeping verifiable specifics (names, quotes, URLs, dates, metrics) faithful to the source. |
| `followSubdomains` | boolean | no | When true, follow links on subdomains of the starting URL's domain. |
| `includeFrames` | boolean | no | When true, iframe contents are included in Markdown before extraction. |
| `instructions` | string | no | Optional extraction guidance, such as which facts to prioritize or how to interpret fields in the schema. |
| `maxAgeMs` | integer | no | Return cached scrape results if a prior scrape for the same parameters is younger than this many milliseconds. Defaults to 7 days (604800000 ms). |
| `maxDepth` | integer | no | Optional maximum link depth from the starting URL (0 = only the starting page). If omitted, there is no crawl depth limit. |
| `maxPages` | integer | no | Maximum number of pages to analyze for extraction. Hard cap: 50. Defaults to 5. |
| `pdf` | string | no |   |
| `schema` | object | yes | JSON Schema for the returned data object. Image fields such as `image_urls` or `product_photos` automatically make page image references available to extraction, so product data and photos can be returned in one call. TypeScript Zod users can pass a JSON Schema generated from a Zod object; Python users can pass the equivalent JSON Schema object. |
| `settleAnimations` | boolean | no | When true, waits briefly for CSS and transition animations to settle before extracting each crawled page. Defaults to false. This adds a bit of latency in exchange for more stable output on animated pages. |
| `stopAfterMs` | integer | no | Soft time budget for the crawl in milliseconds. Min: 10000 (10s). Max: 110000 (110s). Defaults to 80000 (80s), or 110000 (110s) when browser actions are provided. |
| `tags` | array | no | Optional tags for tracking usage. Up to 20 tags, each 1 to 50 characters. |
| `timeoutMS` | integer | no | Optional timeout in milliseconds for the request. If the request takes longer than this value, it will be aborted with a 408 status code. Maximum allowed value is 300000ms (5 minutes). |
| `url` | string | yes | The starting website URL to crawl and extract from. Must include http:// or https://. |
| `waitForMs` | integer | no | Optional browser wait time in milliseconds after initial page load for each crawled page. |

## Output

Shape of the run's `result`, checked against 1 real answer:

```txt
{ url, data: { title }, status, metadata: { numUrls, numFailed, numBlocked, numS... ...
```

The shape is cut here. Signed in, `looot inspect context-dev-web-extract` prints all of it.

## Run it

Every call needs your API token in `LOOOT_TOKEN`; [Sign in](/get-started/sign-in#use-the-token-in-scripts-and-agents) shows how to get one. Each run also needs a new idempotency key, so a retry never pays twice. With `wait: 30` the answer comes back inline when the run ends within 30 seconds. Otherwise you get the running run back: poll `GET /v1/runs/<runId>`.

Example input with placeholder values:

<CodeGroup>

```bash curl
curl -X POST "https://api.looot.ai/v1/runs" \
  -H "Authorization: Bearer $LOOOT_TOKEN" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: $(uuidgen)" \
  -d '{"endpointId":"context-dev-web-extract","input":{"url":"https://example.com","schema":{"type":"object","required":["mission_statement","case_studies"],"properties":{"case_studies":{"type":"array","items":{"type":"object","required":["title","url"],"properties":{"url":{"type":"string"},"title":{"type":"string"}},"additionalProperties":false}},"mission_statement":{"type":"string","description":"The company'\''s stated mission."}},"additionalProperties":false},"tags":["production","team-alpha"]},"wait":30}'
```

```js JavaScript
const response = await fetch("https://api.looot.ai/v1/runs", {
  method: "POST",
  headers: {
    Authorization: `Bearer ${process.env.LOOOT_TOKEN}`,
    "Content-Type": "application/json",
    "Idempotency-Key": crypto.randomUUID(),
  },
  body: JSON.stringify({
    endpointId: "context-dev-web-extract",
    input: {
      url: "https://example.com",
      schema: {
        type: "object",
        required: ["mission_statement", "case_studies"],
        properties: {
          case_studies: {
            type: "array",
            items: {
              type: "object",
              required: ["title", "url"],
              properties: {
                url: {
                  type: "string",
                },
                title: {
                  type: "string",
                },
              },
              additionalProperties: false,
            },
          },
          mission_statement: {
            type: "string",
            description: "The company's stated mission.",
          },
        },
        additionalProperties: false,
      },
      tags: ["production", "team-alpha"],
    },
    wait: 30,
  }),
});
const run = await response.json();
console.log(run.status, run.result);
```

```python Python
import os
import uuid

import requests

response = requests.post(
    "https://api.looot.ai/v1/runs",
    headers={
        "Authorization": f"Bearer {os.environ['LOOOT_TOKEN']}",
        "Idempotency-Key": str(uuid.uuid4()),
    },
    json={
        "endpointId": "context-dev-web-extract",
        "input": {
            "url": "https://example.com",
            "schema": {
                "type": "object",
                "required": ["mission_statement", "case_studies"],
                "properties": {
                    "case_studies": {
                        "type": "array",
                        "items": {
                            "type": "object",
                            "required": ["title", "url"],
                            "properties": {
                                "url": {
                                    "type": "string",
                                },
                                "title": {
                                    "type": "string",
                                },
                            },
                            "additionalProperties": False,
                        },
                    },
                    "mission_statement": {
                        "type": "string",
                        "description": "The company's stated mission.",
                    },
                },
                "additionalProperties": False,
            },
            "tags": ["production", "team-alpha"],
        },
        "wait": 30,
    },
    timeout=90,
)
run = response.json()
print(run["status"], run.get("result"))
```

</CodeGroup>

To let looot pick among every provider of this job instead, send `job:web.extract` as `endpointId`; see [the job page](/reference/jobs/web-extract).
