Skip to content
looot docs
Esc
↑↓navigate↵open⌘Jpreview
On this page

Context API / Extract Structured Website Data

Context.dev tool context-dev-web-extract on looot: input fields, $0.025 per call, output shape, and code to run it with curl, JavaScript or Python.

Crawl a website, use the provided JSON Schema and instructions to prioritize relevant internal links, and extract structured data from the selected pages.

  • Tool id: context-dev-web-extract
  • Provider: Context.dev
  • Job: Extract web (web.extract)
  • Price: $0.025 per call. A call that fails at the provider costs $0.

Inputs

Name Type Required Description
actions array no Optional browser actions executed in order on the requested page after it loads, before links are discovered or additional pages are crawled. Requires a paid plan. When actions are provided and stopAfterMs is omitted, the crawl budget defaults to 110000 ms.
factCheck boolean no When true, every returned value must be grounded in facts stated on the page; fields that cannot be supported by the page are returned as null/empty. When false (default), the model may make reasonable inferences and derivations from the page content (e.g. ideal customer, competitor analysis, recommendations) while keeping verifiable specifics (names, quotes, URLs, dates, metrics) faithful to the source.
followSubdomains boolean no When true, follow links on subdomains of the starting URL’s domain.
includeFrames boolean no When true, iframe contents are included in Markdown before extraction.
instructions string no Optional extraction guidance, such as which facts to prioritize or how to interpret fields in the schema.
maxAgeMs integer no Return cached scrape results if a prior scrape for the same parameters is younger than this many milliseconds. Defaults to 7 days (604800000 ms).
maxDepth integer no Optional maximum link depth from the starting URL (0 = only the starting page). If omitted, there is no crawl depth limit.
maxPages integer no Maximum number of pages to analyze for extraction. Hard cap: 50. Defaults to 5.
pdf string no
schema object yes JSON Schema for the returned data object. Image fields such as image_urls or product_photos automatically make page image references available to extraction, so product data and photos can be returned in one call. TypeScript Zod users can pass a JSON Schema generated from a Zod object; Python users can pass the equivalent JSON Schema object.
settleAnimations boolean no When true, waits briefly for CSS and transition animations to settle before extracting each crawled page. Defaults to false. This adds a bit of latency in exchange for more stable output on animated pages.
stopAfterMs integer no Soft time budget for the crawl in milliseconds. Min: 10000 (10s). Max: 110000 (110s). Defaults to 80000 (80s), or 110000 (110s) when browser actions are provided.
tags array no Optional tags for tracking usage. Up to 20 tags, each 1 to 50 characters.
timeoutMS integer no Optional timeout in milliseconds for the request. If the request takes longer than this value, it will be aborted with a 408 status code. Maximum allowed value is 300000ms (5 minutes).
url string yes The starting website URL to crawl and extract from. Must include http:// or https://.
waitForMs integer no Optional browser wait time in milliseconds after initial page load for each crawled page.

Output

Shape of the run’s result, checked against 1 real answer:

{ url, data: { title }, status, metadata: { numUrls, numFailed, numBlocked, numS... ...

The shape is cut here. Signed in, looot inspect context-dev-web-extract prints all of it.

Run it

Every call needs your API token in LOOOT_TOKEN; Sign in shows how to get one. Each run also needs a new idempotency key, so a retry never pays twice. With wait: 30 the answer comes back inline when the run ends within 30 seconds. Otherwise you get the running run back: poll GET /v1/runs/<runId>.

Example input with placeholder values:

curl -X POST "https://api.looot.ai/v1/runs" \
  -H "Authorization: Bearer $LOOOT_TOKEN" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: $(uuidgen)" \
  -d '{"endpointId":"context-dev-web-extract","input":{"url":"https://example.com","schema":{"type":"object","required":["mission_statement","case_studies"],"properties":{"case_studies":{"type":"array","items":{"type":"object","required":["title","url"],"properties":{"url":{"type":"string"},"title":{"type":"string"}},"additionalProperties":false}},"mission_statement":{"type":"string","description":"The company'\''s stated mission."}},"additionalProperties":false},"tags":["production","team-alpha"]},"wait":30}'
const response = await fetch("https://api.looot.ai/v1/runs", {
  method: "POST",
  headers: {
    Authorization: `Bearer ${process.env.LOOOT_TOKEN}`,
    "Content-Type": "application/json",
    "Idempotency-Key": crypto.randomUUID(),
  },
  body: JSON.stringify({
    endpointId: "context-dev-web-extract",
    input: {
      url: "https://example.com",
      schema: {
        type: "object",
        required: ["mission_statement", "case_studies"],
        properties: {
          case_studies: {
            type: "array",
            items: {
              type: "object",
              required: ["title", "url"],
              properties: {
                url: {
                  type: "string",
                },
                title: {
                  type: "string",
                },
              },
              additionalProperties: false,
            },
          },
          mission_statement: {
            type: "string",
            description: "The company's stated mission.",
          },
        },
        additionalProperties: false,
      },
      tags: ["production", "team-alpha"],
    },
    wait: 30,
  }),
});
const run = await response.json();
console.log(run.status, run.result);
import os
import uuid

import requests

response = requests.post(
    "https://api.looot.ai/v1/runs",
    headers={
        "Authorization": f"Bearer {os.environ['LOOOT_TOKEN']}",
        "Idempotency-Key": str(uuid.uuid4()),
    },
    json={
        "endpointId": "context-dev-web-extract",
        "input": {
            "url": "https://example.com",
            "schema": {
                "type": "object",
                "required": ["mission_statement", "case_studies"],
                "properties": {
                    "case_studies": {
                        "type": "array",
                        "items": {
                            "type": "object",
                            "required": ["title", "url"],
                            "properties": {
                                "url": {
                                    "type": "string",
                                },
                                "title": {
                                    "type": "string",
                                },
                            },
                            "additionalProperties": False,
                        },
                    },
                    "mission_statement": {
                        "type": "string",
                        "description": "The company's stated mission.",
                    },
                },
                "additionalProperties": False,
            },
            "tags": ["production", "team-alpha"],
        },
        "wait": 30,
    },
    timeout=90,
)
run = response.json()
print(run["status"], run.get("result"))

To let looot pick among every provider of this job instead, send job:web.extract as endpointId; see the job page.

Was this page helpful?