Skip to content
looot docs
Esc
↑↓navigate↵open⌘Jpreview
On this page

Context API / Crawl Sitemap

Context.dev tool context-dev-web-scrape-sitemap on looot: input fields, $0.0025 per call, output shape, and code to run it with curl, JavaScript or Python.

Crawl an entire website’s sitemap and return all discovered page URLs. Pass search to have the crawled sitemap filtered down to the pages about a phrase (for example pricing and plans or api authentication docs), most relevant first, a searched crawl scans the whole sitemap and costs 2 credits instead of 1.

Inputs

Name Type Required Description
domain string yes Domain to build a sitemap for. Example: “example.com”
headers string no Optional outbound HTTP headers forwarded only to the target URL, sent as deep-object query params such as headers[X-Custom]=value. When provided, caching is bypassed: the result is neither read from nor written to cache.
maxLinks integer no Maximum number of links to return from the sitemap crawl. Defaults to 10,000. Minimum is 1, maximum is 100,000.
search string no Optional search phrase. When provided, the crawled sitemap is filtered to the pages whose URLs are about that phrase, most relevant first, and the request costs 2 credits instead of 1. Example: “help center and troubleshooting articles”
sitemapUrl string no Optional explicit sitemap URL. When provided, exactly this sitemap is crawled instead of discovering the domain’s sitemaps.
tags array no Optional tags for tracking usage. Up to 20 tags, each 1 to 50 characters.
timeoutMS integer no Optional timeout in milliseconds for the request. If the request takes longer than this value, it will be aborted with a 408 status code. Maximum allowed value is 300000ms (5 minutes).
urlRegex string no Optional RE2-compatible regex pattern. Only URLs matching this pattern are returned and counted against maxLinks. Example: “https?://[/]+/blog/”
zdr string no Set to enabled to bypass shared caches and omit request and response content from retained usage logs. Requires zero data retention to be enabled for your organization (contact support@context.dev), otherwise the request fails with ZDR_NOT_ENABLED. Successful ZDR responses include X-Context-ZDR: true.

Output

Shape of the run’s result, checked against 1 real answer:

{ meta: { errors, sitemapsFetched, sitemapsSkipped, sitemapsDiscovered }, urls: unknown[], domain, success, request_id, key_metadata: { credits_consumed, credits_remaining } }

Run it

Every call needs your API token in LOOOT_TOKEN; Sign in shows how to get one. Each run also needs a new idempotency key, so a retry never pays twice. With wait: 30 the answer comes back inline when the run ends within 30 seconds. Otherwise you get the running run back: poll GET /v1/runs/<runId>.

Example input with placeholder values:

curl -X POST "https://api.looot.ai/v1/runs" \
  -H "Authorization: Bearer $LOOOT_TOKEN" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: $(uuidgen)" \
  -d '{"endpointId":"context-dev-web-scrape-sitemap","input":{"domain":"example.com","tags":["production","team-alpha"],"search":"help center and troubleshooting articles","urlRegex":"^https?://[^/]+/blog/"},"wait":30}'
const response = await fetch("https://api.looot.ai/v1/runs", {
  method: "POST",
  headers: {
    Authorization: `Bearer ${process.env.LOOOT_TOKEN}`,
    "Content-Type": "application/json",
    "Idempotency-Key": crypto.randomUUID(),
  },
  body: JSON.stringify({
    endpointId: "context-dev-web-scrape-sitemap",
    input: {
      domain: "example.com",
      tags: ["production", "team-alpha"],
      search: "help center and troubleshooting articles",
      urlRegex: "^https?://[^/]+/blog/",
    },
    wait: 30,
  }),
});
const run = await response.json();
console.log(run.status, run.result);
import os
import uuid

import requests

response = requests.post(
    "https://api.looot.ai/v1/runs",
    headers={
        "Authorization": f"Bearer {os.environ['LOOOT_TOKEN']}",
        "Idempotency-Key": str(uuid.uuid4()),
    },
    json={
        "endpointId": "context-dev-web-scrape-sitemap",
        "input": {
            "domain": "example.com",
            "tags": ["production", "team-alpha"],
            "search": "help center and troubleshooting articles",
            "urlRegex": "^https?://[^/]+/blog/",
        },
        "wait": 30,
    },
    timeout=90,
)
run = response.json()
print(run["status"], run.get("result"))

To let looot pick among every provider of this job instead, send job:web.site.map as endpointId; see the job page.

Was this page helpful?