Skip to content
looot docs
Esc
↑↓navigate↵open⌘Jpreview
On this page

Schema web extract

Run job:web.extract.schema through the looot API: 1 provider, from $0.000333 per result. Inputs, prices, output shape and code for curl, JavaScript and Python.

Schema web extract.

Job id web.extract.schema, in Scrapers (Web data). As of 2026-10-01, 1 provider serve it through 1 tool. Run job:web.extract.schema and looot picks one of them; with fallback on, a miss moves on to the next. See Jobs.

Inputs

This job has one tool, so it takes that tool’s fields. Any key you send passes through unchanged.

Name Type Required Description
enableWebSearch boolean no When true, the extraction will use web search to find additional data
ignoreInvalidURLs boolean no If invalid URLs are specified in the urls array, they will be ignored. Instead of them failing the entire request, an extract using the remaining valid URLs will be performed, and the invalid URLs will be returned in the invalidURLs field of the response.
ignoreSitemap boolean no When true, sitemap.xml files will be ignored during website scanning
includeSubdomains boolean no When true, subdomains of the provided URLs will also be scanned
prompt string no Prompt to guide the extraction process
schema string no Schema to define the structure of the extracted data. Must conform to JSON Schema.
scrapeOptions string no
showSources boolean no When true, the sources used to extract the data will be included in the response as sources key
urls array yes

Providers and prices

Provider Tool Price
Firecrawl firecrawl-extract $0.000333 per result

A call that fails at the provider costs $0. See What’s free.

Output

Shape of the run’s result, checked against 1 real answer:

{ data: { pageTitle }, status, success, warnings: string[], expiresAt, tokensUsed, creditsUsed, replacement }

Run it

Every call needs your API token in LOOOT_TOKEN; Sign in shows how to get one. Each run also needs a new idempotency key, so a retry never pays twice. With wait: 30 the answer comes back inline when the run ends within 30 seconds. Otherwise you get the running run back: poll GET /v1/runs/<runId>.

Example input with placeholder values:

curl -X POST "https://api.looot.ai/v1/runs" \
  -H "Authorization: Bearer $LOOOT_TOKEN" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: $(uuidgen)" \
  -d '{"endpointId":"job:web.extract.schema","input":{"urls":[]},"wait":30}'
const response = await fetch("https://api.looot.ai/v1/runs", {
  method: "POST",
  headers: {
    Authorization: `Bearer ${process.env.LOOOT_TOKEN}`,
    "Content-Type": "application/json",
    "Idempotency-Key": crypto.randomUUID(),
  },
  body: JSON.stringify({
    endpointId: "job:web.extract.schema",
    input: {
      urls: [],
    },
    wait: 30,
  }),
});
const run = await response.json();
console.log(run.status, run.result);
import os
import uuid

import requests

response = requests.post(
    "https://api.looot.ai/v1/runs",
    headers={
        "Authorization": f"Bearer {os.environ['LOOOT_TOKEN']}",
        "Idempotency-Key": str(uuid.uuid4()),
    },
    json={
        "endpointId": "job:web.extract.schema",
        "input": {
            "urls": [],
        },
        "wait": 30,
    },
    timeout=90,
)
run = response.json()
print(run["status"], run.get("result"))

To pin one provider, send its tool id as endpointId instead of job:web.extract.schema.

Was this page helpful?