---
title: "Firecrawl API / Extract structured data from pages using LLMs"
description: "Firecrawl tool firecrawl-extract on looot: input fields, $0.000333 per result, output shape, and code to run it with curl, JavaScript or Python."
sidebar:
  hidden: true
---

{/* Generated by scripts/generate-api-reference.mjs from data/api-reference.json. Do not edit. */}

Extracts structured data from one or more pages with an LLM. Send urls (up to 100). Priced from Firecrawl's reported credits, which grow with page size.

- **Tool id:** `firecrawl-extract`
- **Provider:** [Firecrawl](/providers/firecrawl)
- **Job:** [Schema web extract](/reference/jobs/web-extract-schema) (`web.extract.schema`)
- **Price:** $0.000333 per result. A call that fails at the provider costs $0.

## Inputs

| Name | Type | Required | Description |
| --- | --- | --- | --- |
| `enableWebSearch` | boolean | no | When true, the extraction will use web search to find additional data |
| `ignoreInvalidURLs` | boolean | no | If invalid URLs are specified in the urls array, they will be ignored. Instead of them failing the entire request, an extract using the remaining valid URLs will be performed, and the invalid URLs will be returned in the invalidURLs field of the response. |
| `ignoreSitemap` | boolean | no | When true, sitemap.xml files will be ignored during website scanning |
| `includeSubdomains` | boolean | no | When true, subdomains of the provided URLs will also be scanned |
| `prompt` | string | no | Prompt to guide the extraction process |
| `schema` | string | no | Schema to define the structure of the extracted data. Must conform to [JSON Schema](https://json-schema.org/). |
| `scrapeOptions` | string | no |   |
| `showSources` | boolean | no | When true, the sources used to extract the data will be included in the response as `sources` key |
| `urls` | array | yes |   |

## Output

Shape of the run's `result`, checked against 1 real answer:

```txt
{ data: { pageTitle }, status, success, warnings: string[], expiresAt, tokensUsed, creditsUsed, replacement }
```

## Run it

Every call needs your API token in `LOOOT_TOKEN`; [Sign in](/get-started/sign-in#use-the-token-in-scripts-and-agents) shows how to get one. Each run also needs a new idempotency key, so a retry never pays twice. With `wait: 30` the answer comes back inline when the run ends within 30 seconds. Otherwise you get the running run back: poll `GET /v1/runs/<runId>`.

Example input with placeholder values:

<CodeGroup>

```bash curl
curl -X POST "https://api.looot.ai/v1/runs" \
  -H "Authorization: Bearer $LOOOT_TOKEN" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: $(uuidgen)" \
  -d '{"endpointId":"firecrawl-extract","input":{"urls":[]},"wait":30}'
```

```js JavaScript
const response = await fetch("https://api.looot.ai/v1/runs", {
  method: "POST",
  headers: {
    Authorization: `Bearer ${process.env.LOOOT_TOKEN}`,
    "Content-Type": "application/json",
    "Idempotency-Key": crypto.randomUUID(),
  },
  body: JSON.stringify({
    endpointId: "firecrawl-extract",
    input: {
      urls: [],
    },
    wait: 30,
  }),
});
const run = await response.json();
console.log(run.status, run.result);
```

```python Python
import os
import uuid

import requests

response = requests.post(
    "https://api.looot.ai/v1/runs",
    headers={
        "Authorization": f"Bearer {os.environ['LOOOT_TOKEN']}",
        "Idempotency-Key": str(uuid.uuid4()),
    },
    json={
        "endpointId": "firecrawl-extract",
        "input": {
            "urls": [],
        },
        "wait": 30,
    },
    timeout=90,
)
run = response.json()
print(run["status"], run.get("result"))
```

</CodeGroup>

To let looot pick among every provider of this job instead, send `job:web.extract.schema` as `endpointId`; see [the job page](/reference/jobs/web-extract-schema).
