---
title: "Tavily Search and Extract API / Initiate a web crawl from a base URL"
description: "Tavily Search and Extract API tool tavily-crawl on looot: input fields, $0.0024 per result, output shape, and code to run it with curl, JavaScript or Python."
sidebar:
  hidden: true
---

{/* Generated by scripts/generate-api-reference.mjs from data/api-reference.json. Do not edit. */}

Tavily Crawl is a graph-based website traversal tool that can explore hundreds of paths in parallel with built-in extraction and intelligent discovery.

- **Tool id:** `tavily-crawl`
- **Provider:** [Tavily Search and Extract API](/providers/tavily)
- **Job:** [Get crawl](/reference/jobs/crawl-get) (`crawl.get`)
- **Price:** $0.0024 per result. A call that fails at the provider costs $0.

## Inputs

| Name | Type | Required | Description |
| --- | --- | --- | --- |
| `allow_external` | boolean | no | Whether to include external domain links in the final results list. |
| `chunks_per_source` | integer | no | Chunks are short content snippets (maximum 500 characters each) pulled directly from the source. Use `chunks_per_source` to define the maximum number of relevant chunks returned per source and to control the `raw_content` length. Chunks will appear in the `raw_content` field as: `&lt;chunk 1&gt; [...] &lt;chunk 2&gt; [...] &lt;chunk 3&gt;`. Available only when `instructions` are provided. Must be between 1 and 5. |
| `exclude_domains` | array | no | Regex patterns to exclude specific domains or subdomains from crawling (e.g., `^private\.example\.com$`). |
| `exclude_paths` | array | no | Regex patterns to exclude URLs with specific path patterns (e.g., `/private/.*`, `/admin/.*`). |
| `extract_depth` | string | no | Advanced extraction retrieves more data, including tables and embedded content, with higher success but may increase latency. `basic` extraction costs 1 credit per 5 successful extractions, while `advanced` extraction costs 2 credits per 5 successful extractions. |
| `format` | string | no | The format of the extracted web page content. `markdown` returns content in markdown format. `text` returns plain text and may increase latency. |
| `include_favicon` | boolean | no | Whether to include the favicon URL for each result. |
| `include_images` | boolean | no | Whether to include images in the crawl results. |
| `include_usage` | boolean | no | Whether to include credit usage information in the response. `NOTE:`The value may be 0 if the total use of /extract and /map have not yet reached minimum requirements. See our [Credits & Pricing documentation](https://docs.tavily.com/documentation/api-credits) for details. |
| `instructions` | string | no | Natural language instructions for the crawler. When specified, the mapping cost increases to 2 API credits per 10 successful pages instead of 1 API credit per 10 pages. Example: "Find all pages about the Python SDK" |
| `limit` | integer | no | Total number of links the crawler will process before stopping. |
| `max_breadth` | integer | no | Max number of links to follow per level of the tree (i.e., per page). |
| `max_depth` | integer | no | Max depth of the crawl. Defines how far from the base URL the crawler can explore. |
| `select_domains` | array | no | Regex patterns to select crawling to specific domains or subdomains (e.g., `^docs\.example\.com$`). |
| `select_paths` | array | no | Regex patterns to select only URLs with specific path patterns (e.g., `/docs/.*`, `/api/v1.*`). |
| `timeout` | number | no | Maximum time in seconds to wait for the crawl operation before timing out. Must be between 10 and 150 seconds. |
| `url` | string | yes | The root URL to begin the crawl. Example: "docs.tavily.com" |

## Output

Shape of the run's `result`, the expected shape, not yet checked against real answers:

```txt
{ usage: { credits }, results: { url, raw_content }[], base_url, request_id, res... ...
```

The shape is cut here. Signed in, `looot inspect tavily-crawl` prints all of it.

## Run it

Every call needs your API token in `LOOOT_TOKEN`; [Sign in](/get-started/sign-in#use-the-token-in-scripts-and-agents) shows how to get one. Each run also needs a new idempotency key, so a retry never pays twice. With `wait: 30` the answer comes back inline when the run ends within 30 seconds. Otherwise you get the running run back: poll `GET /v1/runs/<runId>`.

Example input with placeholder values:

<CodeGroup>

```bash curl
curl -X POST "https://api.looot.ai/v1/runs" \
  -H "Authorization: Bearer $LOOOT_TOKEN" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: $(uuidgen)" \
  -d '{"endpointId":"tavily-crawl","input":{"url":"docs.tavily.com","instructions":"Find all pages about the Python SDK"},"wait":30}'
```

```js JavaScript
const response = await fetch("https://api.looot.ai/v1/runs", {
  method: "POST",
  headers: {
    Authorization: `Bearer ${process.env.LOOOT_TOKEN}`,
    "Content-Type": "application/json",
    "Idempotency-Key": crypto.randomUUID(),
  },
  body: JSON.stringify({
    endpointId: "tavily-crawl",
    input: {
      url: "docs.tavily.com",
      instructions: "Find all pages about the Python SDK",
    },
    wait: 30,
  }),
});
const run = await response.json();
console.log(run.status, run.result);
```

```python Python
import os
import uuid

import requests

response = requests.post(
    "https://api.looot.ai/v1/runs",
    headers={
        "Authorization": f"Bearer {os.environ['LOOOT_TOKEN']}",
        "Idempotency-Key": str(uuid.uuid4()),
    },
    json={
        "endpointId": "tavily-crawl",
        "input": {
            "url": "docs.tavily.com",
            "instructions": "Find all pages about the Python SDK",
        },
        "wait": 30,
    },
    timeout=90,
)
run = response.json()
print(run["status"], run.get("result"))
```

</CodeGroup>

To let looot pick among every provider of this job instead, send `job:crawl.get` as `endpointId`; see [the job page](/reference/jobs/crawl-get).
