Skip to content
looot docs
Esc
↑↓navigate↵open⌘Jpreview
On this page

Tavily Search and Extract API / Initiate a web crawl from a base URL

Tavily Search and Extract API tool tavily-crawl on looot: input fields, $0.0024 per result, output shape, and code to run it with curl, JavaScript or Python.

Tavily Crawl is a graph-based website traversal tool that can explore hundreds of paths in parallel with built-in extraction and intelligent discovery.

Inputs

Name Type Required Description
allow_external boolean no Whether to include external domain links in the final results list.
chunks_per_source integer no Chunks are short content snippets (maximum 500 characters each) pulled directly from the source. Use chunks_per_source to define the maximum number of relevant chunks returned per source and to control the raw_content length. Chunks will appear in the raw_content field as: <chunk 1> [...] <chunk 2> [...] <chunk 3>. Available only when instructions are provided. Must be between 1 and 5.
exclude_domains array no Regex patterns to exclude specific domains or subdomains from crawling (e.g., ^private\.example\.com$).
exclude_paths array no Regex patterns to exclude URLs with specific path patterns (e.g., /private/.*, /admin/.*).
extract_depth string no Advanced extraction retrieves more data, including tables and embedded content, with higher success but may increase latency. basic extraction costs 1 credit per 5 successful extractions, while advanced extraction costs 2 credits per 5 successful extractions.
format string no The format of the extracted web page content. markdown returns content in markdown format. text returns plain text and may increase latency.
include_favicon boolean no Whether to include the favicon URL for each result.
include_images boolean no Whether to include images in the crawl results.
include_usage boolean no Whether to include credit usage information in the response. NOTE:The value may be 0 if the total use of /extract and /map have not yet reached minimum requirements. See our Credits & Pricing documentation for details.
instructions string no Natural language instructions for the crawler. When specified, the mapping cost increases to 2 API credits per 10 successful pages instead of 1 API credit per 10 pages. Example: “Find all pages about the Python SDK”
limit integer no Total number of links the crawler will process before stopping.
max_breadth integer no Max number of links to follow per level of the tree (i.e., per page).
max_depth integer no Max depth of the crawl. Defines how far from the base URL the crawler can explore.
select_domains array no Regex patterns to select crawling to specific domains or subdomains (e.g., ^docs\.example\.com$).
select_paths array no Regex patterns to select only URLs with specific path patterns (e.g., /docs/.*, /api/v1.*).
timeout number no Maximum time in seconds to wait for the crawl operation before timing out. Must be between 10 and 150 seconds.
url string yes The root URL to begin the crawl. Example: “docs.tavily.com”

Output

Shape of the run’s result, the expected shape, not yet checked against real answers:

{ usage: { credits }, results: { url, raw_content }[], base_url, request_id, res... ...

The shape is cut here. Signed in, looot inspect tavily-crawl prints all of it.

Run it

Every call needs your API token in LOOOT_TOKEN; Sign in shows how to get one. Each run also needs a new idempotency key, so a retry never pays twice. With wait: 30 the answer comes back inline when the run ends within 30 seconds. Otherwise you get the running run back: poll GET /v1/runs/<runId>.

Example input with placeholder values:

curl -X POST "https://api.looot.ai/v1/runs" \
  -H "Authorization: Bearer $LOOOT_TOKEN" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: $(uuidgen)" \
  -d '{"endpointId":"tavily-crawl","input":{"url":"docs.tavily.com","instructions":"Find all pages about the Python SDK"},"wait":30}'
const response = await fetch("https://api.looot.ai/v1/runs", {
  method: "POST",
  headers: {
    Authorization: `Bearer ${process.env.LOOOT_TOKEN}`,
    "Content-Type": "application/json",
    "Idempotency-Key": crypto.randomUUID(),
  },
  body: JSON.stringify({
    endpointId: "tavily-crawl",
    input: {
      url: "docs.tavily.com",
      instructions: "Find all pages about the Python SDK",
    },
    wait: 30,
  }),
});
const run = await response.json();
console.log(run.status, run.result);
import os
import uuid

import requests

response = requests.post(
    "https://api.looot.ai/v1/runs",
    headers={
        "Authorization": f"Bearer {os.environ['LOOOT_TOKEN']}",
        "Idempotency-Key": str(uuid.uuid4()),
    },
    json={
        "endpointId": "tavily-crawl",
        "input": {
            "url": "docs.tavily.com",
            "instructions": "Find all pages about the Python SDK",
        },
        "wait": 30,
    },
    timeout=90,
)
run = response.json()
print(run["status"], run.get("result"))

To let looot pick among every provider of this job instead, send job:crawl.get as endpointId; see the job page.

Was this page helpful?