Skip to content
looot docs
Esc
↑↓navigate↵open⌘Jpreview
On this page

Tavily Search and Extract API / Retrieve raw web content from specified URLs

Tavily Search and Extract API tool tavily-extract on looot: input fields, $0.0016 per result, output shape, and code to run it with curl, JavaScript or Python.

Extract web page content from one or more specified URLs using Tavily Extract.

Inputs

Name Type Required Description
chunks_per_source integer no Chunks are short content snippets (maximum 500 characters each) pulled directly from the source. Use chunks_per_source to define the maximum number of relevant chunks returned per source and to control the raw_content length. Chunks will appear in the raw_content field as: <chunk 1> [...] <chunk 2> [...] <chunk 3>. Available only when query is provided. Must be between 1 and 5.
extract_depth string no The depth of the extraction process. advanced extraction retrieves more data, including tables and embedded content, with higher success but may increase latency.basic extraction costs 1 credit per 5 successful URL extractions, while advanced extraction costs 2 credits per 5 successful URL extractions.
format string no The format of the extracted web page content. markdown returns content in markdown format. text returns plain text and may increase latency.
include_favicon boolean no Whether to include the favicon URL for each result.
include_images boolean no Include a list of images extracted from the URLs in the response. Default is false.
include_usage boolean no Whether to include credit usage information in the response. NOTE:The value may be 0 if the total successful URL extractions has not yet reached 5 calls. See our Credits & Pricing documentation for details.
query string no User intent for reranking extracted content chunks. When provided, chunks are reranked based on relevance to this query.
timeout number no Maximum time in seconds to wait for the URL extraction before timing out. Must be between 1.0 and 60.0 seconds. If not specified, default timeouts are applied based on extract_depth: 10 seconds for basic extraction and 30 seconds for advanced extraction.
urls array yes One or more URLs to extract content from (Tavily accepts up to 5 per request).

Output

Shape of the run’s result, checked against 2 real answers:

{ results: { url, title, images: unknown[], raw_content }[], request_id, response_time, failed_results: unknown[] }

Run it

Every call needs your API token in LOOOT_TOKEN; Sign in shows how to get one. Each run also needs a new idempotency key, so a retry never pays twice. With wait: 30 the answer comes back inline when the run ends within 30 seconds. Otherwise you get the running run back: poll GET /v1/runs/<runId>.

Example input with placeholder values:

curl -X POST "https://api.looot.ai/v1/runs" \
  -H "Authorization: Bearer $LOOOT_TOKEN" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: $(uuidgen)" \
  -d '{"endpointId":"tavily-extract","input":{"urls":[]},"wait":30}'
const response = await fetch("https://api.looot.ai/v1/runs", {
  method: "POST",
  headers: {
    Authorization: `Bearer ${process.env.LOOOT_TOKEN}`,
    "Content-Type": "application/json",
    "Idempotency-Key": crypto.randomUUID(),
  },
  body: JSON.stringify({
    endpointId: "tavily-extract",
    input: {
      urls: [],
    },
    wait: 30,
  }),
});
const run = await response.json();
console.log(run.status, run.result);
import os
import uuid

import requests

response = requests.post(
    "https://api.looot.ai/v1/runs",
    headers={
        "Authorization": f"Bearer {os.environ['LOOOT_TOKEN']}",
        "Idempotency-Key": str(uuid.uuid4()),
    },
    json={
        "endpointId": "tavily-extract",
        "input": {
            "urls": [],
        },
        "wait": 30,
    },
    timeout=90,
)
run = response.json()
print(run["status"], run.get("result"))

To let looot pick among every provider of this job instead, send job:extract.get as endpointId; see the job page.

Was this page helpful?