Tavily Search and Extract API / Retrieve raw web content from specified URLs
Tavily Search and Extract API tool tavily-extract on looot: input fields, $0.0016 per result, output shape, and code to run it with curl, JavaScript or Python.
Extract web page content from one or more specified URLs using Tavily Extract.
- Tool id:
tavily-extract - Provider: Tavily Search and Extract API
- Job: Get extract (
extract.get) - Price: $0.0016 per result. A call that fails at the provider costs $0.
Inputs
| Name | Type | Required | Description |
|---|---|---|---|
chunks_per_source |
integer | no | Chunks are short content snippets (maximum 500 characters each) pulled directly from the source. Use chunks_per_source to define the maximum number of relevant chunks returned per source and to control the raw_content length. Chunks will appear in the raw_content field as: <chunk 1> [...] <chunk 2> [...] <chunk 3>. Available only when query is provided. Must be between 1 and 5. |
extract_depth |
string | no | The depth of the extraction process. advanced extraction retrieves more data, including tables and embedded content, with higher success but may increase latency.basic extraction costs 1 credit per 5 successful URL extractions, while advanced extraction costs 2 credits per 5 successful URL extractions. |
format |
string | no | The format of the extracted web page content. markdown returns content in markdown format. text returns plain text and may increase latency. |
include_favicon |
boolean | no | Whether to include the favicon URL for each result. |
include_images |
boolean | no | Include a list of images extracted from the URLs in the response. Default is false. |
include_usage |
boolean | no | Whether to include credit usage information in the response. NOTE:The value may be 0 if the total successful URL extractions has not yet reached 5 calls. See our Credits & Pricing documentation for details. |
query |
string | no | User intent for reranking extracted content chunks. When provided, chunks are reranked based on relevance to this query. |
timeout |
number | no | Maximum time in seconds to wait for the URL extraction before timing out. Must be between 1.0 and 60.0 seconds. If not specified, default timeouts are applied based on extract_depth: 10 seconds for basic extraction and 30 seconds for advanced extraction. |
urls |
array | yes | One or more URLs to extract content from (Tavily accepts up to 5 per request). |
Output
Shape of the run’s result, checked against 2 real answers:
{ results: { url, title, images: unknown[], raw_content }[], request_id, response_time, failed_results: unknown[] }
Run it
Every call needs your API token in LOOOT_TOKEN; Sign in shows how to get one. Each run also needs a new idempotency key, so a retry never pays twice. With wait: 30 the answer comes back inline when the run ends within 30 seconds. Otherwise you get the running run back: poll GET /v1/runs/<runId>.
Example input with placeholder values:
curl -X POST "https://api.looot.ai/v1/runs" \
-H "Authorization: Bearer $LOOOT_TOKEN" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: $(uuidgen)" \
-d '{"endpointId":"tavily-extract","input":{"urls":[]},"wait":30}'const response = await fetch("https://api.looot.ai/v1/runs", {
method: "POST",
headers: {
Authorization: `Bearer ${process.env.LOOOT_TOKEN}`,
"Content-Type": "application/json",
"Idempotency-Key": crypto.randomUUID(),
},
body: JSON.stringify({
endpointId: "tavily-extract",
input: {
urls: [],
},
wait: 30,
}),
});
const run = await response.json();
console.log(run.status, run.result);import os
import uuid
import requests
response = requests.post(
"https://api.looot.ai/v1/runs",
headers={
"Authorization": f"Bearer {os.environ['LOOOT_TOKEN']}",
"Idempotency-Key": str(uuid.uuid4()),
},
json={
"endpointId": "tavily-extract",
"input": {
"urls": [],
},
"wait": 30,
},
timeout=90,
)
run = response.json()
print(run["status"], run.get("result"))To let looot pick among every provider of this job instead, send job:extract.get as endpointId; see the job page.