Exa / Get page contents for URLs
Exa tool exa-contents on looot: input fields, $0.001 per result, output shape, and code to run it with curl, JavaScript or Python.
Fetches text, highlights or a summary for up to 100 URLs or Exa result ids in one call. Send urls or ids and pick text, highlights or summary; livecrawl and maxAgeHours control freshness.
- Tool id:
exa-contents - Provider: Exa
- Job: Extract the article behind a URL (
web.article.extract) - Price: $0.001 per result. A call that fails at the provider costs $0.
Inputs
| Name | Type | Required | Description |
|---|---|---|---|
highlights |
unknown | no | Text snippets Exa’s LLM identifies as most relevant from each page. true for defaults, or an object to steer selection with your own query. |
ids |
array | no | Document ids obtained from a prior /search call, instead of urls. |
livecrawl |
string | no | Deprecated – use maxAgeHours instead (maxAgeHours: 0 replaces livecrawl: “always”). Kept only because the spec still accepts it; do not send alongside maxAgeHours. |
maxAgeHours |
integer | no | Max age of cached content in hours; 0 fetches fresh, -1 always uses cache. Deprecated livecrawl (never/always/fallback/preferred) is superseded by this – do not send both. |
summary |
unknown | no | LLM-generated summary of the webpage, optionally guided by a query or returned structured against a JSON schema. |
text |
unknown | no | Full page text for each result. true for defaults with default settings, or an object for advanced control. |
urls |
array | yes | URLs to crawl. |
Output
Shape of the run’s result, checked against 3 real answers:
{ results: { id, url, text, image, title, author, entities: { id, type, version, properties: { ... } }[] }[], statuses: { id, error: { tag, httpStatusCode }, source, status }[], requestId, searchTime, costDollars: { total, contents: { text } } }
Run it
Every call needs your API token in LOOOT_TOKEN; Sign in shows how to get one. Each run also needs a new idempotency key, so a retry never pays twice. With wait: 30 the answer comes back inline when the run ends within 30 seconds. Otherwise you get the running run back: poll GET /v1/runs/<runId>.
Example input with placeholder values:
curl -X POST "https://api.looot.ai/v1/runs" \
-H "Authorization: Bearer $LOOOT_TOKEN" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: $(uuidgen)" \
-d '{"endpointId":"exa-contents","input":{"urls":["https://arxiv.org/pdf/2307.06435"]},"wait":30}'const response = await fetch("https://api.looot.ai/v1/runs", {
method: "POST",
headers: {
Authorization: `Bearer ${process.env.LOOOT_TOKEN}`,
"Content-Type": "application/json",
"Idempotency-Key": crypto.randomUUID(),
},
body: JSON.stringify({
endpointId: "exa-contents",
input: {
urls: ["https://arxiv.org/pdf/2307.06435"],
},
wait: 30,
}),
});
const run = await response.json();
console.log(run.status, run.result);import os
import uuid
import requests
response = requests.post(
"https://api.looot.ai/v1/runs",
headers={
"Authorization": f"Bearer {os.environ['LOOOT_TOKEN']}",
"Idempotency-Key": str(uuid.uuid4()),
},
json={
"endpointId": "exa-contents",
"input": {
"urls": ["https://arxiv.org/pdf/2307.06435"],
},
"wait": 30,
},
timeout=90,
)
run = response.json()
print(run["status"], run.get("result"))To let looot pick among every provider of this job instead, send job:web.article.extract as endpointId; see the job page.