---
title: "Exa / Get page contents for URLs"
description: "Exa tool exa-contents on looot: input fields, $0.001 per result, output shape, and code to run it with curl, JavaScript or Python."
sidebar:
  hidden: true
---

{/* Generated by scripts/generate-api-reference.mjs from data/api-reference.json. Do not edit. */}

Fetches text, highlights or a summary for up to 100 URLs or Exa result ids in one call. Send urls or ids and pick text, highlights or summary; livecrawl and maxAgeHours control freshness.

- **Tool id:** `exa-contents`
- **Provider:** [Exa](/providers/exa)
- **Job:** [Extract the article behind a URL](/reference/jobs/web-article-extract) (`web.article.extract`)
- **Price:** $0.001 per result. A call that fails at the provider costs $0.

## Inputs

| Name | Type | Required | Description |
| --- | --- | --- | --- |
| `highlights` | unknown | no | Text snippets Exa's LLM identifies as most relevant from each page. true for defaults, or an object to steer selection with your own query. |
| `ids` | array | no | Document ids obtained from a prior /search call, instead of urls. |
| `livecrawl` | string | no | Deprecated -- use maxAgeHours instead (maxAgeHours: 0 replaces livecrawl: "always"). Kept only because the spec still accepts it; do not send alongside maxAgeHours. |
| `maxAgeHours` | integer | no | Max age of cached content in hours; 0 fetches fresh, -1 always uses cache. Deprecated livecrawl (never/always/fallback/preferred) is superseded by this -- do not send both. |
| `summary` | unknown | no | LLM-generated summary of the webpage, optionally guided by a query or returned structured against a JSON schema. |
| `text` | unknown | no | Full page text for each result. true for defaults with default settings, or an object for advanced control. |
| `urls` | array | yes | URLs to crawl. |

## Output

Shape of the run's `result`, checked against 3 real answers:

```txt
{ results: { id, url, text, image, title, author, entities: { id, type, version, properties: { ... } }[] }[], statuses: { id, error: { tag, httpStatusCode }, source, status }[], requestId, searchTime, costDollars: { total, contents: { text } } }
```

## Run it

Every call needs your API token in `LOOOT_TOKEN`; [Sign in](/get-started/sign-in#use-the-token-in-scripts-and-agents) shows how to get one. Each run also needs a new idempotency key, so a retry never pays twice. With `wait: 30` the answer comes back inline when the run ends within 30 seconds. Otherwise you get the running run back: poll `GET /v1/runs/<runId>`.

Example input with placeholder values:

<CodeGroup>

```bash curl
curl -X POST "https://api.looot.ai/v1/runs" \
  -H "Authorization: Bearer $LOOOT_TOKEN" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: $(uuidgen)" \
  -d '{"endpointId":"exa-contents","input":{"urls":["https://arxiv.org/pdf/2307.06435"]},"wait":30}'
```

```js JavaScript
const response = await fetch("https://api.looot.ai/v1/runs", {
  method: "POST",
  headers: {
    Authorization: `Bearer ${process.env.LOOOT_TOKEN}`,
    "Content-Type": "application/json",
    "Idempotency-Key": crypto.randomUUID(),
  },
  body: JSON.stringify({
    endpointId: "exa-contents",
    input: {
      urls: ["https://arxiv.org/pdf/2307.06435"],
    },
    wait: 30,
  }),
});
const run = await response.json();
console.log(run.status, run.result);
```

```python Python
import os
import uuid

import requests

response = requests.post(
    "https://api.looot.ai/v1/runs",
    headers={
        "Authorization": f"Bearer {os.environ['LOOOT_TOKEN']}",
        "Idempotency-Key": str(uuid.uuid4()),
    },
    json={
        "endpointId": "exa-contents",
        "input": {
            "urls": ["https://arxiv.org/pdf/2307.06435"],
        },
        "wait": 30,
    },
    timeout=90,
)
run = response.json()
print(run["status"], run.get("result"))
```

</CodeGroup>

To let looot pick among every provider of this job instead, send `job:web.article.extract` as `endpointId`; see [the job page](/reference/jobs/web-article-extract).
