---
title: Turn a web page into markdown
description: Scrape a URL into clean markdown or structured JSON, read the title and body, and tell a blocked or sign-in page apart from the page's content.
---

`web.scrape.markdown` reads a page and gives back markdown; `web.scrape.structured` gives back
JSON instead. Both take a `url`.

## What it costs

```json MCP
search {"filters": {"capability": "web.scrape.markdown"}, "prefer": "cheapest"}
```

On 2026-09-28 the cheapest listed `web.scrape.markdown` provider priced around $0.0002 per call,
and `web.scrape.structured` around $0.0012 per call. Read the live number before a batch of pages.

1. **Scrape to markdown**

    <CodeGroup>

    ```json MCP
    run {"endpointId": "job:web.scrape.markdown", "input": {"url": "https://example.com/pricing"}, "idempotencyKey": "scrape-example-pricing", "fallback": {"maxAttempts": 3, "maxCostUsd": 0.05}}
    ```

    ```bash REST
    curl -s "https://api.looot.ai/v1/runs?wait=20" \
      -H "Authorization: Bearer $LOOOT_TOKEN" \
      -H "Content-Type: application/json" \
      -d '{
        "endpointId": "job:web.scrape.markdown",
        "input": {"url": "https://example.com/pricing"},
        "idempotencyKey": "scrape-example-pricing",
        "fallback": {"maxAttempts": 3, "maxCostUsd": 0.05}
      }'
    ```

    ```bash CLI
    looot run job:web.scrape.markdown \
      --input '{"url": "https://example.com/pricing"}' \
      --wait
    ```

    </CodeGroup>

    :::note
    looot 1.1.0 does not accept `--fallback` on `looot run`. The CLI call above tries one provider.
    :::

2. **Read the page**

    Read `normalized.fields.markdown` and `normalized.fields.title`. A blocked page, an empty page or a
    sign-in wall is a `miss`, not an error: with [`fallback`](/concepts/fallback), the run moves on to the next scraper.
    It never returns a login page as if it were content.

3. **Get structured JSON**

    `web.scrape.structured` takes the same `url` and returns JSON. It has no
    [`normalized`](/concepts/normalized-output) map yet, so read the shape from `result` and check [`inspect`](/mcp-tools/inspect) on the endpoint for its
    `outputSchema` first if you need a specific field reliably.

    ```json MCP
    run {"endpointId": "job:web.scrape.structured", "input": {"url": "https://example.com/pricing"}, "idempotencyKey": "scrape-example-structured", "fallback": true}
    ```

## What the answer looks like

| Field | What it is |
| --- | --- |
| `status` | `completed`, `failed`, `queued`, `running` |
| [`outcome`](/concepts/outcomes) | `hit`, `weak`, `miss`, `error`, `rejected`, `skipped` or `pending` |
| [`servedProviderId`](/concepts/served-provider) | the provider that answered |
| `actualCost` | what this run charged |
| `normalized.fields` | `web.scrape.markdown` only: `markdown`, `title` |
| `route.summary` | fallback runs only |

## On a miss or error

- Blocked, empty or a sign-in page: `outcome` is `miss`. With `fallback`, the next scraper in the
  route gets a try inside the same hold.
- A URL that does not resolve, or the domain refuses every scraper: report the miss. Don't
  run the same URL again with a new key; a different provider is unlikely to reach a page
  nothing else could load.
- [`insufficient_balance`](/errors/rest-errors#insufficient_balance): the run is blocked and not charged. Use [`top_up`](/mcp-tools/top-up), then retry with a new
  [`idempotencyKey`](/concepts/idempotency).

<Related />
