Skip to content
looot docs
Esc
↑↓navigate↵open⌘Jpreview
On this page

Context.dev / Scrape a web page to markdown

Context.dev tool context-dev-web-scrape-markdown on looot: input fields, $0.0025 per call, output shape, and code to run it with curl, JavaScript or Python.

Scrapes one URL into markdown. Send url with http:// or https://. Options keep links or images, keep only the main content, wait for the page, or fetch it from a given country. YouTube URLs return the video’s title, channel, description and transcript. Returns markdown, url, contentLength and metadata (title, description, language, canonical URL and more).

Inputs

Name Type Required Description
actions string no Optional browser actions executed in array order after the page loads and before content is captured. Requires a paid plan. Send a JSON array in the query parameter. Maximum: 5 actions.
country string no Fetch the target page through a residential proxy in this country (ISO 3166-1 alpha-2). Example: “de”
excludeSelectors string no CSS selectors to remove before conversion to Markdown. Applied after includeSelectors. Exclusion takes precedence: an element matching both is removed. Examples: “nav”, “footer”, “.ad-banner”, “[aria-hidden=true]”.
headers string no Optional outbound HTTP headers forwarded only to the target URL, sent as deep-object query params such as headers[X-Custom]=value. When provided, caching is bypassed: the result is neither read from nor written to cache.
includeFrames boolean no When true, the contents of iframes are rendered to Markdown.
includeHTML boolean no When true, the response also includes an html field with the page HTML the Markdown was converted from, the same body the Scrape HTML endpoint returns for the equivalent request.
includeImages boolean no Include image references in Markdown output
includeLinks boolean no Preserve hyperlinks in Markdown output
includeSelectors string no CSS selectors. When provided, only matching HTML subtrees (and their descendants) are kept before conversion to Markdown. When omitted, the entire document is kept. Examples: “article.main”, “#content”, “[role=main]”.
maxAgeMs string no Return a cached result if a prior scrape for the same parameters exists and is younger than this many milliseconds. Defaults to 1 day (86400000 ms) when omitted. Max is 30 days (2592000000 ms). Set to 0 to always scrape fresh.
pdf string no PDF parsing controls. Use start/end to limit text extraction and embedded-image detection/OCR to an inclusive 1-based page range.
settleAnimations boolean no When true, waits briefly for CSS and transition animations to settle before converting to Markdown. Defaults to false. This adds a bit of latency in exchange for more stable output on animated pages.
shortenBase64Images boolean no Shorten base64-encoded image data in the Markdown output
tags array no Optional tags for tracking usage. Up to 20 tags, each 1 to 50 characters.
timeoutMS integer no Optional timeout in milliseconds for the request. If the request takes longer than this value, it will be aborted with a 408 status code. Maximum allowed value is 300000ms (5 minutes).
url string yes Full URL to scrape into LLM usable Markdown (must include http:// or https:// protocol). Example: “https://www.example.com”
useMainContentOnly boolean no Extract only the main content of the page, excluding headers, footers, sidebars, and navigation
waitForMs string no Optional browser wait time in milliseconds after initial page load before converting the page to Markdown. Min: 0. Max: 30000 (30 seconds).
zdr string no Set to enabled to bypass shared caches and omit request and response content from retained usage logs. Requires zero data retention to be enabled for your organization (contact support@context.dev), otherwise the request fails with ZDR_NOT_ENABLED. Successful ZDR responses include X-Context-ZDR: true.

Output

Shape of the run’s result, checked against 4 real answers:

{ url, success, markdown, metadata: { image, title, author, jsonLd: { @type, @gr... ...

The shape is cut here. Signed in, looot inspect context-dev-web-scrape-markdown prints all of it.

Run it

Every call needs your API token in LOOOT_TOKEN; Sign in shows how to get one. Each run also needs a new idempotency key, so a retry never pays twice. With wait: 30 the answer comes back inline when the run ends within 30 seconds. Otherwise you get the running run back: poll GET /v1/runs/<runId>.

Example input with placeholder values:

curl -X POST "https://api.looot.ai/v1/runs" \
  -H "Authorization: Bearer $LOOOT_TOKEN" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: $(uuidgen)" \
  -d '{"endpointId":"context-dev-web-scrape-markdown","input":{"url":"https://www.example.com","tags":["production","team-alpha"],"country":"de"},"wait":30}'
const response = await fetch("https://api.looot.ai/v1/runs", {
  method: "POST",
  headers: {
    Authorization: `Bearer ${process.env.LOOOT_TOKEN}`,
    "Content-Type": "application/json",
    "Idempotency-Key": crypto.randomUUID(),
  },
  body: JSON.stringify({
    endpointId: "context-dev-web-scrape-markdown",
    input: {
      url: "https://www.example.com",
      tags: ["production", "team-alpha"],
      country: "de",
    },
    wait: 30,
  }),
});
const run = await response.json();
console.log(run.status, run.result);
import os
import uuid

import requests

response = requests.post(
    "https://api.looot.ai/v1/runs",
    headers={
        "Authorization": f"Bearer {os.environ['LOOOT_TOKEN']}",
        "Idempotency-Key": str(uuid.uuid4()),
    },
    json={
        "endpointId": "context-dev-web-scrape-markdown",
        "input": {
            "url": "https://www.example.com",
            "tags": ["production", "team-alpha"],
            "country": "de",
        },
        "wait": 30,
    },
    timeout=90,
)
run = response.json()
print(run["status"], run.get("result"))

To let looot pick among every provider of this job instead, send job:web.scrape.markdown as endpointId; see the job page.

Was this page helpful?