Context API / Scrape HTML
Context.dev tool context-dev-web-scrape-html on looot: input fields, $0.0025 per call, output shape, and code to run it with curl, JavaScript or Python.
Scrapes the given URL and returns the raw HTML content of the page. The base request costs 1 credit; requests with browser actions cost 2 credits.
- Tool id:
context-dev-web-scrape-html - Provider: Context.dev
- Job: Get a web page’s HTML (
web.scrape.html) - Price: $0.0025 per call. A call that fails at the provider costs $0.
Inputs
| Name | Type | Required | Description |
|---|---|---|---|
actions |
string | no | Optional browser actions executed in array order after the page loads and before content is captured. Requires a paid plan. Send a JSON array in the query parameter. Maximum: 5 actions. |
country |
string | no | Fetch the target page through a residential proxy in this country (ISO 3166-1 alpha-2). Example: “de” |
excludeSelectors |
string | no | CSS selectors to remove from the result. Applied after includeSelectors. Exclusion takes precedence: an element matching both is removed. Examples: “nav”, “footer”, “.ad-banner”, “[aria-hidden=true]”. |
headers |
string | no | Optional outbound HTTP headers forwarded only to the target URL, sent as deep-object query params such as headers[X-Custom]=value. When provided, caching is bypassed: the result is neither read from nor written to cache. |
includeFrames |
boolean | no | When true, iframes are rendered inline into the returned HTML. |
includeSelectors |
string | no | CSS selectors. When provided, only matching subtrees (and their descendants) are kept and everything else is dropped. When omitted, the entire document is kept. Examples: “article.main”, “#content”, “[role=main]”. |
maxAgeMs |
string | no | Return a cached result if a prior scrape for the same parameters exists and is younger than this many milliseconds. Defaults to 1 day (86400000 ms) when omitted. Max is 30 days (2592000000 ms). Set to 0 to always scrape fresh. |
pdf |
string | no | PDF parsing controls. Use start/end to limit text extraction and embedded-image detection/OCR to an inclusive 1-based page range. |
settleAnimations |
boolean | no | When true, waits briefly for CSS and transition animations to settle before extracting HTML. Defaults to false. This adds a bit of latency in exchange for more stable output on animated pages. |
tags |
array | no | Optional tags for tracking usage. Up to 20 tags, each 1 to 50 characters. |
timeoutMS |
integer | no | Optional timeout in milliseconds for the request. If the request takes longer than this value, it will be aborted with a 408 status code. Maximum allowed value is 300000ms (5 minutes). |
url |
string | yes | Full URL to scrape (must include http:// or https:// protocol). Example: “https://www.example.com” |
useMainContentOnly |
boolean | no | When true, return only the page’s main content in the HTML response, excluding headers, footers, sidebars, and navigation when detectable. |
waitForMs |
string | no | Optional browser wait time in milliseconds after initial page load. Min: 0. Max: 30000 (30 seconds). |
zdr |
string | no | Set to enabled to bypass shared caches and omit request and response content from retained usage logs. Requires zero data retention to be enabled for your organization (contact support@context.dev), otherwise the request fails with ZDR_NOT_ENABLED. Successful ZDR responses include X-Context-ZDR: true. |
Output
Shape of the run’s result, checked against 1 real answer:
{ url, html, type, success, metadata: { title, favicon, finalUrl, headings: { te... ...
The shape is cut here. Signed in, looot inspect context-dev-web-scrape-html prints all of it.
Run it
Every call needs your API token in LOOOT_TOKEN; Sign in shows how to get one. Each run also needs a new idempotency key, so a retry never pays twice. With wait: 30 the answer comes back inline when the run ends within 30 seconds. Otherwise you get the running run back: poll GET /v1/runs/<runId>.
Example input with placeholder values:
curl -X POST "https://api.looot.ai/v1/runs" \
-H "Authorization: Bearer $LOOOT_TOKEN" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: $(uuidgen)" \
-d '{"endpointId":"context-dev-web-scrape-html","input":{"url":"https://www.example.com","tags":["production","team-alpha"],"country":"de"},"wait":30}'const response = await fetch("https://api.looot.ai/v1/runs", {
method: "POST",
headers: {
Authorization: `Bearer ${process.env.LOOOT_TOKEN}`,
"Content-Type": "application/json",
"Idempotency-Key": crypto.randomUUID(),
},
body: JSON.stringify({
endpointId: "context-dev-web-scrape-html",
input: {
url: "https://www.example.com",
tags: ["production", "team-alpha"],
country: "de",
},
wait: 30,
}),
});
const run = await response.json();
console.log(run.status, run.result);import os
import uuid
import requests
response = requests.post(
"https://api.looot.ai/v1/runs",
headers={
"Authorization": f"Bearer {os.environ['LOOOT_TOKEN']}",
"Idempotency-Key": str(uuid.uuid4()),
},
json={
"endpointId": "context-dev-web-scrape-html",
"input": {
"url": "https://www.example.com",
"tags": ["production", "team-alpha"],
"country": "de",
},
"wait": 30,
},
timeout=90,
)
run = response.json()
print(run["status"], run.get("result"))To let looot pick among every provider of this job instead, send job:web.scrape.html as endpointId; see the job page.