Any Site API / /webparser/parse
Anysite tool anysite-api-webparser-parse on looot: input fields, $0.0036 per call, output shape, and code to run it with curl, JavaScript or Python.
Parse and clean HTML from web page Price: 1 credit
- Tool id:
anysite-api-webparser-parse - Provider: Anysite
- Job: Parse api webparser (
api.webparser.parse) - Price: $0.0036 per call. A call that fails at the provider costs $0.
Inputs
| Name | Type | Required | Description |
|---|---|---|---|
exclude_tags |
array | no | |
extract_contacts |
boolean | no | Extract links, emails, and phone numbers from the page |
extract_minimal |
boolean | no | Use minimal extraction (only links, title, emails, phones if set) |
include_tags |
array | no | |
min_text_block |
integer | no | Minimum text block size for main content detection (in characters) |
only_main_content |
boolean | no | Extract only main content of the page (heuristic algorithm) |
remove_base64_images |
boolean | no | Remove base64-encoded images (reduces output size) |
remove_comments |
boolean | no | Remove HTML comments |
resolve_srcset |
boolean | no | Convert image srcset to src (selects the largest image) |
return_full_html |
boolean | no | Return full HTML document (True) or only body content (False) |
same_origin_links |
boolean | no | Only extract links from the same domain (used with extract_contacts) |
social_links_only |
boolean | no | Only extract social media links (LinkedIn, Twitter/X, Facebook, Instagram, etc.) |
strip_all_tags |
boolean | no | Remove all HTML tags and return plain text only |
timeout |
integer | no | Max scrapping execution timeout (in seconds) |
url |
string | yes | URL of the page to parse. Example: “https://www.example.com” |
x-request-id |
string | no |
Output
The run’s result holds the provider’s answer. Signed in, looot inspect anysite-api-webparser-parse prints its fields.
Run it
Every call needs your API token in LOOOT_TOKEN; Sign in shows how to get one. Each run also needs a new idempotency key, so a retry never pays twice. With wait: 30 the answer comes back inline when the run ends within 30 seconds. Otherwise you get the running run back: poll GET /v1/runs/<runId>.
Example input with placeholder values:
curl -X POST "https://api.looot.ai/v1/runs" \
-H "Authorization: Bearer $LOOOT_TOKEN" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: $(uuidgen)" \
-d '{"endpointId":"anysite-api-webparser-parse","input":{"url":"https://www.example.com"},"wait":30}'const response = await fetch("https://api.looot.ai/v1/runs", {
method: "POST",
headers: {
Authorization: `Bearer ${process.env.LOOOT_TOKEN}`,
"Content-Type": "application/json",
"Idempotency-Key": crypto.randomUUID(),
},
body: JSON.stringify({
endpointId: "anysite-api-webparser-parse",
input: {
url: "https://www.example.com",
},
wait: 30,
}),
});
const run = await response.json();
console.log(run.status, run.result);import os
import uuid
import requests
response = requests.post(
"https://api.looot.ai/v1/runs",
headers={
"Authorization": f"Bearer {os.environ['LOOOT_TOKEN']}",
"Idempotency-Key": str(uuid.uuid4()),
},
json={
"endpointId": "anysite-api-webparser-parse",
"input": {
"url": "https://www.example.com",
},
"wait": 30,
},
timeout=90,
)
run = response.json()
print(run["status"], run.get("result"))To let looot pick among every provider of this job instead, send job:api.webparser.parse as endpointId; see the job page.