Skip to content
looot docs
Esc
↑↓navigate↵open⌘Jpreview
On this page

Any Site API / /webparser/parse

Anysite tool anysite-api-webparser-parse on looot: input fields, $0.0036 per call, output shape, and code to run it with curl, JavaScript or Python.

Parse and clean HTML from web page Price: 1 credit

  • Tool id: anysite-api-webparser-parse
  • Provider: Anysite
  • Job: Parse api webparser (api.webparser.parse)
  • Price: $0.0036 per call. A call that fails at the provider costs $0.

Inputs

Name Type Required Description
exclude_tags array no
extract_contacts boolean no Extract links, emails, and phone numbers from the page
extract_minimal boolean no Use minimal extraction (only links, title, emails, phones if set)
include_tags array no
min_text_block integer no Minimum text block size for main content detection (in characters)
only_main_content boolean no Extract only main content of the page (heuristic algorithm)
remove_base64_images boolean no Remove base64-encoded images (reduces output size)
remove_comments boolean no Remove HTML comments
resolve_srcset boolean no Convert image srcset to src (selects the largest image)
return_full_html boolean no Return full HTML document (True) or only body content (False)
same_origin_links boolean no Only extract links from the same domain (used with extract_contacts)
social_links_only boolean no Only extract social media links (LinkedIn, Twitter/X, Facebook, Instagram, etc.)
strip_all_tags boolean no Remove all HTML tags and return plain text only
timeout integer no Max scrapping execution timeout (in seconds)
url string yes URL of the page to parse. Example: “https://www.example.com”
x-request-id string no

Output

The run’s result holds the provider’s answer. Signed in, looot inspect anysite-api-webparser-parse prints its fields.

Run it

Every call needs your API token in LOOOT_TOKEN; Sign in shows how to get one. Each run also needs a new idempotency key, so a retry never pays twice. With wait: 30 the answer comes back inline when the run ends within 30 seconds. Otherwise you get the running run back: poll GET /v1/runs/<runId>.

Example input with placeholder values:

curl -X POST "https://api.looot.ai/v1/runs" \
  -H "Authorization: Bearer $LOOOT_TOKEN" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: $(uuidgen)" \
  -d '{"endpointId":"anysite-api-webparser-parse","input":{"url":"https://www.example.com"},"wait":30}'
const response = await fetch("https://api.looot.ai/v1/runs", {
  method: "POST",
  headers: {
    Authorization: `Bearer ${process.env.LOOOT_TOKEN}`,
    "Content-Type": "application/json",
    "Idempotency-Key": crypto.randomUUID(),
  },
  body: JSON.stringify({
    endpointId: "anysite-api-webparser-parse",
    input: {
      url: "https://www.example.com",
    },
    wait: 30,
  }),
});
const run = await response.json();
console.log(run.status, run.result);
import os
import uuid

import requests

response = requests.post(
    "https://api.looot.ai/v1/runs",
    headers={
        "Authorization": f"Bearer {os.environ['LOOOT_TOKEN']}",
        "Idempotency-Key": str(uuid.uuid4()),
    },
    json={
        "endpointId": "anysite-api-webparser-parse",
        "input": {
            "url": "https://www.example.com",
        },
        "wait": 30,
    },
    timeout=90,
)
run = response.json()
print(run["status"], run.get("result"))

To let looot pick among every provider of this job instead, send job:api.webparser.parse as endpointId; see the job page.

Was this page helpful?