Skip to content
looot docs
Esc
↑↓navigate↵open⌘Jpreview
On this page

Post fetch

Run job:fetch.post through the looot API: 1 provider, from $0 per call. Inputs, prices, output shape and code for curl, JavaScript and Python.

Post fetch.

Job id fetch.post, in Scrapers (Fetch). As of 2026-10-01, 1 provider serve it through 1 tool. Run job:fetch.post and looot picks one of them; with fallback on, a miss moves on to the next. See Jobs.

Inputs

This job has one tool, so it takes that tool’s fields. Any key you send passes through unchanged.

Name Type Required Description
exclude_selectors array no Array of CSS selectors (1-20 entries, each 1-1000 characters) for elements to remove before extraction, applied before include_selectors scopes what remains, so it also prunes inside selected regions. Entries may themselves use CSS comma-grouping. Entries that match nothing are a no-op, never an error, but URLs that resolve to direct PDF/CSV downloads fail with selector_unsupported. Invalid CSS selector syntax is rejected with a 422. Applied post-fetch: caching and routing are unchanged.
format string no Output format for extracted content. “markdown” (default) is ideal for LLM consumption. “html” returns cleaned semantic HTML. “json” returns a structured document tree. Example: “markdown”
if_modified_since string no Last-Modified validator from a prior fetch of this URL, forwarded verbatim as the If-Modified-Since header on the origin request. Only valid with a single URL, combining with a batch of URLs returns a 400. tf-fetch does not persist validators; the caller owns replaying them. Example: “Wed, 21 Oct 2015 07:28 GMT”
if_none_match string no ETag validator from a prior fetch of this URL, forwarded verbatim as the If-None-Match header on the origin request. Only valid with a single URL, combining with a batch of URLs returns a 400. tf-fetch does not persist validators; the caller owns replaying them. Example: “W/"abc123"”
image_links boolean no Extract all image URLs (<img src>) from each page. Useful for finding visual content or media assets. Image links are returned as absolute URLs in the image_links array of each result. Example: false
include_etag_and_last_modified boolean no Opt-in to receiving etag / last_modified validators (and not_modified detection) on each result. Defaults to false, tf-fetch omits these fields unless requested. Independent of if_none_match / if_modified_since: works with a single URL or a batch. Example: true
include_selectors array no Array of CSS selectors (1-20 entries, each 1-1000 characters) that scope extracted content (text, links, image_links) to elements matching ANY entry, concatenated in document order. Tag selectors cover semantic sections (main, article, nav); entries may themselves use CSS comma-grouping. Selected content is returned verbatim in the requested format (scripts/styles stripped), automatic boilerplate removal is bypassed. Page-level metadata (title, description, language, `autho…
links boolean no Extract all outbound links (<a href>) from each page. Useful for discovering related pages or navigating to specific content. Links are returned as absolute URLs in the links array of each result. Example: false
page_metadata boolean no Return page-head metadata for each page in the page_metadata object of each result: canonical URL, favicon, robots directive, generator, viewport, keywords, all Open Graph (og), Twitter card (twitter), and article tags, and remaining named meta tags under other. Useful for SEO/technical audits and link-preview generation. Example: false
per_url_timeout_ms integer no Wall-clock timeout budget in milliseconds applied independently to each URL. If one URL exceeds this budget, it returns a per-URL timeout error while other URLs in the same request continue. Example: 45000
purpose string no Why these URLs are being fetched, the underlying goal or task the content will be used for. Used to better tailor fetching and extraction to your intent. Example: “Compare pricing tiers across vendors for a procurement report”
ttl integer no Caller freshness tolerance in seconds for the cached entry. Omit (default) for unlimited tolerance, any cached entry is acceptable. Set to 0 to prefer a live fetch; a cached entry is still served if the origin’s Cache-Control: max-age covers its age, or the host is in the small allowlist of operator-pinned never-expire domains. Set to N > 0 to accept a cached entry whose age is below N; the upstream Cache-Control: max-age and the never-expire allowlist may extend (never shorten) t… Example: 0
urls array yes Array of URLs to fetch (1-10). All URLs are fetched in parallel. Each URL is processed independently, if one fails, others still return successfully. Errors are reported per-URL in the errors array.

Providers and prices

Provider Tool Price
TinyFish Search API tinyfish-fetch $0 per call

A call that fails at the provider costs $0. See What’s free.

Output

Shape of the run’s result, checked against 2 real answers:

{ errors: unknown[], results: { url, text, title, author, format, language, fina... ...

The shape is cut here. Signed in, looot inspect tinyfish-fetch prints all of it.

Run it

Every call needs your API token in LOOOT_TOKEN; Sign in shows how to get one. Each run also needs a new idempotency key, so a retry never pays twice. With wait: 30 the answer comes back inline when the run ends within 30 seconds. Otherwise you get the running run back: poll GET /v1/runs/<runId>.

Example input with placeholder values:

curl -X POST "https://api.looot.ai/v1/runs" \
  -H "Authorization: Bearer $LOOOT_TOKEN" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: $(uuidgen)" \
  -d '{"endpointId":"job:fetch.post","input":{"urls":["https://example.com","https://example.org"],"ttl":0,"links":false,"format":"markdown","purpose":"Compare pricing tiers across vendors for a procurement report","image_links":false,"if_none_match":"W/\"abc123\"","page_metadata":false,"exclude_selectors":[".comments",".newsletter-signup"],"if_modified_since":"Wed, 21 Oct 2015 07:28:00 GMT","include_selectors":["article"],"per_url_timeout_ms":45000,"include_etag_and_last_modified":true},"wait":30}'
const response = await fetch("https://api.looot.ai/v1/runs", {
  method: "POST",
  headers: {
    Authorization: `Bearer ${process.env.LOOOT_TOKEN}`,
    "Content-Type": "application/json",
    "Idempotency-Key": crypto.randomUUID(),
  },
  body: JSON.stringify({
    endpointId: "job:fetch.post",
    input: {
      urls: ["https://example.com", "https://example.org"],
      ttl: 0,
      links: false,
      format: "markdown",
      purpose: "Compare pricing tiers across vendors for a procurement report",
      image_links: false,
      if_none_match: "W/\"abc123\"",
      page_metadata: false,
      exclude_selectors: [".comments", ".newsletter-signup"],
      if_modified_since: "Wed, 21 Oct 2015 07:28:00 GMT",
      include_selectors: ["article"],
      per_url_timeout_ms: 45000,
      include_etag_and_last_modified: true,
    },
    wait: 30,
  }),
});
const run = await response.json();
console.log(run.status, run.result);
import os
import uuid

import requests

response = requests.post(
    "https://api.looot.ai/v1/runs",
    headers={
        "Authorization": f"Bearer {os.environ['LOOOT_TOKEN']}",
        "Idempotency-Key": str(uuid.uuid4()),
    },
    json={
        "endpointId": "job:fetch.post",
        "input": {
            "urls": ["https://example.com", "https://example.org"],
            "ttl": 0,
            "links": False,
            "format": "markdown",
            "purpose": "Compare pricing tiers across vendors for a procurement report",
            "image_links": False,
            "if_none_match": "W/\"abc123\"",
            "page_metadata": False,
            "exclude_selectors": [".comments", ".newsletter-signup"],
            "if_modified_since": "Wed, 21 Oct 2015 07:28:00 GMT",
            "include_selectors": ["article"],
            "per_url_timeout_ms": 45000,
            "include_etag_and_last_modified": True,
        },
        "wait": 30,
    },
    timeout=90,
)
run = response.json()
print(run["status"], run.get("result"))

To pin one provider, send its tool id as endpointId instead of job:fetch.post.

Was this page helpful?