Skip to content
looot docs
Esc
↑↓navigate↵open⌘Jpreview
On this page

TinyFish Fetch API / Fetch and extract content from URLs

TinyFish Search API tool tinyfish-fetch on looot: input fields, $0 per call, output shape, and code to run it with curl, JavaScript or Python.

Fetches web pages, renders JavaScript-heavy pages when needed, and returns clean extracted content in your preferred format. Submit up to 10 URLs, get back structured content. Per-URL failures appear in errors[] and do not fail the entire request. Per-URL error codes (in errors[].error): - target_http_error, target server returned a non-2xx HTTP status other than 404/410; the raw status code is in errors[].status - page_not_found, target URL returned HTTP 404 or 410; the raw status code is in errors[].status - target_unreachable, connection refused, TLS failure, DNS failure, or other network error - timeout, request timed out - proxy_error, proxy tunnel failure - bot_blocked, bot-challenge page detected (Cloudflare, etc.) - empty_content, page loaded but no extractable text was found - login_required, the target redirected to a login or account wall; retrying without credentials will not help - content_too_large, the document exceeds the 20MB size limit for generic documents (PDF and CSV have their own 50MB budgets); not retried in the browser - invalid_url, malformed URL or SSRF-blocked address - invalid_redirect_url, redirect target rejected before fetch - conditional_unsupported, conditional requests (if_none_match / if_modified_since) are supported on the fast path only; this URL requires browser rendering - selector_not_matched, no elements matching any include_selectors entry remained after exclude_selectors was applied; the error carries unmatched_selectors plus candidate_selectors retry hints (a partial miss is not an error, it’s reported on the result’s unmatched_selectors) - selector_unsupported, include_selectors / exclude_selectors sent for a URL that resolves to a direct PDF/CSV download (no HTML to scope)

  • Tool id: tinyfish-fetch
  • Provider: TinyFish Search API
  • Job: Post fetch (fetch.post)
  • Price: $0 per call. A call that fails at the provider costs $0.

Inputs

Name Type Required Description
exclude_selectors array no Array of CSS selectors (1-20 entries, each 1-1000 characters) for elements to remove before extraction, applied before include_selectors scopes what remains, so it also prunes inside selected regions. Entries may themselves use CSS comma-grouping. Entries that match nothing are a no-op, never an error, but URLs that resolve to direct PDF/CSV downloads fail with selector_unsupported. Invalid CSS selector syntax is rejected with a 422. Applied post-fetch: caching and routing are unchanged.
format string no Output format for extracted content. “markdown” (default) is ideal for LLM consumption. “html” returns cleaned semantic HTML. “json” returns a structured document tree. Example: “markdown”
if_modified_since string no Last-Modified validator from a prior fetch of this URL, forwarded verbatim as the If-Modified-Since header on the origin request. Only valid with a single URL, combining with a batch of URLs returns a 400. tf-fetch does not persist validators; the caller owns replaying them. Example: “Wed, 21 Oct 2015 07:28 GMT”
if_none_match string no ETag validator from a prior fetch of this URL, forwarded verbatim as the If-None-Match header on the origin request. Only valid with a single URL, combining with a batch of URLs returns a 400. tf-fetch does not persist validators; the caller owns replaying them. Example: “W/"abc123"”
image_links boolean no Extract all image URLs (<img src>) from each page. Useful for finding visual content or media assets. Image links are returned as absolute URLs in the image_links array of each result. Example: false
include_etag_and_last_modified boolean no Opt-in to receiving etag / last_modified validators (and not_modified detection) on each result. Defaults to false, tf-fetch omits these fields unless requested. Independent of if_none_match / if_modified_since: works with a single URL or a batch. Example: true
include_selectors array no Array of CSS selectors (1-20 entries, each 1-1000 characters) that scope extracted content (text, links, image_links) to elements matching ANY entry, concatenated in document order. Tag selectors cover semantic sections (main, article, nav); entries may themselves use CSS comma-grouping. Selected content is returned verbatim in the requested format (scripts/styles stripped), automatic boilerplate removal is bypassed. Page-level metadata (title, description, language, `autho…
links boolean no Extract all outbound links (<a href>) from each page. Useful for discovering related pages or navigating to specific content. Links are returned as absolute URLs in the links array of each result. Example: false
page_metadata boolean no Return page-head metadata for each page in the page_metadata object of each result: canonical URL, favicon, robots directive, generator, viewport, keywords, all Open Graph (og), Twitter card (twitter), and article tags, and remaining named meta tags under other. Useful for SEO/technical audits and link-preview generation. Example: false
per_url_timeout_ms integer no Wall-clock timeout budget in milliseconds applied independently to each URL. If one URL exceeds this budget, it returns a per-URL timeout error while other URLs in the same request continue. Example: 45000
purpose string no Why these URLs are being fetched, the underlying goal or task the content will be used for. Used to better tailor fetching and extraction to your intent. Example: “Compare pricing tiers across vendors for a procurement report”
ttl integer no Caller freshness tolerance in seconds for the cached entry. Omit (default) for unlimited tolerance, any cached entry is acceptable. Set to 0 to prefer a live fetch; a cached entry is still served if the origin’s Cache-Control: max-age covers its age, or the host is in the small allowlist of operator-pinned never-expire domains. Set to N > 0 to accept a cached entry whose age is below N; the upstream Cache-Control: max-age and the never-expire allowlist may extend (never shorten) t… Example: 0
urls array yes Array of URLs to fetch (1-10). All URLs are fetched in parallel. Each URL is processed independently, if one fails, others still return successfully. Errors are reported per-URL in the errors array.

Output

Shape of the run’s result, checked against 2 real answers:

{ errors: unknown[], results: { url, text, title, author, format, language, fina... ...

The shape is cut here. Signed in, looot inspect tinyfish-fetch prints all of it.

Run it

Every call needs your API token in LOOOT_TOKEN; Sign in shows how to get one. Each run also needs a new idempotency key, so a retry never pays twice. With wait: 30 the answer comes back inline when the run ends within 30 seconds. Otherwise you get the running run back: poll GET /v1/runs/<runId>.

Example input with placeholder values:

curl -X POST "https://api.looot.ai/v1/runs" \
  -H "Authorization: Bearer $LOOOT_TOKEN" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: $(uuidgen)" \
  -d '{"endpointId":"tinyfish-fetch","input":{"urls":["https://example.com","https://example.org"],"ttl":0,"links":false,"format":"markdown","purpose":"Compare pricing tiers across vendors for a procurement report","image_links":false,"if_none_match":"W/\"abc123\"","page_metadata":false,"exclude_selectors":[".comments",".newsletter-signup"],"if_modified_since":"Wed, 21 Oct 2015 07:28:00 GMT","include_selectors":["article"],"per_url_timeout_ms":45000,"include_etag_and_last_modified":true},"wait":30}'
const response = await fetch("https://api.looot.ai/v1/runs", {
  method: "POST",
  headers: {
    Authorization: `Bearer ${process.env.LOOOT_TOKEN}`,
    "Content-Type": "application/json",
    "Idempotency-Key": crypto.randomUUID(),
  },
  body: JSON.stringify({
    endpointId: "tinyfish-fetch",
    input: {
      urls: ["https://example.com", "https://example.org"],
      ttl: 0,
      links: false,
      format: "markdown",
      purpose: "Compare pricing tiers across vendors for a procurement report",
      image_links: false,
      if_none_match: "W/\"abc123\"",
      page_metadata: false,
      exclude_selectors: [".comments", ".newsletter-signup"],
      if_modified_since: "Wed, 21 Oct 2015 07:28:00 GMT",
      include_selectors: ["article"],
      per_url_timeout_ms: 45000,
      include_etag_and_last_modified: true,
    },
    wait: 30,
  }),
});
const run = await response.json();
console.log(run.status, run.result);
import os
import uuid

import requests

response = requests.post(
    "https://api.looot.ai/v1/runs",
    headers={
        "Authorization": f"Bearer {os.environ['LOOOT_TOKEN']}",
        "Idempotency-Key": str(uuid.uuid4()),
    },
    json={
        "endpointId": "tinyfish-fetch",
        "input": {
            "urls": ["https://example.com", "https://example.org"],
            "ttl": 0,
            "links": False,
            "format": "markdown",
            "purpose": "Compare pricing tiers across vendors for a procurement report",
            "image_links": False,
            "if_none_match": "W/\"abc123\"",
            "page_metadata": False,
            "exclude_selectors": [".comments", ".newsletter-signup"],
            "if_modified_since": "Wed, 21 Oct 2015 07:28:00 GMT",
            "include_selectors": ["article"],
            "per_url_timeout_ms": 45000,
            "include_etag_and_last_modified": True,
        },
        "wait": 30,
    },
    timeout=90,
)
run = response.json()
print(run["status"], run.get("result"))

To let looot pick among every provider of this job instead, send job:fetch.post as endpointId; see the job page.

Was this page helpful?