Firecrawl API / Crawl multiple URLs based on options
Firecrawl tool firecrawl-crawl on looot: input fields, $0.005 per result, output shape, and code to run it with curl, JavaScript or Python.
Crawls a site from a start URL and returns its pages. limit sets the page count, up to 100; set it on every call, since Firecrawl’s own default is much higher. maxDiscoveryDepth is capped at 5 and maxConcurrency at 10. Priced per page crawled, from Firecrawl’s reported credits.
- Tool id:
firecrawl-crawl - Provider: Firecrawl
- Job: Submit web crawl (
web.crawl.submit) - Price: $0.005 per result. A call that fails at the provider costs $0.
Inputs
| Name | Type | Required | Description |
|---|---|---|---|
allowExternalLinks |
boolean | no | Allows the crawler to follow links to external websites. |
allowSubdomains |
boolean | no | Allows the crawler to follow links to subdomains of the main domain. |
crawlEntireDomain |
boolean | no | Allows the crawler to follow internal links to sibling or parent URLs, not just child paths. false: Only crawls deeper (child) URLs. → e.g. /features/feature-1 → /features/feature-1/tips ✅ → Won’t follow /pricing or / ❌ true: Crawls any internal links, including siblings and parents. → e.g. /features/feature-1 → /pricing, /, etc. ✅ Use true for broader internal coverage beyond nested paths. |
delay |
number | no | Delay in seconds between scrapes. This helps respect website rate limits. |
excludePaths |
array | no | URL pathname regex patterns that exclude matching URLs from the crawl. For example, if you set “excludePaths”: [“blog/.*”] for the base URL firecrawl.dev, any results matching that pattern will be excluded, such as https://www.firecrawl.dev/blog/firecrawl-launch-week-1-recap. |
ignoreQueryParameters |
boolean | no | Do not re-scrape the same path with different (or none) query parameters |
includePaths |
array | no | URL pathname regex patterns that include matching URLs in the crawl. Only the paths that match the specified patterns will be included in the response. For example, if you set “includePaths”: [“blog/.*”] for the base URL firecrawl.dev, only results matching that pattern will be included, such as https://www.firecrawl.dev/blog/firecrawl-launch-week-1-recap. |
limit |
integer | no | Maximum number of pages to crawl. Default limit is 10000. Example: 10 |
maxConcurrency |
integer | no | Maximum number of concurrent scrapes. This parameter allows you to set a concurrency limit for this crawl. If not specified, the crawl adheres to your team’s concurrency limit. Example: 10 |
maxDiscoveryDepth |
integer | no | Maximum depth to crawl based on discovery order. The root site and sitemapped pages has a discovery depth of 0. For example, if you set it to 1, and you set ignoreSitemap, you will only crawl the entered URL and all URLs that are linked on that page. Example: 2 |
prompt |
string | no | A prompt to use to generate the crawler options (all the parameters below) from natural language. Explicitly set parameters will override the generated equivalents. |
scrapeOptions |
string | no | |
sitemap |
string | no | Sitemap mode when crawling. If you set it to ‘skip’, the crawler will ignore the website sitemap and only crawl the entered URL and discover pages from there onwards. |
url |
string | yes | The base URL to start crawling from |
webhook |
string | no | A webhook specification object. |
zeroDataRetention |
boolean | no | If true, this will enable zero data retention for this crawl. To enable this feature, please contact help@firecrawl.dev |
Output
Shape of the run’s result, checked against 2 real answers:
{ data: { markdown, metadata: { url, title, favicon, cachedAt, language, scrapeI... ...
The shape is cut here. Signed in, looot inspect firecrawl-crawl prints all of it.
Run it
Every call needs your API token in LOOOT_TOKEN; Sign in shows how to get one. Each run also needs a new idempotency key, so a retry never pays twice. With wait: 30 the answer comes back inline when the run ends within 30 seconds. Otherwise you get the running run back: poll GET /v1/runs/<runId>.
Example input with placeholder values:
curl -X POST "https://api.looot.ai/v1/runs" \
-H "Authorization: Bearer $LOOOT_TOKEN" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: $(uuidgen)" \
-d '{"endpointId":"firecrawl-crawl","input":{"url":"https://example.com","limit":10,"maxConcurrency":10,"maxDiscoveryDepth":2},"wait":30}'const response = await fetch("https://api.looot.ai/v1/runs", {
method: "POST",
headers: {
Authorization: `Bearer ${process.env.LOOOT_TOKEN}`,
"Content-Type": "application/json",
"Idempotency-Key": crypto.randomUUID(),
},
body: JSON.stringify({
endpointId: "firecrawl-crawl",
input: {
url: "https://example.com",
limit: 10,
maxConcurrency: 10,
maxDiscoveryDepth: 2,
},
wait: 30,
}),
});
const run = await response.json();
console.log(run.status, run.result);import os
import uuid
import requests
response = requests.post(
"https://api.looot.ai/v1/runs",
headers={
"Authorization": f"Bearer {os.environ['LOOOT_TOKEN']}",
"Idempotency-Key": str(uuid.uuid4()),
},
json={
"endpointId": "firecrawl-crawl",
"input": {
"url": "https://example.com",
"limit": 10,
"maxConcurrency": 10,
"maxDiscoveryDepth": 2,
},
"wait": 30,
},
timeout=90,
)
run = response.json()
print(run["status"], run.get("result"))To let looot pick among every provider of this job instead, send job:web.crawl.submit as endpointId; see the job page.