Firecrawl API / Scrape multiple URLs and optionally extract information using an LLM
Firecrawl tool firecrawl-batch-scrape on looot: input fields, $0.005 per result, output shape, and code to run it with curl, JavaScript or Python.
Scrapes several URLs in one job and can extract information with an LLM. Send urls (up to 100); maxConcurrency is capped at 10. Priced per page scraped, from Firecrawl’s reported credits.
- Tool id:
firecrawl-batch-scrape - Provider: Firecrawl
- Job: Submit web scrape bulk (
web.scrape.bulk.submit) - Price: $0.005 per result. A call that fails at the provider costs $0.
Inputs
| Name | Type | Required | Description |
|---|---|---|---|
actions |
array | no | Actions to perform on the page before grabbing the content |
blockAds |
boolean | no | Enables ad-blocking and cookie popup blocking. |
excludeTags |
array | no | Tags to exclude from the output. |
formats |
array | no | Output formats to include in the response. You can specify one or more formats, either as strings (e.g., 'markdown') or as objects with additional options (e.g., \{ type: 'json', schema: \{...\} \}). Some formats require specific options to be set. Example: ['markdown', \{ type: 'json', schema: \{...\} \}]. |
headers |
string | no | Headers to send with the request. Can be used to send cookies, user-agent, etc. |
ignoreInvalidURLs |
boolean | no | If invalid URLs are specified in the urls array, they will be ignored. Instead of them failing the entire request, a batch scrape using the remaining valid URLs will be created, and the invalid URLs will be returned in the invalidURLs field of the response. |
includeTags |
array | no | Tags to include in the output. |
location |
string | no | Location settings for the request. When specified, this will use an appropriate proxy if available and emulate the corresponding language and timezone settings. Defaults to ‘US’ if not specified. |
maxAge |
integer | no | Returns a cached version of the page if it is younger than this age in milliseconds. If a cached version of the page is older than this value, the page will be scraped. If you do not need extremely fresh data, enabling this can speed up your scrapes by 500%. Defaults to 2 days. |
maxConcurrency |
integer | no | Maximum number of concurrent scrapes. This parameter allows you to set a concurrency limit for this batch scrape. If not specified, the batch scrape adheres to your team’s concurrency limit. Example: 10 |
mobile |
boolean | no | Set to true if you want to emulate scraping from a mobile device. Useful for testing responsive pages and taking mobile screenshots. |
onlyMainContent |
boolean | no | Only return the main content of the page excluding headers, navs, footers, etc. |
parsers |
array | no | Controls how files are processed during scraping. When “pdf” is included (default), the PDF content is extracted and converted to markdown format, with billing based on the number of pages (1 credit per page). When an empty array is passed, the PDF file is returned in base64 encoding with a flat rate of 1 credit total. |
proxy |
string | no | Specifies the type of proxy to use. - basic: Proxies for scraping sites with none to basic anti-bot solutions. Fast and usually works. - enhanced: Enhanced proxies for scraping sites with advanced anti-bot solutions. Slower, but more reliable on certain sites. Billed at the same credit cost as basic. - auto: Firecrawl will automatically retry scraping with enhanced proxies if the basic proxy fails. Enhanced proxies carry no credit surcharge, so either way only the regular cost is… |
removeBase64Images |
boolean | no | Removes all base 64 images from the output, which may be overwhelmingly long. The image’s alt text remains in the output, but the URL is replaced with a placeholder. |
skipTlsVerification |
boolean | no | Skip TLS certificate verification when making requests |
storeInCache |
boolean | no | If true, the page will be stored in the Firecrawl index and cache. Setting this to false is useful if your scraping activity may have data protection concerns. Using some parameters associated with sensitive scraping (actions, headers) will force this parameter to be false. |
timeout |
integer | no | Timeout in milliseconds for the request. |
urls |
array | yes | |
waitFor |
integer | no | Specify a delay in milliseconds before fetching the content, allowing the page sufficient time to load. |
webhook |
string | no | A webhook specification object. |
zeroDataRetention |
boolean | no | If true, this will enable zero data retention for this batch scrape. To enable this feature, please contact help@firecrawl.dev |
Output
The run’s result holds the provider’s answer. Signed in, looot inspect firecrawl-batch-scrape prints its fields.
Run it
Every call needs your API token in LOOOT_TOKEN; Sign in shows how to get one. Each run also needs a new idempotency key, so a retry never pays twice. With wait: 30 the answer comes back inline when the run ends within 30 seconds. Otherwise you get the running run back: poll GET /v1/runs/<runId>.
Example input with placeholder values:
curl -X POST "https://api.looot.ai/v1/runs" \
-H "Authorization: Bearer $LOOOT_TOKEN" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: $(uuidgen)" \
-d '{"endpointId":"firecrawl-batch-scrape","input":{"urls":[],"maxConcurrency":10},"wait":30}'const response = await fetch("https://api.looot.ai/v1/runs", {
method: "POST",
headers: {
Authorization: `Bearer ${process.env.LOOOT_TOKEN}`,
"Content-Type": "application/json",
"Idempotency-Key": crypto.randomUUID(),
},
body: JSON.stringify({
endpointId: "firecrawl-batch-scrape",
input: {
urls: [],
maxConcurrency: 10,
},
wait: 30,
}),
});
const run = await response.json();
console.log(run.status, run.result);import os
import uuid
import requests
response = requests.post(
"https://api.looot.ai/v1/runs",
headers={
"Authorization": f"Bearer {os.environ['LOOOT_TOKEN']}",
"Idempotency-Key": str(uuid.uuid4()),
},
json={
"endpointId": "firecrawl-batch-scrape",
"input": {
"urls": [],
"maxConcurrency": 10,
},
"wait": 30,
},
timeout=90,
)
run = response.json()
print(run["status"], run.get("result"))To let looot pick among every provider of this job instead, send job:web.scrape.bulk.submit as endpointId; see the job page.