Bright Data / Start a structured dataset scrape
Bright Data tool brightdata-scrape-structured on looot: input fields, $0.0015 per result, output shape, and code to run it with curl, JavaScript or Python.
Starts a managed-scraper job for one or more URLs and returns a snapshot id at once. Send input, dataset_id and format.
- Tool id:
brightdata-scrape-structured - Provider: Bright Data
- Job: Scrape a URL into structured JSON (
web.scrape.structured) - Price: $0.0015 per result. A call that fails at the provider costs $0.
Inputs
| Name | Type | Required | Description |
|---|---|---|---|
input |
array | yes | array of input objects, usually [{url}] |
dataset_id |
string | yes | Bright Data dataset id. Only datasets that return one record per input URL are allowed: LinkedIn person, company, job post and post, Instagram profile and post, TikTok profile, X profile, YouTube channel, Amazon product, ChatGPT answer. The per-dataset rows (brightdata-*) are easier to use. |
format |
string | no | json | ndjson | csv. Example: “json” |
Output
Shape of the run’s result, checked against 4 real answers:
{ snapshot_id } | { id, url, fbid, logo, name, about, image, input: { url }, posts: { id, url, caption, datetime, image_url, content_type, post_hashtags }[], alumni, slogan, account, founded, similar: { Links, title, location, subtitle }[], updates: { date, text, time, title, images: string[], repost: { images, videos: string[], tagged_users, external_links, repost_hangtags, tagged_companies: { ... }[] }, videos: string[], post_id, post_url, text_html, likes_count, tagged_people, comments_count, tagged_companies }[], website, pronouns, biography, employees: { img, link, title }[], followers, f... ...
The shape is cut here. Signed in, looot inspect brightdata-scrape-structured prints all of it.
Run it
Every call needs your API token in LOOOT_TOKEN; Sign in shows how to get one. Each run also needs a new idempotency key, so a retry never pays twice. With wait: 30 the answer comes back inline when the run ends within 30 seconds. Otherwise you get the running run back: poll GET /v1/runs/<runId>.
Example input with placeholder values:
curl -X POST "https://api.looot.ai/v1/runs" \
-H "Authorization: Bearer $LOOOT_TOKEN" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: $(uuidgen)" \
-d '{"endpointId":"brightdata-scrape-structured","input":{"dataset_id":"<dataset_id>","input":[],"format":"json"},"wait":30}'const response = await fetch("https://api.looot.ai/v1/runs", {
method: "POST",
headers: {
Authorization: `Bearer ${process.env.LOOOT_TOKEN}`,
"Content-Type": "application/json",
"Idempotency-Key": crypto.randomUUID(),
},
body: JSON.stringify({
endpointId: "brightdata-scrape-structured",
input: {
dataset_id: "<dataset_id>",
input: [],
format: "json",
},
wait: 30,
}),
});
const run = await response.json();
console.log(run.status, run.result);import os
import uuid
import requests
response = requests.post(
"https://api.looot.ai/v1/runs",
headers={
"Authorization": f"Bearer {os.environ['LOOOT_TOKEN']}",
"Idempotency-Key": str(uuid.uuid4()),
},
json={
"endpointId": "brightdata-scrape-structured",
"input": {
"dataset_id": "<dataset_id>",
"input": [],
"format": "json",
},
"wait": 30,
},
timeout=90,
)
run = response.json()
print(run["status"], run.get("result"))To let looot pick among every provider of this job instead, send job:web.scrape.structured as endpointId; see the job page.