---
title: "ElevenLabs / Speech with character timings (for captions)"
description: "ElevenLabs tool elevenlabs-text-to-speech on looot: input fields, $0.00008 per result, output shape, and code to run it with curl, JavaScript or Python."
sidebar:
  hidden: true
---

{/* Generated by scripts/generate-api-reference.mjs from data/api-reference.json. Do not edit. */}

Reads text aloud and returns the speech (base64) plus the timing of every character, for captions and lip sync. Send text (up to 2,000 characters); voice_id, model_id and output_format are optional. For a plain MP3 download, use elevenlabs-text-to-speech-audio.

- **Tool id:** `elevenlabs-text-to-speech`
- **Provider:** [ElevenLabs](/providers/elevenlabs)
- **Job:** [Turn text into speech with character timings](/reference/jobs/audio-speech-timestamps) (`audio.speech.timestamps`)
- **Price:** $0.00008 per result. A call that fails at the provider costs $0.

## Inputs

| Name | Type | Required | Description |
| --- | --- | --- | --- |
| `language_code` | string | no | Optional ISO 639-1 language code, for example en or fr. |
| `model_id` | string | no | Voice model. eleven_multilingual_v2 and eleven_v3 cost $0.08 per 1,000 characters; eleven_flash_v2_5 and eleven_turbo_v2_5 cost $0.04. Defaults to eleven_multilingual_v2. |
| `text` | string | yes | The words to speak, up to 2,000 characters (about two minutes of speech). |
| `voice_settings` | object | no | Optional: stability, similarity_boost, style, use_speaker_boost, speed. |
| `output_format` | string | no | Audio format. Defaults to mp3_22050_32, which keeps long texts under the 1 MB reply limit. |
| `voice_id` | string | no | ElevenLabs voice id. Defaults to Rachel (21m00Tcm4TlvDq8ikWAM). |

## Output

Shape of the run's `result`, the expected shape, not yet checked against real answers:

```txt
{ audio_base64: string, alignment: { characters: string[], character_start_times_seconds: number[], character_end_times_seconds: number[] } }
```

## Run it

Every call needs your API token in `LOOOT_TOKEN`; [Sign in](/get-started/sign-in#use-the-token-in-scripts-and-agents) shows how to get one. Each run also needs a new idempotency key, so a retry never pays twice. With `wait: 30` the answer comes back inline when the run ends within 30 seconds. Otherwise you get the running run back: poll `GET /v1/runs/<runId>`.

Example input with placeholder values:

<CodeGroup>

```bash curl
curl -X POST "https://api.looot.ai/v1/runs" \
  -H "Authorization: Bearer $LOOOT_TOKEN" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: $(uuidgen)" \
  -d '{"endpointId":"elevenlabs-text-to-speech","input":{"text":"<text>"},"wait":30}'
```

```js JavaScript
const response = await fetch("https://api.looot.ai/v1/runs", {
  method: "POST",
  headers: {
    Authorization: `Bearer ${process.env.LOOOT_TOKEN}`,
    "Content-Type": "application/json",
    "Idempotency-Key": crypto.randomUUID(),
  },
  body: JSON.stringify({
    endpointId: "elevenlabs-text-to-speech",
    input: {
      text: "<text>",
    },
    wait: 30,
  }),
});
const run = await response.json();
console.log(run.status, run.result);
```

```python Python
import os
import uuid

import requests

response = requests.post(
    "https://api.looot.ai/v1/runs",
    headers={
        "Authorization": f"Bearer {os.environ['LOOOT_TOKEN']}",
        "Idempotency-Key": str(uuid.uuid4()),
    },
    json={
        "endpointId": "elevenlabs-text-to-speech",
        "input": {
            "text": "<text>",
        },
        "wait": 30,
    },
    timeout=90,
)
run = response.json()
print(run["status"], run.get("result"))
```

</CodeGroup>

To let looot pick among every provider of this job instead, send `job:audio.speech.timestamps` as `endpointId`; see [the job page](/reference/jobs/audio-speech-timestamps).
