Eleven Scribe v2
ElevenLabs · Text · Audio → Text
Transcribe speech from audio or video, with word timings and speaker labels.
eleven-speech-to-textWorkflow: elevenlabs/v1/speech-to-textOverview
Alongside the unified model APIs, we expose compatibility APIs that take each vendor's original parameters exactly as the vendor defines them — nothing renamed, nothing reshaped on the way through.
That makes them the shortest path onto Picsart if you already work with the vendor directly: the request bodies you've already written keep working as they are. You keep their parameter names and defaults, and you get the vendor's full parameter set rather than the subset that is shared across every model.
The cost is that a request is written for one vendor — switching models later means rewriting it, and results come back in the vendor's own shape. When you'd rather write once and change models freely, use the unified API.
Make a request
Call the workflow with ai.apis.run() in TypeScript, or hit /workflows/{workflow}/execute directly — params is passed through untouched either way.
These endpoints are addressed by workflow name rather than model id: the one in this model's header, since one model runs as one workflow. Authentication is unchanged — your Picsart API key as a bearer token (see Authentication).
import { createClient, ApiRunMode } from '@picsart/ai-sdk';
const ai = createClient({
apiKey: process.env.PICSART_API_KEY,
apiUrl: 'https://api.picsart.com',
});
// Calls the 'elevenlabs/v1/speech-to-text' workflow directly — params are sent as-is.
const { result, usage } = await ai.apis.run('elevenlabs/v1/speech-to-text', {
audio_url: "https://cdn.example.com/audio/input.mp3"
}, {
mode: ApiRunMode.SYNC,
});
console.log(result); // workflow-specific output
console.log(usage?.credits); // credits chargedParameters
9 parameters, sent inside params. These are the vendor's own names, so they line up one-for-one with the vendor's documentation. Required ones must be supplied; the rest fall back to their defaults.
| Parameter | Type | Required | Default | Details |
|---|---|---|---|---|
audio_urlURL of the audio or video file to transcribe | string | yes | — | e.g. https://cdn.example.com/audio/input.mp3 |
model_idThe transcription model to use | string | no | scribe_v2 | scribe_v2 |
language_codeISO 639-1 or 639-3 language code. Omit to let the model detect the language | string | no | — | e.g. en |
diarizeAnnotate which speaker is talking in each word | boolean | no | — | — |
num_speakersHow many speakers are in the audio. Omit to let the model decide | number | no | — | 1–32 |
timestamps_granularityGranularity of the returned timestamps | string | no | — | wordcharacter |
tag_audio_eventsTag non-speech events such as (laughter) in the transcript | boolean | no | — | — |
temperatureSampling temperature. Higher values transcribe more freely | number | no | — | 0–2 |
seedSeed for deterministic transcription | number | no | — | 0–2147483647 |
Response
Over HTTP the output arrives inside a status envelope, at response.result. ai.apis.run() unwraps that envelope for you and resolves to { result, usage } instead. Either way the resultitself is the vendor's own shape.
{
"result": {
"text": "…",
"language_code": "…",
"language_probability": 0,
"words": [
{
"text": "…",
"type": "…",
"start": 0,
"end": 0,
"speaker_id": "wf_abc123",
"characters": [
{}
]
}
],
"audio_duration_secs": 0,
"transcription_id": "wf_abc123"
},
"usage": {
"credits": 1
}
}