Eleven Scribe v2
ElevenLabs · Text · Audio → Text
Transcribe speech from audio or video, with word timings and speaker labels.
eleven-speech-to-textWorkflow: elevenlabs/v1/speech-to-textAuthentication
Send your Picsart API key as a bearer token on every request.
Authorization: Bearer <PICSART_API_KEY>Make a request
One endpoint serves every model: POST to /workflows/v1/models/execute with the model id and its input in the body. The nesting is deliberate — the outer paramsis the envelope every workflow takes, and the inner one is this model's own input. Those inner names are the same unified ones the SDK takes, so a call translates one-to-one between the two — no per-vendor naming to learn.
curl -X POST 'https://api.picsart.com/workflows/v1/models/execute' \
-H 'Authorization: Bearer $PICSART_API_KEY' \
-H 'Content-Type: application/json' \
-d '{
"params": {
"model": "eleven-speech-to-text",
"params": {
"audioUrl": "https://cdn.example.com/input.mp3"
}
}
}'Parameters
7 parameters, sent inside the inner paramsobject. Identical to the SDK's — one parameter set, whichever transport you use.
| Parameter | Type | Required | Default | Details |
|---|---|---|---|---|
audioUrlAudio or Video | file (audio) | yes | — | — |
languageLanguage (ISO code, optional) | text | no | — | e.g. "e.g. en, rus — omit to auto-detect" |
diarizeLabel Speakers | boolean | no | false | — |
numSpeakersSpeakers | range | no | 1 | 1–32 (step 1) |
timestampsGranularityTiming Detail | enum | no | word | word, character |
tagAudioEventsTag Audio Events | boolean | no | false | — |
seedSeed | range | no | — | 0–2147483647 (step 1) |
Response
The body arrives in the gateway's status envelope, so the output is at response.result — alongside the task id, its status and the credits charged. Asset models return url plus an items array; text models return text instead.
{
"status": "success",
"response": {
"id": "task_abc123",
"status": "COMPLETED",
"result": {
"text": "…generated text…",
"model": "eleven-speech-to-text"
},
"usage": {
"credits": 1
}
}
}