Eleven Scribe v2
ElevenLabs · Text · Audio → Text
Transcribe speech from audio or video, with word timings and speaker labels.
Model ID:
eleven-speech-to-textWorkflow: elevenlabs/v1/speech-to-textTranscriptionSpeaker Labels
Install the SDK
The SDK targets Node 20+ and ships with TypeScript types.
npm
npm install @picsart/ai-sdkAuthenticate
Create a client once with your API key — it's sent as a bearer token on every request.
TypeScript
import { createClient } from '@picsart/ai-sdk';
const ai = createClient({
apiKey: process.env.PICSART_API_KEY, // sent as: Authorization: Bearer <key>
apiUrl: 'https://api.picsart.com',
});Run the model
Call the model with its parameters. This example uses the required ones.
TypeScript
import { createClient } from '@picsart/ai-sdk';
const ai = createClient({
apiKey: process.env.PICSART_API_KEY,
apiUrl: 'https://api.picsart.com',
});
const { text } = await ai.generateText('eleven-speech-to-text', {
audioUrl: "https://cdn.example.com/input.mp3"
});
console.log(text);Parameters
7 parameters. Required ones must be supplied; the rest fall back to their defaults. The API takes these same names over HTTP.
| Parameter | Type | Required | Default | Details |
|---|---|---|---|---|
audioUrlAudio or Video | file (audio) | yes | — | — |
languageLanguage (ISO code, optional) | text | no | — | e.g. "e.g. en, rus — omit to auto-detect" |
diarizeLabel Speakers | boolean | no | false | — |
numSpeakersSpeakers | range | no | 1 | 1–32 (step 1) |
timestampsGranularityTiming Detail | enum | no | word | word, character |
tagAudioEventsTag Audio Events | boolean | no | false | — |
seedSeed | range | no | — | 0–2147483647 (step 1) |
Output
generateText() resolves to the generated text plus the raw response.
Result
{
"text": "…generated text…",
"model": "eleven-speech-to-text"
}