Picsart API Platform
← All models

Eleven Scribe v2

ElevenLabs · Text · Audio → Text

Transcribe speech from audio or video, with word timings and speaker labels.

Model ID: eleven-speech-to-textWorkflow: elevenlabs/v1/speech-to-text
TranscriptionSpeaker Labels
Try on Playground ↗

Install the SDK

The SDK targets Node 20+ and ships with TypeScript types.

npm
npm install @picsart/ai-sdk

Authenticate

Create a client once with your API key — it's sent as a bearer token on every request.

TypeScript
import { createClient } from '@picsart/ai-sdk';

const ai = createClient({
  apiKey: process.env.PICSART_API_KEY, // sent as: Authorization: Bearer <key>
  apiUrl: 'https://api.picsart.com',
});

Run the model

Call the model with its parameters. This example uses the required ones.

TypeScript
import { createClient } from '@picsart/ai-sdk';

const ai = createClient({
  apiKey: process.env.PICSART_API_KEY,
  apiUrl: 'https://api.picsart.com',
});

const { text } = await ai.generateText('eleven-speech-to-text', {
  audioUrl: "https://cdn.example.com/input.mp3"
});

console.log(text);

Parameters

7 parameters. Required ones must be supplied; the rest fall back to their defaults. The API takes these same names over HTTP.

ParameterTypeRequiredDefaultDetails
audioUrl
Audio or Video
file (audio)yes——
language
Language (ISO code, optional)
textno—e.g. "e.g. en, rus — omit to auto-detect"
diarize
Label Speakers
booleannofalse—
numSpeakers
Speakers
rangeno11–32 (step 1)
timestampsGranularity
Timing Detail
enumnowordword, character
tagAudioEvents
Tag Audio Events
booleannofalse—
seed
Seed
rangeno—0–2147483647 (step 1)

Output

generateText() resolves to the generated text plus the raw response.

Result
{
  "text": "…generated text…",
  "model": "eleven-speech-to-text"
}