Picsart API Platform
← All models

Eleven Scribe v2

ElevenLabs · Text · Audio → Text

Transcribe speech from audio or video, with word timings and speaker labels.

Model ID: eleven-speech-to-textWorkflow: elevenlabs/v1/speech-to-text
TranscriptionSpeaker Labels
Try on Playground ↗

Overview

Alongside the unified model APIs, we expose compatibility APIs that take each vendor's original parameters exactly as the vendor defines them — nothing renamed, nothing reshaped on the way through.

That makes them the shortest path onto Picsart if you already work with the vendor directly: the request bodies you've already written keep working as they are. You keep their parameter names and defaults, and you get the vendor's full parameter set rather than the subset that is shared across every model.

The cost is that a request is written for one vendor — switching models later means rewriting it, and results come back in the vendor's own shape. When you'd rather write once and change models freely, use the unified API.

Make a request

Call the workflow with ai.apis.run() in TypeScript, or hit /workflows/{workflow}/execute directly — params is passed through untouched either way.

These endpoints are addressed by workflow name rather than model id: the one in this model's header, since one model runs as one workflow. Authentication is unchanged — your Picsart API key as a bearer token (see Authentication).

ts
import { createClient, ApiRunMode } from '@picsart/ai-sdk';

const ai = createClient({
  apiKey: process.env.PICSART_API_KEY,
  apiUrl: 'https://api.picsart.com',
});

// Calls the 'elevenlabs/v1/speech-to-text' workflow directly — params are sent as-is.
const { result, usage } = await ai.apis.run('elevenlabs/v1/speech-to-text', {
  audio_url: "https://cdn.example.com/audio/input.mp3"
}, {
  mode: ApiRunMode.SYNC,
});

console.log(result); // workflow-specific output
console.log(usage?.credits); // credits charged

Parameters

9 parameters, sent inside params. These are the vendor's own names, so they line up one-for-one with the vendor's documentation. Required ones must be supplied; the rest fall back to their defaults.

ParameterTypeRequiredDefaultDetails
audio_url
URL of the audio or video file to transcribe
stringyes—
e.g. https://cdn.example.com/audio/input.mp3
model_id
The transcription model to use
stringnoscribe_v2
scribe_v2
language_code
ISO 639-1 or 639-3 language code. Omit to let the model detect the language
stringno—
e.g. en
diarize
Annotate which speaker is talking in each word
booleanno——
num_speakers
How many speakers are in the audio. Omit to let the model decide
numberno—
1–32
timestamps_granularity
Granularity of the returned timestamps
stringno—
wordcharacter
tag_audio_events
Tag non-speech events such as (laughter) in the transcript
booleanno——
temperature
Sampling temperature. Higher values transcribe more freely
numberno—
0–2
seed
Seed for deterministic transcription
numberno—
0–2147483647

Response

Over HTTP the output arrives inside a status envelope, at response.result. ai.apis.run() unwraps that envelope for you and resolves to { result, usage } instead. Either way the resultitself is the vendor's own shape.

json
{
  "result": {
    "text": "…",
    "language_code": "…",
    "language_probability": 0,
    "words": [
      {
        "text": "…",
        "type": "…",
        "start": 0,
        "end": 0,
        "speaker_id": "wf_abc123",
        "characters": [
          {}
        ]
      }
    ],
    "audio_duration_secs": 0,
    "transcription_id": "wf_abc123"
  },
  "usage": {
    "credits": 1
  }
}