Picsart API Platform
← All models

Gemini 2.5 Pro TTS

Premium

Google · Audio · Text → Speech

Premium Gemini TTS with richer expressiveness and multi-speaker support.

Model ID: gemini-2.5-pro-ttsWorkflow: gemini/v1/audios
Multilingual30 VoicesMulti-Speaker
Try on Playground ↗

Overview

Alongside the unified model APIs, we expose compatibility APIs that take each vendor's original parameters exactly as the vendor defines them — nothing renamed, nothing reshaped on the way through.

That makes them the shortest path onto Picsart if you already work with the vendor directly: the request bodies you've already written keep working as they are. You keep their parameter names and defaults, and you get the vendor's full parameter set rather than the subset that is shared across every model.

The cost is that a request is written for one vendor — switching models later means rewriting it, and results come back in the vendor's own shape. When you'd rather write once and change models freely, use the unified API.

Make a request

Call the workflow with ai.apis.run() in TypeScript, or hit /workflows/{workflow}/execute directly — params is passed through untouched either way.

These endpoints are addressed by workflow name rather than model id: the one in this model's header, since one model runs as one workflow. Authentication is unchanged — your Picsart API key as a bearer token (see Authentication).

ts
import { createClient, ApiRunMode } from '@picsart/ai-sdk';

const ai = createClient({
  apiKey: process.env.PICSART_API_KEY,
  apiUrl: 'https://api.picsart.com',
});

// Calls the 'gemini/v1/audios' workflow directly — params are sent as-is.
const { result, usage } = await ai.apis.run('gemini/v1/audios', {
  text: "Hello, welcome to our product demo.",
  model: "gemini-2.5-flash-tts"
}, {
  mode: ApiRunMode.SYNC,
});

console.log(result); // workflow-specific output
console.log(usage?.credits); // credits charged

Async (submit & poll)

This model can run longer than the sync limit (~20s). In TypeScript, ai.apis.run(…, { mode: ApiRunMode.ASYNC }) polls for you; over HTTP, submit and poll yourself.

ts
import { createClient, ApiRunMode } from '@picsart/ai-sdk';

const ai = createClient({
  apiKey: process.env.PICSART_API_KEY,
  apiUrl: 'https://api.picsart.com',
});

// mode: ASYNC submits the job and polls under the hood — you just await.
const { result } = await ai.apis.run('gemini/v1/audios', {
  text: "Hello, welcome to our product demo.",
  model: "gemini-2.5-flash-tts"
}, {
  mode: ApiRunMode.ASYNC,
});

console.log(result);

Parameters

7 parameters, sent inside params. These are the vendor's own names, so they line up one-for-one with the vendor's documentation. Required ones must be supplied; the rest fall back to their defaults.

ParameterTypeRequiredDefaultDetails
text
Text to synthesize into speech (single part). Ignored when `parts` is given.
stringno——
style
Delivery direction for the whole text; shorthand for a single part's `style`. Honoured by the 3.8 models.
stringno——
parts
Ordered spoken parts with optional per-part speaker and style. Use instead of `text` for directed or multi-speaker speech.
SpeechPartDto[]no—
max 200 items
└text
Verbatim text to speak. On 3.8, angle-bracket vocal events (<laugh>, <sigh>, <breath>, <short pause>) are honoured inline.
stringyes——
└speaker
Speaker name; must match a `multiSpeakerVoiceConfigs` entry. Required on every part of a multi-speaker request.
stringno——
└style
Natural-language delivery direction for this part, e.g. "frightened, breathy whisper; a scared little girl". Sent as speechMetadata.style, never spoken.
stringno——
model
TTS model to use
stringyes—
gemini-2.5-flash-ttsgemini-2.5-pro-ttsgemini-3.8-flash-ttsgemini-3.8-flash-lite-tts
voiceName
Voice name for single-speaker TTS
stringno—
PuckKoreCharonFenrirAoedeLedaZephyrOrusAutonoeCallirrhoeDespinaErinomeGacruxLaomedeiaPulcherrimaSulafatVindemiatrixAchernarAchirdAlgenibAlgiebaAlnilamEnceladusIapetusRasalgethiSadachbiaSadaltagerSchedarUmbrielZubenelgenubi
multiSpeakerVoiceConfigs
Multi-speaker voice configs (max 2). When provided, voiceName is ignored.
SpeakerVoiceConfigDto[]no—
max 2 items
└speaker
Speaker identifier (e.g. "Speaker 1")
stringyes——
└voiceName
Voice name for this speaker
stringyes—
PuckKoreCharonFenrirAoedeLedaZephyrOrusAutonoeCallirrhoeDespinaErinomeGacruxLaomedeiaPulcherrimaSulafatVindemiatrixAchernarAchirdAlgenibAlgiebaAlnilamEnceladusIapetusRasalgethiSadachbiaSadaltagerSchedarUmbrielZubenelgenubi
options
Options controlling safety checks and drive integration
objectno——
└safety_checks
Safety check settings
objectno——
└enabled
Whether to run content moderation. Defaults to true.
booleanno——
└drive
Save result to Picsart Drive
objectno——
└name
File name in Picsart Drive
stringyes——
└attributes
Custom attributes to attach to the file
objectno——
└folder
Target folder in Picsart Drive
objectno——
└inputs_transformation
Input transformation settings
objectno——
└downscale_oversized_images
Whether to downscale oversized input images. Defaults to false.
booleanno——

Response

Over HTTP the output arrives inside a status envelope, at response.result. ai.apis.run() unwraps that envelope for you and resolves to { result, usage } instead. Either way the resultitself is the vendor's own shape.

json
{
  "result": {
    "audioUrls": [
      {
        "url": "https://cdn.picsart.com/…/result.mp3",
        "mimeType": "audio/mpeg",
        "driveFile": {}
      }
    ]
  },
  "usage": {
    "credits": 0
  }
}