Shunya LabsShunya LabsPlayground
Docs

Text-to-Speech (TTS)

Radiance is Shunya's speech synthesis family. 46 speaker voices across 23 Indic languages and English. Every voice can speak every language. 11 expression styles. Voice cloning from a 3-6 second reference clip.

First: get an access token

The examples below send Authorization: Bearer $ACCESS_TOKEN. The speech APIs accept only a short-lived access token — never your API key directly.

From the Shunya Playground (recommended). Open API keys in the Playground, click Generate token next to your API key, and copy it — then set it:

shell
export ACCESS_TOKEN="eyJhbGciOiJSUzI1NiIs…paste-here"

Or mint it from your API key — the path for production, where your app refreshes the token as it nears expiry (the response carries expires_in):

shell
export ACCESS_TOKEN=$(curl -s -X POST https://app.shunyalabs.ai/api/auth/token \
  -H "api-key: $SHUNYALABS_API_KEY" | jq -r .token)

Radiance models

Pass the model id in the model field. The product names are for reading; the ids are what the API accepts (the earlier zero-* ids still work).

ModelModel idLanguagesUse it for
Radiance Indicradiance-indic23 Indic languages + EnglishThe default. All 46 named voices, expression styles and voice cloning.
Radiance Orientalradiance-orientalJapanese, KoreanNative Japanese and Korean speech.
Radiance Universalradiance-universalExtended world languages (see Voices & languages)French, German, Spanish, Arabic and other non-Indic languages.

How it fits together

Batch vs Streaming

Two synthesis modes are available. Same model, same voices, different transport and different "when does the first byte of audio leave the server."

Batch
HTTP POST, returns a complete file

Send text via HTTP POST and receive a complete audio file in a single response.

  • Pre-rendered voice prompts for IVR and telephony systems.
  • Notification audio, order updates, alerts, reminders.
  • Podcast, audiobook, and long-form content generation.
  • Any use case where audio does not need to start playing before synthesis is complete.
Transport
HTTP POST
Endpoint
https://ttsv2.shunyalabs.ai/v1/audio/speech
Auth
Bearer <ACCESS_TOKEN>
Required
text, model, voice
Default format
mp3
1POST text→2Server synthesizes→3Receive audio
Streaming
WebSocket, chunks arrive in real time

Open a persistent WebSocket connection and receive audio chunks in real time as synthesis happens.

  • Voice agents and conversational AI requiring sub-second audio start.
  • IVR and telephony pipelines.
  • Real-time audio playback in applications.
  • Any use case where audio must begin playing before synthesis of the full text is complete.
Transport
WebSocket
Endpoint
wss://ttsv2.shunyalabs.ai/v1/realtime
also: /ws/tts, /ws
Config
TTSConfig
Default format
mp3
1Connect→2Receive chunks→3Done

Source: Shunyalabs TTS Developer Documentation v1.0 (March 2026), §2.1 Batch overview and §3.1 Streaming overview, text reproduced verbatim.

Endpoints

ModeEndpointDefault format
BatchPOST https://ttsv2.shunyalabs.ai/v1/audio/speechmp3
Streamingwss://ttsv2.shunyalabs.ai/v1/realtimemp3 (pcm recommended)
HealthGET https://ttsv2.shunyalabs.ai/health-

Your first synthesis

shell
curl -X POST https://ttsv2.shunyalabs.ai/v1/audio/speech \
  -H "Authorization: Bearer $ACCESS_TOKEN" \
  -H "Content-Type: application/json" \
  -d '{"model":"radiance-indic","input":"Hello, how are you today?","voice":"Varun"}' \
  --output hello.mp3
python
import asyncio
from shunyalabs import AsyncShunyaClient
from shunyalabs.tts import TTSConfig

async def main():
    async with AsyncShunyaClient() as client:
        result = await client.tts.synthesize(
            "Hello, how are you today?",
            config=TTSConfig(model="radiance-indic", voice="Varun"),
        )
        result.save("hello.mp3")

asyncio.run(main())
python
from openai import OpenAI

client = OpenAI(
    api_key=os.environ["ACCESS_TOKEN"],
    base_url="https://ttsv2.shunyalabs.ai/v1",
)
response = client.audio.speech.create(
    model="radiance-indic",
    input="Hello, how are you today?",
    voice="Varun",
    response_format="mp3",
)
response.stream_to_file("output.mp3")

Key features

Required fields, at a glance

json
{
  "model": "radiance-indic",         // required
  "input": "Your text here",     // required, up to 10,000 chars
  "voice": "Varun",              // required, see Voices page for full list
  "response_format": "mp3",      // optional, default mp3
  "speed": 1.0,                  // optional, 0.25 to 4.0
  "language": "en",              // optional, ISO code for preprocessing
  "trim_silence": false,         // optional, tight audio when true
  "quality": "low",              // optional, low | medium | high
  "latex": false,                // optional, speak LaTeX math aloud
  "background_audio": null,      // optional, ambient preset or base64
  "background_volume": 0.1,      // optional, 0.0 to 1.0
  "reference_wav": "...",        // optional, base64 for voice cloning
  "reference_text": "..."        // optional, transcript for voice cloning
}