Shunya LabsShunya LabsPlayground
Docs

Automated Speech Recognition (ASR)

Horizon is Shunya's Automated Speech Recognition family. One API surface for batch and streaming, a choice of four models tuned for different domains, and an intelligence layer that adds diarization, emotion, intent, and more on top of the transcript.

First: get an access token

The examples below send Authorization: Bearer $ACCESS_TOKEN. The speech APIs accept only a short-lived access token — never your API key directly.

From the Shunya Playground (recommended). Open API keys in the Playground, click Generate token next to your API key, and copy it — then set it:

shell
export ACCESS_TOKEN="eyJhbGciOiJSUzI1NiIs…paste-here"

Or mint it from your API key — the path for production, where your app refreshes the token as it nears expiry (the response carries expires_in):

shell
export ACCESS_TOKEN=$(curl -s -X POST https://app.shunyalabs.ai/api/auth/token \
  -H "api-key: $SHUNYALABS_API_KEY" | jq -r .token)

How it fits together

Endpoints

ModeEndpointUse for
BatchPOST https://asrv2prod.shunyalabs.ai/v1/audio/transcriptionsUploaded files, post-processing, async jobs.
Streamingwss://asrv2prod.shunyalabs.ai/v1/realtimeLive transcription, voice agents, IVR.
HealthGET https://asrv2prod.shunyalabs.ai/healthLiveness checks. No auth.
LanguagesGET https://asrv2prod.shunyalabs.ai/languagesReturns supported language names, ISO codes, and scripts.
Speakers/v1/speakers/*Register, list, identify, delete voice profiles for speaker identification.

Batch vs Streaming

Same models, same intelligence layer, different transports. Pick batch when you have a complete audio file in hand. Pick streaming when audio is arriving live and you want partial transcripts as the speaker is still talking.

Batch
HTTP POST, transcribe an audio file

Accepts multipart/form-data. Required fields: file (or url) and model.

  • horizon-indic, Horizon Indic. General Indian languages (Hindi, Tamil, Telugu, Kannada, Marathi, Bengali, etc.)
  • horizon-medical, Horizon Medical. Medical/clinical audio, auto-applies medical terminology correction
  • horizon-codeswitch, Horizon Code-Switch. Code-switched speech (Hinglish, Tanglish, etc.), auto-restores English words to Latin script
  • horizon-universal, Horizon Universal. 99-language model, English, European, Asian, and African languages
Endpoint
POST /v1/audio/transcriptions
Host
asrv2prod.shunyalabs.ai
Content type
multipart/form-data
Auth
Bearer <ACCESS_TOKEN>
Required
file (or url) and model
Default response
verbose_json
1Upload file or url→2Server transcribes→3Single JSON response
Streaming
WebSocket, real-time partials + finals

Real-time streaming transcription over WebSocket. Supports binary mode (raw PCM/ulaw/alaw bytes) and JSON mode (base64-encoded audio frames).

  • ulaw: G.711 mu-law (8-bit), Telephony (8 kHz)
  • alaw: G.711 A-law (8-bit), Telephony (8 kHz)
  • int16: 16-bit signed PCM, General recording
  • float32: 32-bit IEEE float, Pre-processed audio
Endpoint
wss://asrv2prod.shunyalabs.ai/v1/realtime
Init message
JSON config (first frame)
Sample rate
16000 Hz (default)
Chunk size
2.0 s (default)
Silence threshold
0.8 s (default)
1Connect to /ws→2Send JSON init→3Stream audio frames→4Send "END"→5Receive events until done
readypartialfinalfinal_refinedutterance_enderror

Source: Shunyalabs ASR Gateway API Reference (31 March), "Base Call" and "WebSocket Streaming API" sections, reproduced verbatim.

Pick a model

Every request takes a model field. There are four to choose from:

Horizon Indic horizon-indic

Hindi, Tamil, Telugu, Kannada, Marathi, Bengali and 50+ Indian languages. The default for Indic content.

View model →
Horizon Universal horizon-universal

204-language model. English, European, Asian, African. Auto-detects language when you don't know it.

View model →
Horizon Medical horizon-medical

Clinical / medical speech with drug, procedure, and diagnosis vocabulary. HIPAA-cleared. Auto-applies medical terminology correction.

View model →
Horizon Code-Switch horizon-codeswitch

Native handling of mixed Hindi-English (Hinglish), Tamil-English (Tanglish) and similar blends. Automatic code-switch restoration.

View model →

Your first request

Minimum viable: file, model, bearer token. Add language_code whenever you already know the language — it is what decides which model actually runs.

shell
curl -X POST https://asrv2prod.shunyalabs.ai/v1/audio/transcriptions \
  -H "Authorization: Bearer $ACCESS_TOKEN" \
  -F "[email protected]" \
  -F "model=horizon-indic" \
  -F "language_code=hi"

Or pass a URL instead of uploading a file:

shell
curl -X POST https://asrv2prod.shunyalabs.ai/v1/audio/transcriptions \
  -H "Authorization: Bearer $ACCESS_TOKEN" \
  -F "url=https://example.com/call.wav" \
  -F "model=horizon-indic" \
  -F "language_code=hi"
language_code is the routing lever
Pass an ISO 639-1 code (hi, en) or the full English name (Hindi) and the request is routed straight to that language. Leave it out, or send auto, and the service runs language identification first and routes on what it detects. Detection is good but not free: on short, noisy or code-mixed audio it can pick the wrong language, and every downstream stage then inherits that choice. If you know the language, say so — it is the single highest-leverage field on this endpoint.