Shunya LabsShunya LabsPlayground
Docs

Automated Speech Recognition features

The intelligence layer on top of Horizon. Enable each with a boolean flag, none of them are required, and they combine freely. All of these return their results in the same JSON response as the transcript.

First: get an access token

The examples below send Authorization: Bearer $ACCESS_TOKEN. The speech APIs accept only a short-lived access token — never your API key directly.

From the Shunya Playground (recommended). Open API keys in the Playground, click Generate token next to your API key, and copy it — then set it:

shell
export ACCESS_TOKEN="eyJhbGciOiJSUzI1NiIs…paste-here"

Or mint it from your API key — the path for production, where your app refreshes the token as it nears expiry (the response carries expires_in):

shell
export ACCESS_TOKEN=$(curl -s -X POST https://app.shunyalabs.ai/api/auth/token \
  -H "api-key: $SHUNYALABS_API_KEY" | jq -r .token)
Easier to browse?
Open the Intelligence overview for a card-based UI with jump links and collapsible request/response examples for every feature.

1. Diarization

"Who spoke when." Adds speaker: SPEAKER_XX to every segment, a top-level speakers array, and a speaker_turns array of raw turn boundaries. The top-level text stays the plain transcript — it carries no speaker tags, so read the speaker off segments or speaker_turns.

num_speakers=N turns diarization on as well, and tells it how many voices to expect.

Request:

shell
curl -X POST https://asrv2prod.shunyalabs.ai/v1/audio/transcriptions \
  -H "Authorization: Bearer $ACCESS_TOKEN" \
  -F "[email protected]" \
  -F "model=horizon-indic" \
  -F "enable_diarization=true"

Response:

json
{
  "text": "नमस्ते, आप कैसे हैं? मैं ठीक हूँ, धन्यवाद।",
  "segments": [
    { "start": 0.5, "end": 3.2, "text": "नमस्ते, आप कैसे हैं?", "speaker": "SPEAKER_00", "confidence": 0.95 },
    { "start": 4.1, "end": 6.8, "text": "मैं ठीक हूँ, धन्यवाद।", "speaker": "SPEAKER_01", "confidence": 0.98 }
  ],
  "speakers": ["SPEAKER_00", "SPEAKER_01"],
  "speaker_turns": [
    { "start": 0.5, "end": 3.2, "speaker": "SPEAKER_00" },
    { "start": 4.1, "end": 6.8, "speaker": "SPEAKER_01" }
  ]
}
  • Segments are capped at 30 seconds each to maintain transcription quality.
  • Works on any number of speakers, but best with 2-6 distinct voices.
  • Speaker labels are per-request. SPEAKER_00 in one call is not the same person as SPEAKER_00 in the next.

2. Speaker identification

Adds speakers_identified, one voiceprint summary per diarized speaker. Requires diarization. It does not rename anyone: the labels in segments stay SPEAKER_XX, and a registered name does not come back in the transcription response. Registration and the transcript are separate today — the endpoint stores a voiceprint, and transcription reports the voiceprints it found, but nothing joins the two.

Step 1: register a speaker (use a 5-15 second clip of the speaker alone, no background music):

shell
curl -X POST https://asrv2prod.shunyalabs.ai/v1/speakers/register \
  -H "Authorization: Bearer $ACCESS_TOKEN" \
  -F "name=Priya" \
  -F "file=@priya_sample.wav" \
  -F "project=support_team"

Step 2: transcribe with identification on:

shell
curl -X POST https://asrv2prod.shunyalabs.ai/v1/audio/transcriptions \
  -H "Authorization: Bearer $ACCESS_TOKEN" \
  -F "[email protected]" \
  -F "model=horizon-indic" \
  -F "enable_diarization=true" \
  -F "enable_speaker_identification=true" \
  -F "project=support_team"

Response:

json
{
  "speakers": ["SPEAKER_00", "SPEAKER_01"],
  "speakers_identified": [
    { "speaker": "SPEAKER_00", "turns": 1, "voiceprint_dim": 256 },
    { "speaker": "SPEAKER_01", "turns": 1, "voiceprint_dim": 256 }
  ]
}

On the register and delete endpoints, project namespaces a voice library and defaults to playground. Remove a profile with DELETE /v1/speakers/delete, passing the same name and project.

3. Emotion diarization

Detects the dominant emotion in the audio. Adds a top-level emotion object and an emotion_summary, plus emotion and emotion_confidence on every segment. Labels are short codes, not words — neu, hap, sad. Works independently of speaker diarization but is commonly used together.

Request:

shell
curl -X POST https://asrv2prod.shunyalabs.ai/v1/audio/transcriptions \
  -H "Authorization: Bearer $ACCESS_TOKEN" \
  -F "[email protected]" \
  -F "model=horizon-indic" \
  -F "enable_diarization=true" \
  -F "enable_emotion_diarization=true"

Response:

json
{
  "emotion": { "label": "hap", "score": 0.5839 },
  "emotion_summary": {
    "dominant_emotion": "neu",
    "emotion_distribution": { "neu": 50.0, "hap": 50.0 },
    "avg_confidence": 0.4773
  },
  "segments": [
    { "start": 0.08, "end": 5.263, "text": "...", "speaker": "SPEAKER_00", "emotion": "neu", "emotion_confidence": 0.4005 },
    { "start": 5.263, "end": 18.9, "text": "...", "speaker": "SPEAKER_01", "emotion": "hap", "emotion_confidence": 0.554 }
  ]
}

4. Intent detection

Classifies the overall transcript intent. Pass intent_choices to constrain to your taxonomy — the label comes back as one of yours, copied verbatim — or leave it off for a free-form phrase such as Request NACH registration. nlp_analysis.intent is a plain string. There is no confidence or reasoning field alongside it.

Request:

shell
curl -X POST https://asrv2prod.shunyalabs.ai/v1/audio/transcriptions \
  -H "Authorization: Bearer $ACCESS_TOKEN" \
  -F "[email protected]" \
  -F "model=horizon-indic" \
  -F "enable_intent_detection=true" \
  -F 'intent_choices=["complaint","inquiry","service_request","compliment"]'

Response:

json
{
  "nlp_analysis": {
    "intent": "service_request"
  }
}

5. Sentiment analysis

Overall sentiment of the transcript. nlp_analysis.sentiment is a plain string — positive, neutral or negative. There is no numeric score and no explanation field.

Request:

shell
curl -X POST https://asrv2prod.shunyalabs.ai/v1/audio/transcriptions \
  -H "Authorization: Bearer $ACCESS_TOKEN" \
  -F "[email protected]" \
  -F "model=horizon-indic" \
  -F "enable_sentiment_analysis=true"

Response:

json
{
  "nlp_analysis": {
    "sentiment": "negative"
  }
}

6. Summarization

Concise summary of the transcript. summary_max_length is an approximate cap in characters, not words, and it is guidance rather than a hard truncation — the summary finishes its sentence rather than being cut off. Measured: a request for 50 came back at 69 characters, one for 150 at 163.

Request:

shell
curl -X POST https://asrv2prod.shunyalabs.ai/v1/audio/transcriptions \
  -H "Authorization: Bearer $ACCESS_TOKEN" \
  -F "[email protected]" \
  -F "model=horizon-indic" \
  -F "enable_summarization=true" \
  -F "summary_max_length=150"

Response:

json
{
  "nlp_analysis": {
    "summary": "Customer called about a vehicle breakdown. Agent confirmed the complaint was registered and promised a technician within the hour."
  }
}

7. Keyterm normalization

Cleans up domain-specific terms the ASR model might render informally, emi → EMI, nach mandate → NACH mandate. Preserves the original language. Optionally focus on a specific glossary with keyterm_keywords.

Request:

shell
curl -X POST https://asrv2prod.shunyalabs.ai/v1/audio/transcriptions \
  -H "Authorization: Bearer $ACCESS_TOKEN" \
  -F "[email protected]" \
  -F "model=horizon-indic" \
  -F "enable_keyterm_normalization=true" \
  -F 'keyterm_keywords=["EMI","NACH mandate","bounce charge"]'

Response:

json
{
  "text": "सर आपका केवाईसी अपडेट नहीं हुआ है और ईएमआई की बाउंस चार्ज लगी है",
  "nlp_analysis": {
    "keyterms": ["KYC", "EMI", "Bounce charge"],
    "normalized_text": "सर आपका KYC अपडेट नहीं हुआ है और EMI की बाउंस चार्ज लगी है"
  }
}

This runs after recognition and writes both nlp_analysis.keyterms and nlp_analysis.normalized_text; the raw text is left as spoken. To influence what the recogniser actually hears — which is what you want for a name it has never encountered — use keyword boosting instead, or both together.

8. Translation

Translation is a separate endpoint, POST /v1/translate. Transcribe first, then send the transcript text with a target language — an ISO 639-1 code (en, hi) or full name (English).

Request:

shell
curl -X POST https://asrv2prod.shunyalabs.ai/v1/translate \
  -H "Authorization: Bearer $ACCESS_TOKEN" \
  -H "Content-Type: application/json" \
  -d '{"text": "नमस्ते, यह एक ज़रूरी कॉल है।", "target_language": "en"}'

Response:

json
{ "translation": "Hello, this is an urgent call." }

9. Profanity hashing

Masks profane words in-place in both the top-level text and each segment's text. The first letter survives and the rest becomes asterisks, so damn comes back as d***.

Request:

shell
curl -X POST https://asrv2prod.shunyalabs.ai/v1/audio/transcriptions \
  -H "Authorization: Bearer $ACCESS_TOKEN" \
  -F "[email protected]" \
  -F "model=horizon-indic" \
  -F "enable_profanity_hashing=true"

Response:

json
{
  "text": "This a d*** s**** situation and the b****** hung up on me.",
  "segments": [
    { "start": 0.0, "end": 4.13, "text": "This a d*** s**** situation and the b****** hung up on me." }
  ]
}

10. Custom keyword redaction (hash_keywords)

Regex-based masking of specific terms or phrases. No LLM, fast, deterministic. Use for PII and domain-sensitive tokens.

Request:

shell
curl -X POST https://asrv2prod.shunyalabs.ai/v1/audio/transcriptions \
  -H "Authorization: Bearer $ACCESS_TOKEN" \
  -F "[email protected]" \
  -F "model=horizon-indic" \
  -F 'hash_keywords=["account number","card number","OTP","aadhaar"]'

Response:

json
{
  "text": "Please confirm the **** ending 4321 and the **** we sent you yesterday.",
  "segments": [
    { "start": 0.0, "end": 6.4, "text": "Please confirm the **** ending 4321 and the **** we sent you yesterday." }
  ]
}

Matching is literal, so the terms have to be in the same language and script as the transcript. The English list above masks nothing in a Hindi transcript — redacting प्लांट takes प्लांट in the list, not plant.

Masking covers text and segments, not the timing arrays
Both maskers rewrite the top-level text and each segment's text. With response_format=verbose_json the unmasked words are still present in the top-level words array, and in each segment's nbest alternatives. If you are redacting for storage, drop those arrays or don't request them.

11. Word timestamps

Per-word timing with a confidence score. response_format=verbose_json adds one top-level words array covering the whole file — the words are not nested inside segments. Each entry is word, start, end and confidence. Segments carry their own confidence and an nbest list instead. Runs in-pipeline, with no external call and negligible latency overhead.

Request:

shell
curl -X POST https://asrv2prod.shunyalabs.ai/v1/audio/transcriptions \
  -H "Authorization: Bearer $ACCESS_TOKEN" \
  -F "[email protected]" \
  -F "model=horizon-indic" \
  -F "response_format=verbose_json"

Response:

json
{
  "text": "टेलीविजन रिपोर्टों में प्लांट से निकलने वाला सफेद धुआं दिखाया गया है",
  "segments": [
    {
      "start": 0.0,
      "end": 5.1,
      "text": "टेलीविजन रिपोर्टों में प्लांट से निकलने वाला सफेद धुआं दिखाया गया है",
      "confidence": 0.9906
    }
  ],
  "words": [
    { "word": "टेलीविजन", "start": 0.08, "end": 0.797, "confidence": 0.9876 },
    { "word": "रिपोर्टों", "start": 0.797, "end": 1.434, "confidence": 0.9873 },
    { "word": "में", "start": 1.434, "end": 1.594, "confidence": 0.9961 }
  ]
}

12. Keyword boosting

Bias recognition toward terms the model has never seen — drug names, product names, people, places, acronyms. Unlike keyterm normalization (§7), which rewrites text after recognition, boosting acts during decoding, so it can recover a word the recogniser would otherwise never produce at all. The corrected term appears in text itself.

Send a short, specific list — the terms that matter for this call, not your whole catalogue. It works on the batch endpoint and on a real-time connection, where the terms go in the opening frame and apply for the life of the connection.

Request:

shell
curl -X POST https://asrv2prod.shunyalabs.ai/v1/audio/transcriptions \
  -H "Authorization: Bearer $ACCESS_TOKEN" \
  -F "[email protected]" \
  -F "language_code=en" \
  -F "boost_phrases=Empagliflozin||NACH mandate||Bengaluru"

Without the lexicon an unfamiliar drug name loses its opening syllable — “Mpagliflozin”. With it, the term comes back intact.

boost_weight applies to Indian languages only
For Indian-language audio, boost_weight controls how hard the decoder is pushed. The default of 10 is also the maximum, and higher values are clamped — above it the recogniser starts inserting boosted terms that were never spoken, which is worse than the mis-recognition you were fixing. For English and auto the terms still apply, but the strength setting has no effect.

Combining features

Every feature above can be enabled on the same request. Here's a full contact-centre configuration:

shell
curl -X POST https://asrv2prod.shunyalabs.ai/v1/audio/transcriptions \
  -H "Authorization: Bearer $ACCESS_TOKEN" \
  -F "[email protected]" \
  -F "model=horizon-indic" \
  -F "language_code=hi" \
  -F "enable_diarization=true" \
  -F "enable_speaker_identification=true" \
  -F "project=support_team" \
  -F "enable_emotion_diarization=true" \
  -F "enable_intent_detection=true" \
  -F 'intent_choices=["complaint","inquiry","service_request"]' \
  -F "enable_summarization=true" \
  -F "enable_sentiment_analysis=true" \
  -F 'hash_keywords=["account number","card number","OTP"]'
Latency trade-off
Each intelligence feature adds an extra processing pass on top of the ASR path. Enable only what you consume, don't pay for a summary you won't read.