Shunya LabsShunya LabsPlayground
Docs

Automated Speech Recognition configuration

Every parameter you can pass to POST /v1/audio/transcriptions, in one place. Two are required (file or url, and model); the rest have safe defaults.

First: get an access token

The examples below send Authorization: Bearer $ACCESS_TOKEN. The speech APIs accept only a short-lived access token — never your API key directly.

From the Shunya Playground (recommended). Open API keys in the Playground, click Generate token next to your API key, and copy it — then set it:

shell
export ACCESS_TOKEN="eyJhbGciOiJSUzI1NiIs…paste-here"

Or mint it from your API key — the path for production, where your app refreshes the token as it nears expiry (the response carries expires_in):

shell
export ACCESS_TOKEN=$(curl -s -X POST https://app.shunyalabs.ai/api/auth/token \
  -H "api-key: $SHUNYALABS_API_KEY" | jq -r .token)

Required

FieldTypeDescription
filefile uploadAudio file (WAV, MP3, M4A, OGG, FLAC, WebM). One of file or url required.
urlstringPublic audio URL. Can't be combined with file.
modelstringOne of horizon-indic, horizon-universal, horizon-medical, horizon-codeswitch. These are the Horizon Indic, Universal, Medical and Code-Switch models; see Horizon models.

Language & output

ParameterType, defaultDescription
language_codestring, "auto"Language hint. Prefer an ISO 639-1 code (hi, ta, kn, bn, mr, te, gu, pa, ml, or, ur, en) — the canonical form. Full English names (Hindi, Tamil) are also accepted and mapped to the code, and matching is case-insensitive (hi, HI, Hindi are equivalent). Use auto to detect. The codes above are the primary Indic set — ASR serves 204 languages in total (the Indic tier plus Japanese/Korean and global languages); call GET /languages for the routable set with codes and scripts.
response_formatstring, "verbose_json"verbose_json for full response (segments, NLP, timing, language). json for minimal {"text": "..."}: OpenAI-compatible.
output_scriptstring, "auto"Transliterate to a different script without changing language. Devanagari, Bengali, Telugu, Tamil, Kannada, Latin, ITRANS. Deterministic — no language model involved, and no measurable latency cost.
translationseparate endpointNot a transcription parameter. Transcribe, then POST the transcript text to POST /v1/translate with target_language (it returns { "translation": "..." }). For a script change only, use output_script.
Why setting language_code matters
On clips < 5 seconds, language detection is error-prone and the model can "translate" instead of transcribe. Locking the language avoids both.

Example: transcribe Hindi audio and romanise the script to Latin:

shell
curl -X POST https://asrv2prod.shunyalabs.ai/v1/audio/transcriptions \
  -H "Authorization: Bearer $ACCESS_TOKEN" \
  -F "[email protected]" \
  -F "model=horizon-indic" \
  -F "language_code=hi" \
  -F "output_script=Latin"

Returns: "namaste mohammad ji ye ek zaruri call hai" (Hindi pronunciation, written in Latin letters).

Your vocabulary

Recognisers mis-hear words they have never seen — drug names, product names, people, acronyms, places. Two settings fix that, and they work at different moments. Reaching for the wrong one is the usual reason a term keeps coming back wrong.

ParameterType, defaultDescription
boost_phrasesstring, ""Terms to bias recognition toward, separated by || or newlines. Acts during decoding, so it can recover a word the recogniser would otherwise never produce. Match your own spelling and capitalisation — case differences usually still match but not always. Send at most ten terms: that is the largest list measured to cost nothing, and on a live connection anything beyond it is dropped. A short, specific list also works better than a long one.
boost_weightfloat, 10How strongly to bias, for Indian-language audio. Leave it out. Omitting it is exactly the same as sending 10, which is both the default and the maximum — higher values are clamped, because above it the recogniser starts inserting boosted terms that were never spoken. Send 0 to apply your terms with no biasing at all. It has no effect on English or on auto: the terms still apply, but the number does not change the result.
keyterm_keywordsstring, ""JSON array of your own spellings, applied after recognition: wherever the transcript contains one of these misheard or spelled phonetically, it is replaced with your spelling. Only used when enable_keyterm_normalization is on, and only affects nlp_analysis.normalized_text — the raw text is left as spoken.
Which one do I want?
Use boost_phrases when the term must appear correctly in the transcript itself — a prescription, an order number, a customer name. Use keyterm_keywords when you want the transcript left verbatim but need a cleaned-up copy alongside it. They compose: boosting improves what is heard, normalisation tidies how it is written.

Boosting works on both endpoints and in every language. On a real-time connection, put the terms in the opening frame — they apply for the life of the connection, so you send them once rather than per turn:

json
{
  "language_code": "hi",
  "sample_rate": 16000,
  "boost_phrases": "Empagliflozin||NACH mandate||Bengaluru",
  "boost_weight": 30
}

The ready frame echoes boost_phrases as a count, so you can confirm your list was read without your vocabulary being sent back over the wire.

shell
# Without boosting, an unfamiliar drug name loses its first syllable:
#   "The patient was prescribed Mpagliflozin and Dapagliflozin ..."

curl -X POST https://asrv2prod.shunyalabs.ai/v1/audio/transcriptions \
  -H "Authorization: Bearer $ACCESS_TOKEN" \
  -F "[email protected]" \
  -F "language_code=en" \
  -F "boost_phrases=Empagliflozin||Dapagliflozin"

#   "The patient was prescribed Empagliflozin and Dapagliflozin ..."

Segmentation & alignment

ParameterType, defaultDescription
response_format=verbose_jsonbool, falseAdds per-word start, end, and score to every segment. Computed in-pipeline, with no external call and negligible added latency.
enable_diarizationbool, falseSpeaker-level segmentation. Every segment gets a speaker: SPEAKER_XX label; the top-level transcript is prefixed with speaker tags; speakers array lists all unique speakers detected. Segments are capped at 30 s each for quality.
speaker_id + projectbool, falseResolves anonymous SPEAKER_XX labels to registered names. Requires enable_diarization=true and pre-registered voice profiles. project scopes the speaker library, useful for per-customer isolation. Speaker registration API →
emotionbool, falseDetects the dominant emotion in each segment and adds an emotion field. Works alongside standard diarization.

Intelligence layer

ParameterType, defaultDescription
enable_intent_detection + intent_choicesbool, falseClassifies the overall transcript intent. Optionally constrain to a list of allowed intents with intent_choices (JSON array). Result in nlp_analysis.intent: label, confidence, reasoning.
enable_summarization + summary_max_lengthbool, falseGenerates a concise transcript summary. summary_max_length is an approximate word count (default 150). Result in nlp_analysis.summary.
enable_sentiment_analysisbool, falseReturns nlp_analysis.sentiment with a label (positive/negative/neutral), a numeric score, and a short explanation.
enable_keyterm_normalization + keyterm_keywordsbool, falseNormalises domain-specific terms the ASR model might render informally. Optionally focus on specific terms with keyterm_keywords (JSON array). Output preserves the original language. The corrected terms appear in the transcript itself.

Example: classify the call into one of four intents:

shell
curl -X POST https://asrv2prod.shunyalabs.ai/v1/audio/transcriptions \
  -H "Authorization: Bearer $ACCESS_TOKEN" \
  -F "[email protected]" \
  -F "model=horizon-indic" \
  -F "enable_intent_detection=true" \
  -F 'intent_choices=["complaint","inquiry","service_request","compliment"]'

Redaction

ParameterType, defaultDescription
enable_profanity_hashingbool, falseReplaces profane words with **** in-place in both segments[].text and the top-level text. Uses a language model for detection.
hash_keywordsJSON array, noneMasks a specific list of words/phrases using regex (case-insensitive, no LLM). Independent of profanity hashing. Use for account numbers, card numbers, OTP, and any custom sensitive terms.

Example: mask account numbers, card numbers, and OTPs in the transcript:

shell
curl -X POST https://asrv2prod.shunyalabs.ai/v1/audio/transcriptions \
  -H "Authorization: Bearer $ACCESS_TOKEN" \
  -F "[email protected]" \
  -F "model=horizon-indic" \
  -F 'hash_keywords=["account number","card number","OTP"]'

Legacy / compatibility

ParameterType, defaultDescription
taskstring, "transcribe"OpenAI-compatible field. Only "transcribe" is supported today.

Putting it all together

shell
curl -X POST https://asrv2prod.shunyalabs.ai/v1/audio/transcriptions \
  -H "Authorization: Bearer $ACCESS_TOKEN" \
  -F "[email protected]" \
  -F "model=horizon-indic" \
  -F "language_code=hi" \
  -F "enable_diarization=true" \
  -F "enable_speaker_identification=true" \
  -F "project=support_team" \
  -F "enable_emotion_diarization=true" \
  -F "enable_intent_detection=true" \
  -F 'intent_choices=["complaint","inquiry","service_request"]' \
  -F "enable_summarization=true" \
  -F "enable_sentiment_analysis=true" \
  -F "boost_phrases=Empagliflozin||NACH mandate||Bengaluru"
A sane default set
For a contact-centre transcription with agent-assist, start with: model=horizon-indic, enable_diarization=true, enable_intent_detection=true, enable_sentiment_analysis=true. Add response_format=verbose_json if you need precise search.