Voice AI That Works EvenWithout the Internet

Phones, kiosks, and head units shouldn't have to depend on the network for the understanding layer.Edge SLU packages the same models that run in the cloud for on-device inference, no GPU at runtime,graceful fallback when the on-device graph can't close a query.

No GPU

At runtime, on any device

15-100M

Classifier params per language

240+

Concurrent users / L4 GPU (cloud reference)

Instant

Replies in under a fifth of a second

How Edge SLU Works

Voice-to-Intent, Without aTranscript Step

Direct speech-to-intent, skipping the ASR transcript entirely. 15 to 100M parameters per language, running on Android devices with as little as 2GB RAM. Classifier latency under 200ms; end-to-end p95 under 1000ms including a TTS confirmation.

01

On-Device SLU

Speech maps directly to intent on the device. No network round trip is needed to understand what the user said - only the final business action (e.g. calling an account API) needs connectivity.

02

Edge TTS

About 200ms first audio emission on edge SoC, MOS 4.3 in deployed production. 5–15MB per-language voice packs delivered via CDN as a one-time download - doesn't add to app install size or cold-start time, loads asynchronously.

03

Offline-First Understanding

The voice understanding layer runs without the internet. Graceful fallback to a server-side ASR plus LLM stack when the on-device graph can't close a query.

04

In Production Today

Hindi, Indian English, and Bangla on-device for a Tier-1 Indian telco's mobile app, with a roadmap to all 55 Indian languages on the same edge stack - the same Vaani-trained foundation that powers the cloud tier.

Why edge

Cloud Voice AI Was Never Built For Every Environment

Voice assistants fail where connectivity is weakest. Underground parking. Moving vehicles. Hospitals. Factory floors. Retail stores. Public kiosks.Edge SLU keeps understanding speech locally, even when the network doesn't.

NO NETWORKOFFLINEON-DEVICE

Always Available

Works without internet connectivity. Ideal for kiosks, vehicles, mobile apps, and industrial devices.

Footprint

Built For Devices, Not Data Centers.

Edge AI succeeds only if it's small enough to deploy everywhere.

01

Speech Understanding

20–80 MB

Runs locally

Under 200 ms

02

Voice Generation

5–15 MB

Neural speech

Under 50 ms

03

Background Sync

Anonymous telemetry

Optional

Never blocks interaction

04

Human Escalation

Cloud handoff

Automatic

Under 500 ms

Runs on
PhonesVehiclesBanking kiosksHealthcare devicesSmart TVsIndustrial terminalsRetail kiosksHead unitsWearablesPhonesVehiclesBanking kiosksHealthcare devicesSmart TVsIndustrial terminalsRetail kiosksHead unitsWearables

Architecture

How Edge SLU Works

  1. 01
    Speech
  2. 02
    On-device intent model
  3. 03
    Business action
  4. 04
    On-device voice response
  5. 05
    Cloud reasoning
  6. 06
    Done
  1. 01
    Speech
  2. 02
    On-device intent model
  3. 03
    Business action
  4. 04
    On-device voice response
  5. 05
    Cloud reasoning
  6. 06
    Done

Use cases

Built for On-DeviceEnterprise Intelligence

Not every voice interaction can depend on the cloud. Edge SLU enables instant, private, and reliable voice experiences directly on the device, helping enterprises deliver natural interactions even in low-connectivity environments.

Performance Snapshot

Every Component, Measured

Component
Description
Size
RAM
Latency
Intent classifier
15–100M param model, direct audio to intent, one model per language via SDK
20–80 MB
100–300 MB
<200ms
Edge TTS engine
Pre-rendered static templates + on-phone rendering of dynamic variables
5–15 MB
30 MB
<50ms
Streaming module
Async audio streaming to a server-side store for quality monitoring / retraining, non-blocking
<1 MB
10 MB
0ms (async)
Voice bridge
SIP / VoIP bridge to a contact centre when the on-device graph can't close the query
<1 MB
15 MB
<500ms

Related Pillars

Custom SLMsVoice AgentSTTTTSReal Time Translation