Phones, kiosks, and head units shouldn't have to depend on the network for the understanding layer.Edge SLU packages the same models that run in the cloud for on-device inference, no GPU at runtime,graceful fallback when the on-device graph can't close a query.
No GPU
At runtime, on any device
15-100M
Classifier params per language
240+
Concurrent users / L4 GPU (cloud reference)
Instant
Replies in under a fifth of a second
How Edge SLU Works
Direct speech-to-intent, skipping the ASR transcript entirely. 15 to 100M parameters per language, running on Android devices with as little as 2GB RAM. Classifier latency under 200ms; end-to-end p95 under 1000ms including a TTS confirmation.
Speech maps directly to intent on the device. No network round trip is needed to understand what the user said - only the final business action (e.g. calling an account API) needs connectivity.
About 200ms first audio emission on edge SoC, MOS 4.3 in deployed production. 5–15MB per-language voice packs delivered via CDN as a one-time download - doesn't add to app install size or cold-start time, loads asynchronously.
The voice understanding layer runs without the internet. Graceful fallback to a server-side ASR plus LLM stack when the on-device graph can't close a query.
Hindi, Indian English, and Bangla on-device for a Tier-1 Indian telco's mobile app, with a roadmap to all 55 Indian languages on the same edge stack - the same Vaani-trained foundation that powers the cloud tier.
Why edge
Voice assistants fail where connectivity is weakest. Underground parking. Moving vehicles. Hospitals. Factory floors. Retail stores. Public kiosks.Edge SLU keeps understanding speech locally, even when the network doesn't.
Works without internet connectivity. Ideal for kiosks, vehicles, mobile apps, and industrial devices.
Footprint
Edge AI succeeds only if it's small enough to deploy everywhere.
Architecture
Use cases
Not every voice interaction can depend on the cloud. Edge SLU enables instant, private, and reliable voice experiences directly on the device, helping enterprises deliver natural interactions even in low-connectivity environments.
Performance Snapshot
Production reference: in-app menu-less voice for a Tier-1 Indian telco, Hindi, Indian English, and Bangla live in production today.