Real-time Voice Agent
A stateless service that bridges live phone-call audio to a real-time speech model over WebSockets, with the conversation driven by tool calls. Under NDA, so this covers the architecture only.
Architecture
- Bidirectional audio bridge
- Two async loops stream audio between the phone media stream and the model, transcoding between telephony and model audio formats on the fly.
- Ephemeral per-call state
- Each call lives in memory, keyed by its ID, and is discarded once results are delivered. No database on the hot path.
- Tool-driven control
- The model advances the session through function calls, and every tool reply tells it what comes next so it stays on track.
- Quality gates
- Voice-activity gating and transcript validation run before anything is recorded; malformed tool calls go through a staged recovery.
- Watchdogs and self-recovery
- Timers cover silence, slow responses and playback; model or rate-limit errors trigger an automatic retry, and WebSocket keepalives stay on.
- Swappable transports
- Telephony and model connections sit behind protocol interfaces, so either provider can be replaced.
- Pluggable delivery
- Results leave as a multipart callback with exponential back-off, a storage-queue message, or both, each behind a feature flag.
- Two deployment targets
- A Docker image with health checks, or serverless containers with separate environments.
Main processing path
- REST triggerstarts the call
- media streamWebSocket
- audio bridgetranscoding, async loops
- speech modelreal-time
- tool dispatchin-memory state
- deliverycallback or queue
Simplified, with generic component names. Highlighted steps use a model; the rest is software.