Real-time Voice Agent

Role
Backend engineer, co-built
Type
Production, NDA
Focus
Real-time voice architecture

A stateless service that bridges live phone-call audio to a real-time speech model over WebSockets, with the conversation driven by tool calls. Under NDA, so this covers the architecture only.

Architecture

Bidirectional audio bridge
Two async loops stream audio between the phone media stream and the model, transcoding between telephony and model audio formats on the fly.
Ephemeral per-call state
Each call lives in memory, keyed by its ID, and is discarded once results are delivered. No database on the hot path.
Tool-driven control
The model advances the session through function calls, and every tool reply tells it what comes next so it stays on track.
Quality gates
Voice-activity gating and transcript validation run before anything is recorded; malformed tool calls go through a staged recovery.
Watchdogs and self-recovery
Timers cover silence, slow responses and playback; model or rate-limit errors trigger an automatic retry, and WebSocket keepalives stay on.
Swappable transports
Telephony and model connections sit behind protocol interfaces, so either provider can be replaced.
Pluggable delivery
Results leave as a multipart callback with exponential back-off, a storage-queue message, or both, each behind a feature flag.
Two deployment targets
A Docker image with health checks, or serverless containers with separate environments.

Main processing path

  1. REST triggerstarts the call
  2. media streamWebSocket
  3. audio bridgetranscoding, async loops
  4. speech modelreal-time
  5. tool dispatchin-memory state
  6. deliverycallback or queue

Simplified, with generic component names. Highlighted steps use a model; the rest is software.

Stack

  • Python
  • FastAPI
  • WebSockets
  • Twilio Media Streams
  • Real-time speech API
  • NumPy
  • Azure Queues
  • OpenTelemetry
  • Docker