An LLM-powered voice agent is a real-time system that listens to speech, converts it to text, reasons about it with a large language model, and responds through text-to-speech. The most common architecture is the STT, LLM, TTS pipeline, though multimodal speech-to-speech models are emerging as an alternative. VideoSDK's open-source AI Agent SDK lets you assemble either architecture and connect it to real users over web, mobile, or phone in minutes.
Voice is becoming the default interface for AI. Since the arrival of real-time speech models in 2024, developers have moved from clunky IVR menus to agents that actually hold a conversation. The engine behind that shift is the large language model, sitting in the reasoning layer between what a user says and what the agent answers.
If you are building a voice agent in 2026, you are probably deciding between two architectures, choosing STT and TTS providers, and figuring out how to keep latency low enough that the conversation feels natural. This article walks through three concrete LLM examples for voice agents, a customer-support agent, a personal productivity assistant, and a multimodal speech-to-speech agent, and closes with production considerations and a stack-selection framework.

What Makes an LLM-Powered Voice Agent?

Every LLM-powered voice agent runs the same three-stage loop: listen, think, speak. The agent captures audio, converts speech into text, passes that text to an LLM that decides what to say next, and renders the response as audio. The LLM is the "think" stage, and everything else in the system exists to feed it clean input and deliver its output fast.
There are two dominant architectures. The first is the sandwich pipeline: a speech-to-text engine, then the LLM, then a text-to-speech engine. It is modular, lets you swap each component independently, and is what most production deployments use today. The second is the end-to-end speech-to-speech model, a multimodal LLM that consumes audio directly and produces audio directly, skipping the text round trip entirely.
The sandwich approach gives you control, observability, and cost flexibility. The speech-to-speech approach gives you lower perceived latency and richer prosody, because the model hears tone and emotion rather than just words. Most teams start with the sandwich and evaluate speech-to-speech once the conversational logic is proven.

LLM Examples for Voice Agent: Real-World Use Cases

Where are developers actually shipping LLM-driven voice agents? Three patterns dominate: support call centers, personal productivity assistants, and multimodal agents that speak audio natively. Each stresses a different part of the pipeline.

Customer-Support Call Center

A support voice agent transcribes the caller's issue, uses an LLM grounded in company knowledge to suggest or execute a resolution, and replies via TTS in a calm, professional voice. The LLM handles intent classification, retrieval from a knowledge base, and response generation in one pass. Teams deploying these agents report deflection of a large share of tier-one calls, with human agents handling only escalations. The hard parts are policy compliance and latency, not the model itself.

Personal Productivity Assistant

A productivity assistant is hands-free by design. The user says "schedule a review with the design team next Tuesday" and the LLM interprets the request, calls the right calendar function, confirms the slot, and reads back the result. Here the LLM's job is less about generating prose and more about structured tool use: extracting entities, deciding which function to invoke, and confirming actions before executing them.

Example 1: Customer-Support Voice Agent

The support agent is the most instructive example because it exercises every part of the pipeline under real constraints: angry callers, sensitive data, and a hard latency budget. The flow is simple to describe. The caller speaks, turn detection decides when they finished, STT transcribes the utterance, the LLM generates a grounded response, and TTS renders it back, all inside roughly a second.

Choosing STT and TTS Providers

Pick STT and TTS providers on three criteria: latency, accuracy, and cost. For transcription, Deepgram's Nova models and OpenAI's Whisper-family APIs are popular for conversational audio, with Deepgram frequently cited for low latency on streaming input. For speech output, ElevenLabs leads on naturalness for customer-facing voices, while OpenAI TTS and AWS Polly offer lower cost at slightly lower expressiveness. A practical pattern is pairing a fast streaming STT with a mid-tier TTS, then upgrading the voice once the agent logic is stable. Independent comparisons on Artificial Analysis's Speech Arena are a better guide than vendor marketing pages.

Prompt Engineering for Support Context

The system prompt is where support agents succeed or fail. A good support prompt defines the agent's role, the scope of topics it may discuss, escalation rules, and the tone it must maintain. It should explicitly instruct the LLM to refuse actions outside policy, ask clarifying questions when intent is ambiguous, and keep responses short, because every extra sentence is extra TTS time the caller waits through. Grounding the prompt in retrieved knowledge-base content, rather than letting the model improvise, is what keeps answers policy-compliant.

Handling Sensitive Data

Support calls contain payment details, health information, and identity data. Before anything reaches the LLM, the transcript should pass through a redaction layer that masks card numbers and other regulated fields. Choose providers with data-retention controls, and prefer LLM endpoints that do not train on your traffic. If you operate in healthcare or finance, check whether your stack supports HIPAA or regional compliance before launch. VideoSDK's agent infrastructure supports end-to-end encryption on the media path, which covers the audio leg of the journey.

Example 2: Personal Productivity Voice Agent

The productivity assistant looks simpler than the support agent but has its own complexity: it acts. When an LLM can create calendar events or send emails, the conversation design must guarantee the right action happens on the right data. The pipeline is the same sandwich, but the LLM layer is wired to function tools rather than a knowledge base.

Integrating Calendar and Email APIs

The logical flow is: the user speaks a request, the LLM extracts the intent and entities (who, when, what), selects the matching function, and the function calls the calendar or email API. Confirmation matters here. A well-designed assistant reads back the proposed action, "Booking design review for Tuesday at 2 PM, confirm?", and only executes after the user agrees. This two-step confirm-then-act pattern prevents the classic failure mode of an LLM confidently booking the wrong meeting.

Managing Conversation State

Voice conversations are multi-turn, and the LLM needs memory across them. A lightweight session store keyed by the call or room identifier holds the running transcript, extracted entities, and pending actions. For longer-running assistants, a summary-based memory, where older turns are condensed into a compact context block, keeps token usage flat while preserving continuity. VideoSDK's Agent SDK provides session-scoped context management and memory components so you do not have to build this plumbing yourself.

Voice-Specific UX Tips

Voice UX differs from chat UX in three ways. Keep responses short, since callers cannot skim. Design for interruption, so when the user speaks over the agent, the pipeline stops TTS playback immediately and re-enters listening mode. And match pacing to the task: confirmations should be brisk, bad news should be delivered slowly. Turn detection and voice activity detection settings have a large effect on how natural the agent feels, and tuning them is as important as tuning the prompt.

Example 3: Multimodal Voice Agent with Direct Audio Output

Speech-to-speech models change the architecture. Instead of transcribing, reasoning, and synthesizing as separate steps, a multimodal LLM such as OpenAI's Realtime API or Google's Gemini Live consumes the raw audio stream and produces audio output directly. The model hears hesitation, tone, and emotion, and can respond with matching prosody, which makes the sandwich approach sound robotic by comparison.
Architecture Diagram
Prefer speech-to-speech when the experience is emotionally rich or latency-critical: companions, coaching, live tutoring. Prefer the sandwich when you need transcript-level observability, strict cost control, or the ability to swap components per deployment. Many teams run both, using the sandwich for tool-heavy workflows and a speech-to-speech model for free-form conversation. VideoSDK's agent pipeline supports both patterns, including real-time models from OpenAI, Google, and AWS, alongside the sandwich stack.

Production Considerations

Shipping a voice agent is a latency problem wrapped around an LLM problem. A conversational feel requires the full round trip, speech in to speech out, to complete in roughly one second. Budget that time across the pipeline: a few hundred milliseconds for STT, several hundred for LLM first token, and the rest for TTS first audio. Streaming every stage, so TTS starts speaking the first sentence while the LLM still generates the second, is the single biggest perceived-latency win.
Scaling means scaling LLM inference. Managed model endpoints handle this for you; self-hosted inference gives control but demands GPU capacity planning. Monitoring should track transcription accuracy on real audio, end-to-end latency percentiles, and interruption handling, not just uptime. Finally, plan fallbacks: when transcription confidence is low, the agent should ask the user to repeat rather than guess, and a human handoff path should always exist for support deployments.

Choosing the Right Stack

Your stack decision comes down to control versus speed to market. The table below compares the main options.
Decision Open-Source / Self-Hosted Managed Service Winner Condition
LLM reasoning Llama, Mistral on your GPUs OpenAI, Anthropic, Gemini APIs Open-source for data control; managed for speed
STT Whisper self-hosted, Deepgram Deepgram, AssemblyAI, OpenAI Managed for streaming latency
TTS Coqui-style open models ElevenLabs, OpenAI TTS, Cartesia Managed for voice quality
Orchestration Pipecat, VideoSDK Agent SDK VideoSDK Agent Cloud VideoSDK for rooms, telephony, and web delivery in one
Delivery WebRTC you build yourself VideoSDK rooms, SIP telephony VideoSDK unless you need fully custom media
The pragmatic choice for most teams is a managed LLM and STT/TTS layer, orchestrated by an open framework, delivered over managed real-time infrastructure. VideoSDK's AI Agent SDK covers the orchestration and delivery layers, and its Python SDK handles backend integration.

Definitions Glossary

STT (Speech-to-Text): The component that converts spoken audio into text, the first stage of the sandwich pipeline. Streaming STT engines transcribe while the user is still speaking, cutting total response latency.
TTS (Text-to-Speech): The component that renders the LLM's text response as spoken audio. Provider choice drives both the naturalness of the agent's voice and a large share of perceived latency.
Turn Detection: The mechanism that decides when the user has finished speaking and the agent should respond. Poor turn detection is the most common cause of unnatural-feeling voice agents.
Speech-to-Speech Model: A multimodal LLM that consumes audio input and produces audio output directly, skipping the transcription and synthesis stages of the sandwich architecture.
Agent Worker: The Python process that runs a VideoSDK AI agent, managing its session lifecycle inside a VideoSDK room and connecting it to users over web, mobile, or phone.

Key Takeaways

  • The sandwich architecture (STT, LLM, TTS) remains the default for production voice agents because it is modular, observable, and cost-flexible.
  • Speech-to-speech models win on prosody and perceived latency, and suit emotionally rich or free-form conversations.
  • Latency is the defining production constraint: stream every pipeline stage and target a sub-second speech-in to speech-out round trip.
  • Grounding, redaction, and confirm-then-act patterns matter more than model choice for support and productivity agents handling sensitive data.
  • VideoSDK's open-source Agent SDK provides the orchestration, session memory, and real-time delivery layers so you can focus on the LLM logic.

Conclusion

These LLM examples for voice agents, support, productivity, and multimodal, share one lesson: the LLM is the easy part. The craft is in the pipeline around it, the STT and TTS choices, turn detection, latency budgeting, and safety patterns. Start with the sandwich architecture, prove the conversation logic, then evaluate speech-to-speech where the experience demands it. You can explore the full pipeline in the VideoSDK AI Agents docs, browse code samples, or grab the free tier at app.videosdk.live/login and build today. What are you building with VideoSDK? Drop a comment, I'd love to hear what kind of voice agent use case you're working on.

Free $20 Balance for AI Voice Agents & Video Calls

FAQ