The Voice Agent API is VideoSDK's real-time, low-latency interface for building AI-powered phone or web agents. It combines speech-to-text, a large language model, and text-to-speech in a single pipeline, and you can control it via REST endpoints and WebSocket streams. Get started by creating a token, defining an agent, and connecting a client.
Voice-first interfaces are no longer a novelty. Developers building customer support, telehealth, or commerce applications now face a clear expectation from users: they want to speak to systems and get natural, immediate responses. Building this experience from scratch means wrestling with raw audio streams, managing unpredictable network conditions, and stitching together disparate speech-to-text, large language model, and text-to-speech services.
The difficulty does not end at assembly. Latency is the silent killer of voice agent experiences. If the gap between a user finishing a sentence and the agent responding exceeds a few hundred milliseconds, the conversation feels broken. Scaling these systems to handle thousands of concurrent calls adds another layer of complexity, often requiring dedicated infrastructure just to manage media routing and session state.
This guide covers everything you need to know about the Voice Agent API. We will break down the architecture, explain how to authenticate and set up your first agent, compare pipeline architectures, and share production-ready considerations for scaling, security, and monitoring. By the end, you will understand how to build and deploy low-latency AI voice agents using VideoSDK.
What Is the Voice Agent API?
A voice agent API is a developer interface that enables real-time, AI-driven voice conversations between a human and a software agent. Instead of manually wiring separate audio ingestion, transcription, language processing, and speech synthesis services, a voice agent API abstracts these layers into a unified, managed pipeline.
VideoSDK implements the voice agent API through a structured pipeline that connects speech-to-text (STT), a large language model (LLM), and text-to-speech (TTS) components. The API manages the session lifecycle, media transport, and orchestration of these components. Developers define the agent's behavior, choose the AI providers, and connect a client endpoint. VideoSDK handles the rest.
The primary differentiators of VideoSDK's approach are sub-150-millisecond latency, built-in WebRTC transport for sub-second media delivery, and multi-provider flexibility. You are not locked into a single STT or TTS vendor. You can mix and match providers like Deepgram, OpenAI, and ElevenLabs to optimize for accuracy, voice quality, or cost.
Core Components of VideoSDK's Voice Agent API
VideoSDK's voice agent architecture is modular by design. You interact with five core components to bring an agent to life.
STT Layer
The speech-to-text layer is the entry point of the pipeline. It captures incoming audio from the user's device or phone line and transcribes it into text in real time. VideoSDK supports multiple STT providers, including OpenAI Whisper, Deepgram, and Google Cloud STT.
When you configure an agent, you select the STT provider that best fits your use case. Deepgram is often chosen for its low-latency streaming transcription, while Whisper might be preferred for its robustness across diverse accents. The API allows you to specify model parameters, language hints, and endpointing thresholds to control how quickly the system detects the end of a user's speech.
The STT layer also incorporates Voice Activity Detection (VAD) to distinguish between speech and background noise. This is crucial for phone calls where line noise is common. The API allows you to configure endpointing sensitivity, which determines how long the system waits after a user stops speaking before finalizing the transcription. Tuning this parameter is essential for balancing natural conversation flow against premature interruptions.
LLM Layer
Once the user's speech is transcribed, the text is passed to the large language model layer. This is where the agent "thinks." VideoSDK integrates with real-time models like OpenAI Realtime, Google Gemini, and AWS Nova Sonic, as well as standard LLM providers like Anthropic Claude and Cerebras.
You define the agent's persona and instructions through a system prompt. The LLM layer manages the context window, keeping track of the conversation history so the agent can reference earlier statements. Context management is handled automatically by the API, but you can configure the context window size. A larger context window allows the agent to remember more of the conversation, which is useful for complex support scenarios.
If your agent needs to perform actions, like checking an order status or booking a calendar slot, you can define function tools that the LLM can call during the conversation. When the LLM determines it needs to call a function, it emits a structured request. Your backend receives this request, executes the logic, and returns the result to the LLM to inform its response.
TTS Layer
The text-to-speech layer takes the LLM's text response and converts it back into audio. VideoSDK supports streaming TTS output, meaning the agent starts speaking the first words of a response before the LLM has finished generating the entire sentence. This streaming approach is critical for achieving low end-to-end latency.
Supported TTS providers include ElevenLabs, Cartesia, OpenAI TTS, and AWS Polly. You can select different voices, adjust speaking rates, and in some cases use voice cloning to create a custom brand voice. Some TTS providers supported by VideoSDK offer TTS caching. If your agent frequently says the same phrases, like a standard greeting or a compliance disclaimer, caching can significantly reduce latency and API costs. The API manages the audio encoding to match the client's capabilities, whether it is Opus for WebRTC or standard PCM for telephony.
Agent Management Endpoints
You manage your agents through a set of REST endpoints. These endpoints allow you to create an agent, update its configuration, deploy it, and retrieve its details. When you create an agent, you specify the STT, LLM, and TTS providers, the system prompt, and any function tools or knowledge bases. This configuration is stored and associated with a unique agent ID.
Connection Channels
The final component is the connection channel, which determines how the user connects to the agent. For web-based applications, clients connect via WebSocket to establish a bidirectional audio stream. For traditional phone networks, VideoSDK uses a SIP gateway to bridge PSTN calls into a WebRTC room. This flexibility means you can build a single agent and deploy it across web, mobile, and telephony channels.
Setting Up Authentication
Security is non-negotiable when building voice agents that might handle sensitive user data or execute transactions. VideoSDK enforces token-based authentication for all voice agent API interactions.
You must never expose your API secret on the client side. The standard flow involves generating a token server-side using your API key and secret. This token is a short-lived JWT that authenticates a participant's access to a VideoSDK room or agent session.
To set this up, you first retrieve your API key and secret from the VideoSDK dashboard. You then create a token generation endpoint on your backend. When your frontend or telephony system needs to connect to an agent, it requests a token from your backend. You can scope this token to specific actions, limiting what a client can do within a session.
For a detailed walkthrough of generating and scoping tokens, refer to the VideoSDK token generation guide.
Building Your First Voice Agent
Building a voice agent with VideoSDK is a sequential process that moves from configuration to client connection. You do not need to write complex audio processing code. The API handles the media routing, and you focus on defining the agent's behavior.
Step 1: Create an Agent
The first step is to define your agent using the REST Create Agent endpoint. You provide a name for the agent, select your STT, LLM, and TTS providers, and write the system prompt that dictates the agent's behavior. For example, if you are building a customer support bot, your prompt would instruct the agent to greet the caller, ask for their account number, and offer help with billing or technical issues. The API returns an agent ID that you use in subsequent steps.
Step 2: Attach a Knowledge Base
If your agent needs to answer questions based on specific documents, you can attach a knowledge base to enable Retrieval-Augmented Generation (RAG). You upload your documents (PDFs, text files, or URLs) through the API. VideoSDK processes and indexes this content. When the agent receives a question, it searches the knowledge base for relevant context and feeds that context to the LLM to generate an accurate, grounded response. This is essential for agents that need to reference company policies, product manuals, or historical data.
Step 3: Generate a Connection Token
Before a client can connect to the agent, you need to generate a connection token. This is the JWT we discussed in the authentication section. Your backend creates a token scoped to allow the client to join the agent session. This token ensures that only authorized users can interact with your agent and that their permissions are limited to the necessary actions.
Step 4: Connect a Client
The final step is connecting the user. If the user is on a web browser, your frontend application uses the token to establish a WebSocket connection to the VideoSDK agent worker. If the user is calling from a phone, the SIP gateway handles the connection. Once connected, the user's audio is streamed to the STT layer, and the agent's audio is streamed back. The conversation begins.
Here is a visual representation of the end-to-end flow, from client connection to audio response:
A common pitfall during this process is mismatched provider configurations. Ensure that the STT and TTS providers you select in the Create Agent step are actually enabled in your VideoSDK account. Another frequent issue is missing token scopes. If your token does not have permission to join an agent session, the connection will fail silently or with an ambiguous error.
Choosing the Right Architecture
When building a voice agent, you have two primary architectural paths: speech-to-speech and a chained pipeline. Your choice impacts latency, transcript availability, and tool integration capabilities.
Speech-to-Speech (Live Audio)
In a speech-to-speech architecture, the audio goes directly from STT to LLM to TTS without intermediate text storage driving the flow. This is often powered by real-time multimodal models like OpenAI Realtime or Google Gemini Live. The primary advantage is ultra-low latency. The model processes audio tokens directly, reducing the overhead of text conversion. This architecture is ideal for free-flowing conversational agents where natural back-and-forth is more important than perfect transcripts.
Chained Pipeline (STT to LLM to TTS)
A chained pipeline explicitly separates the steps. STT produces text, the LLM processes that text and generates a text response, and TTS converts that text to audio. This architecture is slightly slower due to the sequential nature of the steps, but it offers greater control. You can inspect and modify the transcript before it reaches the LLM, implement deterministic conversational flows using a graph-based approach, and easily log full transcripts for compliance. This is the better choice for structured workflows like appointment booking or loan applications.
| Architecture | Latency | Transcript Control | Best For |
|---|---|---|---|
| Speech-to-Speech | Lowest | Limited | Free-flowing chat, casual assistants |
| Chained Pipeline | Moderate | Full | Compliance, structured flows, tool use |
Production-Ready Considerations
Moving a voice agent from a prototype to a production deployment requires careful planning around scaling, security, and observability.
Scaling and Concurrency
VideoSDK is built on a distributed media processing infrastructure designed to handle thousands of concurrent voice sessions. When you deploy an agent, VideoSDK automatically provisions the necessary agent workers to handle the load. The platform manages load balancing across regions to ensure users connect to the geographically closest server, minimizing latency.
VideoSDK's infrastructure uses an auto-scaling mechanism for agent workers. When a new session is initiated, the platform provisions a worker process to handle the STT, LLM, and TTS orchestration for that specific call. This isolation ensures that a spike in traffic or a poorly behaving agent on one call does not degrade the performance of others. Load balancers distribute these workers across available server clusters. You do not need to manually spin up additional servers as your call volume grows. However, you should monitor your usage through the dashboard and set up alerts for concurrency limits.
Security and Compliance
Voice agents often process sensitive information, from personal health data to financial details. VideoSDK provides several layers of security. All media streams are encrypted in transit using SRTP. For sessions requiring the highest level of privacy, such as telehealth consultations, you can enable End-to-End Encryption (E2EE). This ensures that the audio streams are encrypted on the client device and can only be decrypted by the intended recipient, meaning not even VideoSDK's servers can access the raw audio. This is a critical feature for HIPAA compliance.
Token expiry is a critical security mechanism. Always set short expiry times on your connection tokens, typically ranging from a few minutes to an hour, depending on the expected call duration. Implement role-based access control to ensure that clients can only perform actions they are explicitly authorized for. If you are handling health data, ensure your deployment is configured to meet HIPAA or GDPR requirements.
Monitoring and Analytics
You cannot improve what you cannot measure. VideoSDK provides Agent Analytics endpoints that give you visibility into your agent's performance. You can retrieve call metrics, including duration, latency breakdowns, and error rates. If a call fails, the analytics API helps you trace the failure back to a specific pipeline component, like a TTS provider timeout.
You can also retrieve full conversation transcripts. This is invaluable for quality assurance, debugging agent behavior, and fine-tuning your system prompts. VideoSDK can send webhook events for various agent lifecycle stages. For example, you can receive a webhook when a session starts, when it ends, or when a specific function tool is invoked. This allows you to build reactive systems around your voice agents. You can trigger a CRM update the moment a support call concludes, or send a transcript to your data warehouse for long-term storage and analysis.
Deployment Models
VideoSDK offers flexibility in how you deploy your voice agents. The standard cloud-hosted API is the fastest way to get started and is managed entirely by VideoSDK. This is suitable for most applications.
For privacy-critical use cases or organizations with strict data residency requirements, VideoSDK supports self-hosted deployment options. You can run the agent infrastructure on your own cloud account or on-premise servers using Docker or Kubernetes. This ensures that audio data never leaves your controlled environment.
For more details on managing agents and retrieving analytics, see the VideoSDK REST API reference.
Real-World Use Cases
The Voice Agent API is versatile enough to power a wide range of applications across different industries.
Customer Support Phone Bots
Businesses receive high volumes of repetitive support calls. A voice agent can handle tier-one support, answering FAQs, resetting passwords, and routing complex issues to human agents. The agent connects via the SIP gateway, greets the caller, and uses a knowledge base to resolve issues. VideoSDK's low latency ensures the caller does not experience awkward pauses, making the interaction feel like a natural conversation.
Interactive Voice-Enabled Shopping Assistants
E-commerce platforms can deploy web-based voice agents to help customers find products. A user can say, "I am looking for running shoes under $100," and the agent can query the product catalog, describe options, and even guide the user through checkout. The multi-modal capabilities of the API allow the agent to simultaneously display visual elements on the screen while speaking.
Tele-Health Appointment Triage
Healthcare providers can use voice agents to handle appointment scheduling and basic triage. When a patient calls, the agent can ask about their symptoms, schedule an appointment with the appropriate specialist, and send confirmation details. The chained pipeline architecture is ideal here, as it allows for deterministic flows and ensures all conversations are transcribed and stored for medical compliance.
Automated Outbound Reminder Calls
Clinics and service businesses can use the voice agent API to make outbound calls. The agent can call a patient to remind them of an upcoming appointment or call a customer to notify them that their order is ready for pickup. The agent can handle rescheduling on the spot if the user is unavailable at the proposed time. This reduces no-shows and frees up administrative staff.
Definitions Glossary
Voice Agent API: VideoSDK's unified endpoint for AI-driven voice interactions, managing the STT, LLM, and TTS pipeline.
Turn Detection: The mechanism that decides when a user has finished speaking, allowing the agent to begin its response.
RAG (Retrieval-Augmented Generation): A technique that adds external documents to an LLM's context to generate grounded, accurate responses.
SIP Gateway: A bridge that connects traditional PSTN phone numbers to a WebRTC room, enabling phone-to-agent calls.
Agent Token: A short-lived JWT that authorizes a client to join and interact within a voice agent session.
Key Takeaways
- VideoSDK provides a single, low-latency Voice Agent API that removes the need to stitch separate STT, LLM, and TTS services together manually.
- Authentication is strictly token-based; you must scope tokens to the exact actions a client needs to perform.
- Choose speech-to-speech architecture for real-time, casual conversations, and chained pipelines for deterministic, compliance-heavy workflows.
- Production deployments require proactive attention to scaling, security configurations, and analytics monitoring.
- The API is versatile enough to work for both web-based assistants and traditional PSTN phone agents.
Conclusion
Building a production-ready voice agent does not have to mean rebuilding media infrastructure and stitching together disparate AI services. VideoSDK's Voice Agent API gives you a unified, low-latency pipeline to deploy intelligent voice interactions across web and telephony channels. By handling the complexities of WebRTC transport, turn detection, and provider orchestration, VideoSDK lets you focus on defining your agent's logic and user experience. Start building your voice agent today by signing up for a free account at app.videosdk.live/login. If you have questions or want to share what you are building, join the VideoSDK Discord community to connect with other developers.
Free $20 Balance for AI Voice Agents & Video Calls
FAQ
