An LLM for speech-to-text replaces the traditional acoustic-plus-language-model pipeline with a single neural decoder that transcribes spoken audio into text. These models leverage LLM contextual understanding and multilingual capabilities for higher accuracy on noisy or code-switched audio. VideoSDK integrates LLM-powered STT into its AI voice agent pipeline for real-time transcription, and developers can self-host open-source options like FunASR, IBM Granite Speech, and FireRedASR2 through vLLM.
Speech recognition has spent decades trapped behind a two-stage architecture: an acoustic model that converts sound into phonemes, followed by a separate language model that cleans up the output. That approach worked, but it left a persistent gap. The acoustic model had no idea what words meant, so it could not use context to fix errors. The language model fixed some errors but could not go back and re-listen to the audio.
Large language models are collapsing that gap. By training a single transformer to handle both acoustic understanding and text generation, researchers are building speech-to-text systems that understand context the same way a text LLM understands a paragraph. The result is higher accuracy on accented speech, noisy environments, code-switching between languages, and domain-specific jargon.
For developers building transcription services, AI voice agents, or real-time captioning systems, this shift matters because it changes what you deploy and how you deploy it. This guide covers the architecture, the leading open-source models, deployment patterns, and the performance trade-offs you need to know.
What is an LLM for Speech-to-Text?
An LLM for speech-to-text is defined as a large language model that has been adapted or fine-tuned to accept audio input and produce text transcripts as output. Instead of relying on a dedicated acoustic model that maps audio features to phonetic units, the LLM itself serves as the decoder, processing audio representations directly through its transformer architecture.
LLM-based speech-to-text works by first converting raw audio into a sequence of discrete tokens or continuous embeddings that the LLM can process. These audio representations are fed into the model alongside any text prompts, and the model autoregressively generates the transcript token by token, just as a text LLM generates a response to a text prompt.
This approach differs fundamentally from classic automatic speech recognition (ASR) systems, which chain together separate components for acoustic modeling, pronunciation modeling, and language modeling. In an LLM-based system, the language understanding is baked into the decoder itself, which means the model can leverage semantic context, world knowledge, and conversational flow to disambiguate unclear speech.
VideoSDK's AI voice agent architecture uses a similar pipeline concept, where speech-to-text is the first stage in a chain that feeds into an LLM for reasoning and then a TTS engine for response generation.
How LLM-based STT Differs from Traditional ASR
Traditional ASR systems follow a modular pipeline: audio features go into an acoustic model that predicts phoneme or character probabilities, those predictions pass through a pronunciation lexicon, and a separate n-gram or neural language model rescoring step cleans up the final text. Each component is trained independently, and errors compound across stages.
LLM-based speech-to-text replaces that entire chain with a single end-to-end model. The audio encoder converts sound into representations the LLM can consume, and the LLM decoder generates text directly. There is no separate language model rescoring step because the LLM already has deep language understanding built in.
The biggest benefit is contextual accuracy. A traditional ASR system hearing "I need to schedule a meeting with Dr. Chen next Tuesday" might transcribe "Chen" as "chin" if the acoustic signal is noisy. An LLM-based system uses its understanding of names, scheduling context, and conversational patterns to get it right.
The trade-off is compute. LLM decoders are larger and more expensive to run than dedicated acoustic models. A traditional ASR model might have 100 million parameters, while an LLM-based STT system might have 2 billion or more. That means higher GPU memory requirements, more latency per inference, and more careful deployment planning.
Core Components of an LLM-Based Speech-to-Text System
Every LLM-based speech-to-text system shares four core components, and understanding each one is essential for choosing models, tuning performance, and debugging transcription errors in production.
Audio Front-End
The audio front-end handles everything that happens before audio reaches the LLM. This includes feature extraction (converting raw waveforms into mel-spectrograms or other time-frequency representations), voice activity detection to identify speech segments and discard silence, and language identification to route multilingual audio to the right model or prompt configuration. A well-tuned front-end dramatically reduces the load on the LLM decoder by ensuring it only processes actual speech.
LLM Decoder
The decoder is the heart of the system. Autoregressive decoders generate text one token at a time, which produces high-quality output but introduces sequential latency. Non-autoregressive decoders generate multiple tokens in parallel, trading some accuracy for lower latency. Most current LLM-based STT models use autoregressive decoding because the quality advantage is significant, but non-autoregressive variants like IBM Granite Speech 2b-nar are pushing the boundaries of what parallel decoding can achieve.
Tokenizer and Prompt Engineering
Audio tokenization converts continuous audio features into discrete tokens the LLM can process. Some models use dedicated audio tokenizers that compress audio into a fixed vocabulary, while others use continuous embeddings injected directly into the transformer. Prompt engineering also plays a role: you can prepend text context (like domain vocabulary or speaker names) to guide the model toward more accurate domain-specific transcription.
Post-Processing
Raw LLM output often lacks punctuation, speaker labels, and precise timestamps. Post-processing modules add these elements. Some models, like FireRedASR2, bundle punctuation and timestamp prediction into the model itself. Others require external post-processing for speaker diarization, which identifies who spoke when in multi-speaker audio.
Popular Open-Source LLM-Powered STT Models
The open-source landscape for LLM-based speech-to-text has expanded rapidly. Here are the models developers should know about in 2026.
FunASR with vLLM Integration
FunASR, developed by Alibaba DAMO Academy, is a comprehensive speech recognition toolkit that combines acoustic models, language models, and post-processing into a unified framework. When integrated with vLLM for inference acceleration, FunASR achieves significantly higher throughput than standard PyTorch execution. It supports Chinese and English with strong performance on accented and noisy audio. The toolkit includes built-in voice activity detection, punctuation restoration, and timestamp prediction, making it a practical all-in-one solution for batch transcription workloads.
Qwen3-ASR Family
The Qwen3-ASR models extend Alibaba's Qwen LLM family into speech recognition. These models leverage the Qwen transformer architecture, which has demonstrated strong multilingual text capabilities, and adapt it for audio input. The Qwen3-ASR family supports multiple languages and benefits from the broad world knowledge encoded in the base Qwen models. This makes them particularly effective for domain-specific transcription where the model needs to understand technical vocabulary, proper nouns, or industry jargon.
IBM Granite Speech (2b, 2b-plus, 2b-nar)
IBM has open-sourced several Granite Speech variants under its Granite model family. The 2b variant is a 2-billion-parameter autoregressive model that balances accuracy and inference cost. The 2b-plus variant increases capacity for higher accuracy on challenging audio. The 2b-nar variant is non-autoregressive, generating multiple output tokens in parallel for lower latency at the cost of some transcription accuracy. All three are designed for enterprise deployment with permissive licensing, making them attractive for organizations that need self-hosted STT without vendor lock-in.
FireRedASR2
FireRedASR2 is an all-in-one ASR system that integrates voice activity detection, language identification, speech recognition, and punctuation restoration into a single model. This eliminates the need for a separate audio front-end pipeline, which simplifies deployment considerably. FireRedASR2-LLM variants combine the ASR model with an LLM decoder for enhanced contextual understanding. The integrated design reduces the number of components you need to maintain and debug in production.
vLLM Native Speech-to-Text Support
vLLM, originally built for text LLM inference, has expanded to support multimodal models including speech-to-text. This means you can serve LLM-based STT models using the same inference infrastructure you already use for text generation. vLLM's PagedAttention memory management and continuous batching apply to audio token processing as well, which improves throughput when handling multiple concurrent transcription streams.
Choosing the Right Model for Your Use-Case
Selecting the right LLM for speech-to-text depends on five factors: latency requirements, accuracy needs, language coverage, hardware budget, and licensing constraints. No single model wins on all dimensions.
| Model | Best For | Languages | Latency Profile | License |
|---|---|---|---|---|
| FunASR + vLLM | High-throughput batch transcription | Chinese, English | Medium (autoregressive) | Apache 2.0 |
| Qwen3-ASR | Multilingual domain-specific transcription | Multi-language | Medium-high (larger model) | Apache 2.0 |
| IBM Granite Speech 2b | Balanced enterprise deployment | English, multilingual | Low-medium (2B params) | Apache 2.0 |
| IBM Granite Speech 2b-nar | Real-time low-latency transcription | English, multilingual | Low (non-autoregressive) | Apache 2.0 |
| FireRedASR2 | Simplified all-in-one deployment | Chinese, English | Medium (integrated pipeline) | Open source |
If your priority is real-time transcription for a live voice agent, the non-autoregressive Granite Speech 2b-nar variant offers the lowest latency. If you need maximum accuracy for offline batch processing of recorded meetings, Qwen3-ASR or FunASR with vLLM acceleration will serve you better. If you want the simplest deployment with the fewest moving parts, FireRedASR2's integrated pipeline eliminates several components you would otherwise need to wire together.
For developers building AI voice agents with VideoSDK, the STT model you choose feeds directly into the agent's response latency. VideoSDK's Python SDK supports connecting custom STT pipelines to VideoSDK rooms, so you can swap models without changing your agent's core logic.
Deployment Patterns
LLM-based speech-to-text systems can be deployed in three primary patterns, each suited to different latency and throughput requirements.
Offline Batch Transcription
Batch transcription processes complete audio files after recording. This pattern is ideal for transcribing podcasts, meeting recordings, voicemails, or call center logs where real-time output is not required. You can use larger models with higher accuracy because latency is not a constraint. Batch processing also allows you to queue jobs and maximize GPU utilization through large batch sizes. FunASR with vLLM is well-suited here because vLLM's continuous batching keeps the GPU saturated as new jobs arrive.
Streaming Inference with Chunked Audio
Streaming inference processes audio in fixed-size chunks (typically 200 to 500 milliseconds) as they arrive. The model produces partial transcripts that are refined as more audio context becomes available. This pattern is used for live captioning and real-time transcription where users need to see text appear as someone speaks. The challenge is handling chunk boundaries: if a word spans two chunks, the model needs context from both to transcribe it correctly. Most streaming implementations use a sliding window with overlap to mitigate this.
Real-Time WebSocket Service
For interactive applications like AI voice agents, a WebSocket-based service maintains a persistent connection between the client and the inference server. Audio frames stream in one direction, and transcript tokens stream back in the other. This is the pattern VideoSDK uses in its AI voice agent architecture, where the agent worker maintains a session inside a VideoSDK room and processes speech in real time. vLLM, FunASR, and FireRedASR2-LLM all support this deployment pattern with appropriate configuration.

Performance Considerations
LLM-based speech-to-text introduces performance characteristics that differ from traditional ASR in several important ways.
Latency sources in an LLM STT pipeline include audio feature extraction time, audio tokenization overhead, autoregressive token generation (the dominant factor), and post-processing. The autoregressive decoding step scales linearly with output length, so longer utterances take proportionally more time. GPU memory bandwidth and KV cache efficiency directly affect how fast each token is generated.
Memory footprint is significantly larger than traditional ASR. A 2-billion-parameter model in FP16 precision requires roughly 4 GB of GPU memory for weights alone, plus additional memory for the KV cache during inference. The KV cache grows with sequence length, so transcribing a 30-minute audio file requires more memory than transcribing a 30-second clip. vLLM's PagedAttention helps manage this by allocating KV cache memory in pages rather than contiguous blocks, reducing fragmentation.
Batch size scaling determines how many concurrent transcription requests a single GPU can handle. Larger batch sizes improve throughput but increase memory consumption. The optimal batch size depends on your GPU's memory capacity, the model size, and the average audio length per request.
Accuracy versus speed trade-offs are the central tension. Autoregressive models produce the most accurate transcripts but are slower. Non-autoregressive variants like Granite Speech 2b-nar sacrifice some accuracy for faster generation. Quantization (INT8 or INT4) reduces memory and speeds up inference but may degrade accuracy on challenging audio.
According to Artificial Analysis Speech Arena benchmarks, the gap between the best LLM-based STT models and traditional ASR systems is most pronounced on noisy audio, accented speech, and code-switched conversations. On clean studio-quality audio, the difference narrows considerably, which means the extra compute cost of LLM-based STT may not be justified for all use cases.
Practical Implementation Steps
Building an LLM-based speech-to-text pipeline involves six main steps. Here is a narrative walkthrough of the process.
Step 1: Select Your Model and Download Weights
Choose a model based on your latency, accuracy, and language requirements using the decision matrix above. Download the model weights from the official repository (Hugging Face, ModelScope, or IBM's GitHub). Verify the checksums to ensure the weights are not corrupted. Pay attention to the model's supported sample rate, which is typically 16 kHz for speech models.
Step 2: Set Up Your Environment
Install Python 3.10 or later, then install the required libraries. For vLLM-based deployment, you need the vLLM package along with CUDA-compatible GPU drivers. For FunASR, install the FunASR toolkit and its dependencies. Ensure your GPU has sufficient memory for the model size you chose. A 2-billion-parameter model in FP16 needs at least 8 GB of GPU memory for comfortable inference with a reasonable batch size.
Step 3: Prepare Audio Preprocessing
Normalize your audio to the model's expected sample rate (usually 16 kHz, mono channel). Apply voice activity detection to segment audio into speech chunks and discard silence. If your model supports language identification, configure it to auto-detect or specify the language explicitly. Consistent preprocessing is critical: feeding audio at the wrong sample rate or with unexpected channel configurations is one of the most common causes of poor transcription quality.
Step 4: Launch the Inference Server
Start the inference server using vLLM or your chosen framework. Key configuration parameters include tensor-parallel size (how many GPUs to split the model across), GPU memory utilization (what fraction of GPU memory to allocate), and maximum sequence length. For streaming deployments, configure the chunk size and overlap window. The server exposes an API endpoint (REST or WebSocket) that your application will call.
Step 5: Consume the API from Your Application
Your application sends audio data to the inference server's API endpoint. For batch transcription, send the complete audio file and receive the full transcript. For streaming transcription, open a WebSocket connection and send audio chunks as they arrive, receiving partial transcripts in return. Handle connection errors, reconnection logic, and backpressure if the server cannot keep up with the audio stream rate.
Step 6: Post-Process the Output
Apply punctuation restoration if your model does not produce it natively. Add speaker labels using a separate diarization model if your use case involves multi-speaker audio. Align timestamps to the transcript for applications that need word-level or sentence-level timing precision. If you are feeding the transcript into an LLM for downstream reasoning (as in a voice agent pipeline), format the output to match what the downstream LLM expects.
Common Pitfalls and Troubleshooting
Token-size overflow on long audio. LLM decoders have a maximum context length. If your audio produces more tokens than the model's limit, the decoder will truncate or error. Mitigate this by chunking long audio into segments that fit within the context window, with overlap to handle words that span chunk boundaries.
Language-ID mismatches. If your language identification module guesses the wrong language, the model may produce garbage output or transcribe in the wrong script. For known-language deployments, disable auto-detection and specify the language explicitly. For multilingual deployments, ensure your language ID model covers all languages your users might speak.
GPU out-of-memory on large batch sizes. When concurrent requests spike, the KV cache can exceed available GPU memory. Configure a maximum batch size and queue excess requests rather than crashing. vLLM's PagedAttention helps, but it does not eliminate the limit. Monitor GPU memory usage and set alerts for approaching capacity.
Inconsistent timestamps across streaming chunks. Streaming transcription produces partial transcripts that may have inconsistent timestamp references when stitched together. Use a global time reference (the session start time) and compute relative offsets for each chunk. Some models provide built-in timestamp alignment, but verify that the timestamps are consistent across chunk boundaries.
Future Trends
The trajectory of LLM-based speech-to-text points toward several developments that will reshape how developers build transcription pipelines.
Multimodal LLMs that handle audio, video, and text in a single model are already emerging. These models can transcribe speech while also understanding visual context, which opens up applications like transcribing meetings with slide awareness or transcribing video content with scene understanding.
On-device quantization is making it feasible to run small LLM-based STT models on mobile devices and edge hardware. INT4 quantization of a 2-billion-parameter model can reduce memory requirements enough to run on modern smartphones, enabling offline transcription without cloud dependency.
Real-time translation is converging with STT. Instead of transcribing speech and then translating the text, future models will generate translated transcripts directly from audio input, reducing latency for cross-lingual communication.
Integration with AI voice agents is perhaps the most impactful trend. VideoSDK's AI voice agent platform already supports connecting custom STT pipelines to agent sessions. As LLM-based STT models become faster and more accurate, they will replace traditional ASR as the default input layer for voice agents, enabling agents that understand context, nuance, and multilingual input with far greater fidelity.
Definitions Glossary
LLM for Speech-to-Text: A large language model adapted to accept audio input and generate text transcripts, replacing the traditional multi-stage ASR pipeline with a single neural decoder.
Autoregressive Decoder: A decoding strategy where the model generates output tokens one at a time, each token conditioned on all previously generated tokens, producing high-quality text but introducing sequential latency.
Non-Autoregressive Decoder: A decoding strategy where multiple output tokens are generated in parallel, reducing latency at the cost of some transcription accuracy compared to autoregressive decoding.
Audio Tokenizer: A component that converts continuous audio features (like mel-spectrograms) into discrete tokens or embeddings that an LLM transformer can process as input.
Voice Activity Detection (VAD): A preprocessing step that identifies segments of an audio stream containing speech and discards silence, reducing the computational load on downstream models.
Key Takeaways
- LLM-based speech-to-text replaces the traditional acoustic-model-plus-language-model pipeline with a single neural decoder that leverages contextual understanding for higher accuracy on challenging audio.
- Open-source models like FunASR, Qwen3-ASR, IBM Granite Speech, and FireRedASR2 give developers production-ready options without vendor lock-in, each with distinct strengths in latency, accuracy, and language coverage.
- Deployment patterns range from offline batch transcription (maximum accuracy, no latency constraint) to real-time WebSocket streaming (lowest latency, required for interactive voice agents).
- Performance trade-offs center on the tension between autoregressive quality and non-autoregressive speed, with GPU memory and KV cache management being the primary infrastructure constraints.
- VideoSDK integrates LLM-powered STT into its AI voice agent pipeline, allowing developers to connect custom speech-to-text models to real-time agent sessions through the Python SDK.
Conclusion
LLMs are reshaping speech-to-text by collapsing a multi-stage pipeline into a single context-aware decoder. The open-source landscape in 2026 offers real choices: FunASR for throughput, Qwen3-ASR for multilingual domain accuracy, IBM Granite Speech for balanced enterprise deployment, and FireRedASR2 for simplified all-in-one pipelines. The right choice depends on your latency budget, accuracy requirements, language coverage, and hardware constraints. For developers building AI voice agents, VideoSDK's agent platform lets you plug these models directly into real-time sessions. You can start building for free at app.videosdk.live/login. What are you building with LLM-based speech-to-text? Drop a comment, I'd love to hear what transcription use case you're working on.
Free $20 Balance for AI Voice Agents & Video Calls
FAQ
