Flutter text to speech lets your app convert written text into spoken audio using the fluttertts plugin, which bridges Flutter and native platform TTS engines on Android, iOS, Web, macOS, and Windows. You control speech rate, pitch, volume, language, and voice selection through a single Dart API. For apps that also need real-time voice or video, VideoSDK's Flutter SDK pairs naturally with TTS for richer communication experiences. Start with the [fluttertts package on pub.dev](https://pub.dev/packages/flutter_tts) and follow the setup steps below.
Imagine building a navigation app for visually impaired users. Every turn instruction needs to be spoken aloud, clearly and at the right moment. Or think about a language-learning app where pronunciation is the whole point. In both cases, your Flutter app needs to talk. That is where Flutter text to speech comes in.
Adding speech synthesis to a Flutter app used to mean writing separate native code for Android and iOS. The flutter_tts plugin changed that by wrapping both platforms behind one Dart API. In this guide, you will learn what the plugin does, how to set it up across platforms, how to manage voices and languages, how to control speech parameters, and how to debug the most common issues developers hit in production.
What Is Flutter Text to Speech?
Flutter text to speech is defined as the ability to convert written text into spoken audio within a Flutter application. The mechanism relies on a TTS plugin that acts as a bridge between your Dart code and the native speech synthesis engines built into Android, iOS, and other supported platforms.
Flutter itself does not ship with a built-in speech synthesis engine. Instead, developers use community-maintained packages that expose native TTS capabilities through a Dart-facing API. The most widely adopted package is flutter_tts, which has been the go-to solution for Flutter speech synthesis for years. It supports Android, iOS, Web, macOS, and Windows, making it one of the most platform-complete TTS packages available.
The flutter_tts plugin works by sending your text string to the platform's native TTS engine. On Android, that engine is typically the Android TextToSpeech framework. On iOS, it uses the AVSpeechSynthesizer class from AVFoundation. On Web, it leverages the browser's SpeechSynthesis API. Your Flutter code never directly touches these native APIs. The plugin handles all the marshalling, and you receive callbacks for speech start, completion, and errors through Dart streams.
For developers building apps that combine TTS with real-time communication, VideoSDK's Flutter SDK provides video and audio calling capabilities that can complement speech synthesis features in accessibility, education, and assistant apps.
Core Features of the Flutter TTS Plugin
The flutter_tts plugin gives you a broad set of capabilities that cover nearly every common speech synthesis use case. Understanding these features before you start building helps you architect your audio layer correctly from day one.
Speech Control Methods
The plugin provides methods to speak a given text string, stop ongoing speech, pause speech mid-utterance, and resume from the paused position. Pause and resume are particularly useful for long-form content like article narration or book reading, where users may need to step away and return.
Language and Voice Selection
You can query the plugin for all available languages on the device, check whether a specific language is supported, and set the active language before speaking. The plugin also exposes available voices, letting you pick a specific voice by name or identifier. This is essential for apps that serve multiple locales or need a consistent voice across sessions.
Speech Parameter Control
Three parameters shape the listening experience. Speech rate controls how fast words are spoken. Pitch adjusts the tone, making the voice higher or lower. Volume sets the output loudness. All three accept numeric values within defined ranges, and you can adjust them dynamically between utterances.
Platform Coverage
The plugin supports Android, iOS, Web, macOS, and Windows. This broad coverage means you can write your TTS logic once in Dart and deploy it across mobile, desktop, and web targets without platform-specific branching in most cases.
File Synthesis
On supported platforms, the plugin can synthesize speech directly to an audio file instead of playing it through the speaker in real time. This is valuable for pre-generating audio content, caching frequently used phrases, or building offline-first experiences.
Setting Up Flutter Text to Speech
Setting up the flutter_tts plugin involves three phases: adding the dependency to your project, configuring platform-specific requirements, and handling common pitfalls that catch developers off guard.
Adding the Dependency
You add the flutter_tts package to your project by including it in your pubspec file dependencies. After saving the file, run the package get command to download and link the plugin. Once installed, import the package into any Dart file where you need speech synthesis. The plugin exposes a single class that you instantiate and reuse throughout your app's lifecycle.
Android Configuration
Android requires a minimum SDK version that the plugin specifies in its documentation. You need to ensure your app's Android Gradle configuration meets or exceeds that minimum. Additionally, Android 11 and above require a specific manifest entry that declares intent queries for the text-to-speech engine. Without this entry, the plugin may fail to find any available TTS engine on the device. You also need to verify that your Kotlin version is compatible with the plugin version, as version mismatches between the plugin's native Kotlin code and your project's Kotlin Gradle plugin are a frequent source of build failures.
iOS Configuration
iOS requires a minimum deployment target that the plugin enforces. You set this in your Xcode project or through your Flutter iOS configuration. Unlike Android, iOS does not require special permission entries for speech synthesis because AVSpeechSynthesizer does not need microphone or speech recognition permissions. However, if your app also uses speech recognition alongside TTS, you must add the speech recognition permission string to your Info.plist.
Web Configuration
Web support relies on the browser's built-in SpeechSynthesis API. No additional setup is needed beyond including the package. However, browser support varies. Chrome and Edge provide full support, while some mobile browsers have limited or inconsistent TTS behavior. Always test on your target browsers.
Common Pitfalls
The most frequent setup errors are: forgetting the Android manifest queries tag for TTS intent, using an incompatible Kotlin version, setting an Android minSdk below the plugin requirement, and failing to initialize the plugin instance before calling speak. Each of these produces a specific error that the troubleshooting section below addresses.
Platform-Specific Considerations
While the flutter_tts plugin abstracts most platform differences, several behaviors vary across Android, iOS, and Web in ways that affect your app design.
On iOS, pause and resume are fully supported through the native AVSpeechSynthesizer API. On Android, pause and resume support depends on the TTS engine installed on the device. Some Android engines support pausing, while others do not, which means you should design your UI to degrade gracefully if pause is unavailable.
On Web, the plugin provides progress updates that indicate how far through the utterance the browser has spoken. This is useful for building word-by-word highlighting in reading apps. Android and iOS do not provide equivalent progress callbacks at the same granularity.
Volume control on iOS is tied to the system media volume, while Android allows independent volume control through the plugin. On Web, volume control depends on the browser implementation and may not always produce audible differences across the full range.
File synthesis is available on Android and iOS but is not supported on Web. If your app needs offline audio generation, you should detect the platform at runtime and fall back to real-time speech on Web.
Managing Voices and Languages
Every device has a different set of installed TTS voices and supported languages. Your app should never assume a specific voice or language is available. Instead, query the device at runtime and adapt.
The plugin exposes methods to retrieve all available voices, each with a name and a locale identifier. You can also check whether a specific language code is available before attempting to speak in that language. The recommended approach is to call these methods shortly after initializing the plugin, store the results, and present voice options to the user if your app allows voice selection.
When a user selects a language your app supports but the device does not have installed, you should handle this gracefully. On Android, you can direct the user to install the missing language data from the device's TTS settings. On iOS, additional voices can be downloaded from Settings under Accessibility and Spoken Content. Your app should detect the missing language, inform the user, and provide a fallback to the default voice.
A practical tip: cache the list of available voices after the first query rather than calling the method repeatedly. Voice lists rarely change during a session, and repeated queries add unnecessary overhead.
Controlling Speech Parameters
Speech rate, pitch, and volume are the three levers you have for shaping how synthesized speech sounds. Each serves a distinct purpose depending on your app's context.
Speech Rate
Speech rate controls how quickly the TTS engine reads text. A slower rate benefits accessibility scenarios where users need time to process each word. A faster rate suits power users who want information quickly, such as in a voice assistant reading notifications. The plugin accepts a numeric value, and the default represents normal speed. Setting the rate too high can make speech unintelligible, especially on lower-quality TTS engines.
Pitch
Pitch adjusts the fundamental frequency of the synthesized voice. Higher pitch values produce a more animated, energetic tone, while lower values sound calmer and more authoritative. Pitch is particularly useful in gaming and storytelling apps where different characters need distinct vocal characteristics. However, extreme pitch values can produce robotic or distorted output, so test across the devices you target.
Volume
Volume controls the output loudness of the synthesized speech. This is independent of the device's system volume on Android but tied to media volume on iOS. For apps that mix TTS with other audio, such as background music or notification sounds, volume control lets you duck the TTS audio to an appropriate level without affecting other audio streams.
The best practice is to expose these parameters as user-adjustable settings in your app's preferences, with sensible defaults that work for the majority of users.
Handling Playback Lifecycle
Managing the speech playback lifecycle correctly is what separates a polished app from a frustrating one. Users expect speech to start when they tap a button, stop when they navigate away, and not conflict with other audio on their device.
Start and Stop
Always ensure that any ongoing speech is stopped before starting a new utterance. If you call speak while a previous utterance is still playing, behavior varies by platform. On iOS, the new utterance queues behind the current one. On Android, the behavior depends on the engine. To avoid confusion, call stop before speak unless you intentionally want queueing behavior.
Pause and Resume
Use pause when users need to temporarily halt speech, such as when a phone call comes in or the user switches to another app. Resume picks up from the paused position. Because pause support is inconsistent on Android, implement a fallback that stops and re-starts from the beginning if pause is not available.
Handling Interruptions
When a phone call or system notification interrupts your app, the native TTS engine may stop speech automatically. You should listen for completion or error callbacks and update your UI state accordingly. On iOS, you can observe audio session interruption notifications to pause speech proactively. On Android, the TTS engine typically handles interruptions internally, but you should still reset your UI state when the interruption ends.
App Backgrounding
When your app goes to the background, ongoing speech may continue or stop depending on the platform and engine. iOS generally stops speech when the app is backgrounded unless you configure an audio session for background audio. Android behavior varies by engine. If background speech is important for your use case, test thoroughly and configure the appropriate background modes.
Architecture Overview
Understanding how the layers interact helps you debug issues and reason about performance. Your Flutter app creates an instance of the TTS plugin class in Dart. When you call the speak method, the plugin sends the text and parameters through a platform channel to the native layer. The native layer invokes the platform's TTS engine, which synthesizes audio and routes it to the device's audio output. Callbacks flow back through the same channel to your Dart code.
This architecture means that the quality and availability of voices depends entirely on what the device provides. The plugin is a thin transport layer. If a user's device has no TTS engine installed, no amount of plugin configuration will produce speech. Your app should always verify engine availability before attempting to speak.
Performance and Quality Tips
Speech synthesis performance matters most in real-time apps where users expect immediate audio feedback. Here are practical strategies to keep latency low and quality high.
Pre-load the TTS engine by initializing the plugin instance early in your app's lifecycle, ideally during splash screen or app initialization. This avoids the first-speak delay that occurs when the engine loads lazily. For frequently used phrases, consider using file synthesis to pre-generate audio files and play them back through a standard audio player, which eliminates synthesis latency entirely.
Cache the list of available voices and languages after the first query. Repeated queries add overhead and provide no benefit since the list rarely changes during a session. When switching languages, set the language before calling speak to avoid the engine reconfiguring mid-utterance.
Test on low-end Android devices with budget TTS engines. High-end devices with Google's neural TTS voices produce excellent quality, but many users have older devices with basic engine voices. If your app targets emerging markets, assume the lowest common denominator and design your UX to work with basic TTS quality.
For apps that need premium voice quality or cloud-based TTS with providers like ElevenLabs or Google Cloud TTS, consider integrating a server-side synthesis pipeline. VideoSDK's AI voice agents provide a production-ready architecture for connecting cloud TTS providers to real-time communication sessions, which can complement your Flutter app's on-device TTS for scenarios requiring studio-quality voices.
Common Issues and Debugging Strategies
Even with correct setup, developers encounter recurring issues with Flutter text to speech. Here are the most frequent problems and how to diagnose them.
Engine Not Found on Android
If the plugin reports no available TTS engine, the device may not have one installed or the manifest queries tag is missing. Verify that your Android manifest includes the intent query for the TTS service. On the device, check Settings under Languages and Input to confirm a TTS engine is installed and enabled.
Language Not Installed
When you set a language and the engine does not support it, speech may fall back to the default language or produce no output. Always check language availability before setting it. On Android, direct users to download language data from the TTS settings. On iOS, additional voices can be installed from the Accessibility settings.
Permission Errors
TTS itself does not require special permissions on most platforms. However, if your app combines TTS with speech recognition or microphone access, missing permissions for those features can cause cascading failures. Verify that all required permission strings are present in your Info.plist and Android manifest.
Speech Not Playing on Web
On Web, the SpeechSynthesis API requires a user gesture before it can start speaking in some browsers. If speech does not start automatically, ensure the first speak call is triggered by a user interaction such as a button tap.
Plugin Initialization Errors
If you see errors about the plugin not being initialized, ensure you have created an instance of the TTS class before calling any methods. The plugin is not a static utility. Each instance manages its own state and callbacks.
Choosing the Right TTS Package for Your Project
While flutter_tts is the most popular Flutter text to speech package, several alternatives exist. Choosing the right one depends on your platform targets and feature needs.
The flutter_tts package offers the broadest platform support (Android, iOS, Web, macOS, Windows) and the most comprehensive feature set, including file synthesis, voice selection, and parameter control. It is the default choice for most projects.
The texttospeech package is a lighter alternative with a simpler API but fewer features. It may suit projects that only need basic speech on mobile platforms and do not require Web or desktop support.
The texttospeechplus package is a community fork that adds additional features on top of the original. It can be useful if you need specific capabilities that fluttertts does not offer, but you should verify maintenance activity and compatibility before committing.
The gallitextto_speech package focuses on Nepali language support and is niche. It is only relevant if your app specifically targets Nepali speech synthesis.
For most developers, fluttertts remains the strongest choice due to its platform coverage, active maintenance, and feature completeness. Evaluate alternatives only when you have a specific requirement that fluttertts cannot meet.
Real-World Use Cases
Flutter text to speech powers several categories of apps that developers build today.
Accessibility Narration
Apps for visually impaired users rely on TTS to read screen content aloud. A news app might speak article text when the user taps a headline. A navigation app might speak turn-by-turn directions. In both cases, clear speech at a controllable rate is essential. The flutter_tts plugin's language and voice selection features let these apps serve users across multiple locales.
In-App Voice Assistants
Productivity and smart home apps use TTS to provide spoken responses to user queries. The app processes a command, generates a text response, and speaks it back. Combined with speech-to-text for input, this creates a hands-free interaction loop. For apps that need real-time voice conversations with AI, VideoSDK's AI voice agent architecture provides a production-ready pipeline connecting STT, LLM, and TTS providers.
Language-Learning Apps
Education apps use TTS to demonstrate pronunciation of words and phrases in the target language. The ability to set specific languages and adjust speech rate makes the plugin ideal for this use case. Learners can slow down speech to hear each syllable clearly, then speed it up as they progress. File synthesis can pre-generate pronunciation audio for offline lessons.
Definitions Glossary
TTS (Text to Speech): The process of converting written text into spoken audio using a speech synthesis engine. In Flutter, this is accessed through plugins like flutter_tts that bridge to native platform engines.
flutter_tts Plugin: The most widely used Flutter package for text to speech, providing a Dart API that wraps Android TextToSpeech, iOS AVSpeechSynthesizer, and the Web SpeechSynthesis API.
Platform Channel: Flutter's mechanism for communicating between Dart code and native platform code. The flutter_tts plugin uses platform channels to send text and parameters to native TTS engines.
Speech Rate: A numeric parameter controlling how fast the TTS engine reads text. Higher values produce faster speech, useful for power users, while lower values aid accessibility.
Voice Selection: The ability to choose a specific synthesized voice from the available voices on a device. Voices are identified by name and locale, and availability varies by platform and installed engine data.
File Synthesis: A feature that generates speech audio as a file on disk rather than playing it through the speaker in real time. Useful for caching, offline playback, and pre-generating content.
Key Takeaways
- The flutter_tts plugin is the standard solution for Flutter text to speech, supporting Android, iOS, Web, macOS, and Windows through a single Dart API.
- Always query available voices and languages at runtime rather than assuming specific ones are present on every device.
- Platform differences matter: pause and resume work reliably on iOS but inconsistently on Android, and Web requires user gestures to start speech in some browsers.
- Pre-loading the TTS engine and caching voice lists are the two most effective performance optimizations for production apps.
- For apps needing premium cloud-based voices, consider pairing on-device TTS with a server-side synthesis pipeline using VideoSDK's AI voice agent architecture.
Conclusion
Adding speech synthesis to your Flutter app transforms the user experience for accessibility, education, and voice assistant use cases. The fluttertts plugin gives you a single Dart API that works across five platforms, with control over voices, languages, rate, pitch, and volume. The setup is straightforward once you handle the platform-specific manifest and configuration requirements, and the debugging strategies above cover the issues most developers encounter. For apps that combine TTS with real-time video or audio communication, VideoSDK's Flutter SDK provides the calling infrastructure that pairs naturally with speech features. What are you building with Flutter text to speech? Drop a comment below, and check out the [fluttertts package on pub.dev](https://pub.dev/packages/flutter_tts) to get started. You can also join the VideoSDK Discord community to discuss real-time communication integrations with fellow developers.
Free $20 Balance for AI Voice Agents & Video Calls
FAQ
