[Name]: text
[Name]: text
Your key is stored only in this browser and sent directly to Google's API. Get one free at Google AI Studio.
Two minutes from text to lifelike speech.
Add your key. Open settings Settings and paste a free Gemini API key. Voice previews work even without one.
Write your text. Type, paste, or drop a .txt / .md file straight onto the editor.
Press Play. Speech is generated, then plays. Press again to pause - or to cancel while it's still generating.
Shortcut: Ctrl + Enter generates from anywhere.
Load a .txt, .md, .json, or .csv document. Audio files are handed to the Speech to Text page automatically.
Voice-type straight into the editor - your words appear at the cursor, ready to generate.
Its own sidebar page: record or upload audio, review the transcript, get an AI overview, then send it to the editor, copy, or download as a .txt. Optional speaker labels (ready to re-voice) and [MM:SS] timestamps.
One click to polish wording, summarize for audio, or auto-add expressive cues.
Write [Name]: line and every character gets their own voice. Two speakers become one seamless conversation.
Drop pauses, sighs, whispers, and more exactly where you want them.
Every voice has an editable personality prompt on the Voices page. Flip the toggle on a voice card, describe the delivery you want - "an old storyteller by a campfire" - and that voice performs it every time. It saves as you type, Reset to default brings back the original, and the player shows · Acting while it's on. Off by default.
Previews on the Voices page play instantly from bundled official samples - free, offline, no key needed.
Put these tags right in your text - the voice performs them naturally:
[short pause]
[medium pause]
[long pause]
[sigh]
[laughing]
[uhm]
[whispering]
[shouting]
[sarcasm]
Need an exact silence? [pause: 1.5s] inserts one to the millisecond.
And you can write in 25+ languages - Tamil, Hindi, Japanese, and more - the voice adapts automatically.
lock Everything stays in this browser and goes only to Google's API. No account. No server.
Last updated: 31 July 2026
By using VoxGemini, you agree to these terms:
VoxGemini runs entirely in your web browser. It has no servers of its own and never stores your documents, generated audio, or API keys anywhere outside your browser.
VoxGemini is a direct interface to Google's APIs (Cloud Text-to-Speech and the Gemini API), using your own API key. Your use of this application must comply with Google's API Terms of Service, and any usage quotas or charges on your key are between you and Google.
The voice preview clips bundled with the app are official sample recordings from Google's Gemini-TTS documentation and remain Google's content. They are included for preview purposes only.
VoxGemini is provided "as is" without warranty of any kind. We are not responsible for API quotas, costs, service availability, or data handling by Google or other third parties.
Last updated: 31 July 2026
Here is exactly where your data goes - and where it doesn't:
Your documents, generated audio clips, presets, acting directions, and API key are saved only in this browser (localStorage and IndexedDB). Nothing is ever uploaded to a VoxGemini server - there isn't one.
When you generate speech or use AI Assist, your API key and text are sent directly from your browser to Google's APIs: the Cloud Text-to-Speech endpoint (texttospeech.googleapis.com) and the Gemini API (generativelanguage.googleapis.com). Google's handling of that data is governed by Google's own privacy policy and API terms.
The page loads styling and fonts from CDNs (Tailwind, Google Fonts) and voice avatar photos from Unsplash. Like any web request, those services receive your IP address and browser details when the assets load. None of your documents, audio, or keys are ever sent to them.
VoxGemini uses no tracking cookies, no analytics, and no ads. Removing your data is always in your hands: Forget key in Settings, Clear all in History, or simply clear this site's browser storage.