Audio & Accessibility
Voice Studio
Turn text or PDFs into listenable narration with local extraction, English OCR, neural speech, and audio export.
What it includes
Text-to-speech built for long documents
- Text, TXT, and PDF input with selectable page ranges
- Local PDF text extraction and English OCR fallback for scanned pages
- Kokoro neural speech plus available browser voices
- Sequential long-form narration with retry and cancellation
- Locally encoded WAV and MP3 downloads
Useful for: PDF listening · Audiobook drafts · Proof-listening · Accessibility support.
The neural model downloads on first use. Available browser voices, maximum practical document size, and export performance vary by browser and device.
Workflow
How to use Voice Studio
- Open Voice Studio and enter text, import a TXT file, or choose a PDF.
- For a PDF, select the pages and extract the text; scanned English pages can use local OCR.
- Choose Kokoro or an available browser voice and adjust the speech controls.
- Generate the narration, review playback, and retry any incomplete section.
- Download the completed narration as WAV or MP3.
Privacy
Your document and generated audio stay out of Analytics.
PDF extraction, English OCR, and Kokoro speech generation run in your browser. Google Analytics may measure broad page usage, but it does not receive your documents, entered text, rendered PDF pages, or generated audio. Kokoro downloads its model from the upstream model repository on first use; browser speech may use services supplied by your browser or operating system.
Read the Tools privacy details