Type a script. Pick a voice. Export a TikTok-ready vertical video with burned-in word-by-word captions that hit the beat. All on-device. No account. No watermark.
American and British English. Curated, character-driven, ready to ship.
On desktop, Whisper-base transcribes the generated audio back to find exact word timings, so captions highlight in sync with the voice. Phones skip that 400 MB download and use faster phoneme-based timing, which is close but not word-exact.
Six preset looks covering common creator aesthetics: bold pop, minimal underline, kinetic one-word, and more. Burned into the vertical export.
Step 1. Type your script. Paste the copy for your TikTok, Reel, or Short. Insert pause markers like [pause:500] anywhere you want a breath.
Step 2. Pick a voice. Audition any of the 28 curated Kokoro-82M voices. Each has its own personality: narrator, friend, reporter, warm, clinical.
Step 3. Generate and align. Kokoro synthesises the speech locally. On desktop, Whisper-base then listens back to the result and extracts word-level timing, so captions land precisely on each word. On phones, timing is derived from phonemes instead, which keeps the download small at some cost in precision.
Step 4. Style and export. Choose one of six caption styles. Export as a 9:16 vertical MP4 with captions burned in. Upload straight to TikTok, Reels, Shorts.
Everything happens in your browser. The models download once and are cached forever, after which every future video is local and private. Speech alone is ~300 MB. Adding the desktop Whisper aligner takes it to ~700 MB.
CapCut and Canva have auto-captions, but their TTS locks behind subscriptions and every export routes through their servers. Speechify and ElevenLabs do premium voices but don't ship caption alignment or vertical video export. Free alternatives usually watermark the output or cap daily generations.
PixVoice sits where those lanes don't overlap: neural-quality TTS (Kokoro-82M), burned-in karaoke captions (Whisper word alignment), vertical export, and zero gatekeeping. The trade-off is a ~700 MB first-run model download. After that, every export is free and local forever.
Yes, fully free, forever. No account, no signup, no credit limits, no watermark. The models download once into your browser cache and run locally after that.
No. Kokoro-82M (TTS) and Whisper-base (caption alignment) run entirely in your browser via WebAssembly. There's no upload endpoint on our side, so technically we couldn't see your script even if we wanted to.
Vertical 9:16, ready for TikTok, Instagram Reels, and YouTube Shorts. Chrome and Edge export MP4. Firefox does not support MP4 recording, so it exports WebM instead, which some platforms want converted before upload. Six caption styling presets cover the most common creator looks.
No, and we do not use it. PixVoice runs on WebAssembly only. The WebGPU path is deliberately disabled because it caused device-loss crashes in production. WebAssembly is the slower backend, but it is the reliable one, and it runs everywhere Chrome, Edge, or Firefox runs.
Partly, and we would rather say so up front. Voice generation works on phones. Word-level Whisper alignment does not: we skip that 400 MB model on mobile because it exhausts memory on most phones, so captions fall back to phoneme-based timing. That is close, but not word-exact. For the tightest karaoke sync, use a desktop or laptop.
28 curated Kokoro voices across American and British English. Each tagged and character-driven: narrator, warm, reporter, clinical, and more.