Notepad, except the text can speak in any voice you have cloned.
- Clone a voice — record 20–30 seconds with your microphone, or pick an audio file.
- Type what you want it to say.
- Press Speak (or Ctrl+Enter). The speech plays as soon as it is ready; Save Audio keeps a copy.
- Everything runs locally. No accounts, no telemetry, no cloud. The network is only used to download the models you approve, and offline mode blocks even that.
Voices are like fonts: pick one from the voice menu, type, and speak. Everything else — choosing the clean part of a recording, writing down its words (local Whisper), splitting long text, loading the engine, stitching the sentences together — happens underneath.
- Engines: Qwen3-TTS 1.7B Base (default, voice cloning from a short sample) and optional Chatterbox-Turbo.
- Transcription: faster-whisper, CPU by default.
- Shell: Tauri 2 (Rust) + React/TypeScript; Python workers for audio and inference; FFmpeg for media; SQLite for metadata.
Status: v0.2.0 standalone Linux desktop app. See docs/TEST_REPORT.md for what has been verified on Linux with an NVIDIA GPU.
Aimed at Pop!_OS / Ubuntu 24.04-class desktops with an NVIDIA GPU.
| Component | Notes |
|---|---|
| NVIDIA driver (CUDA 12.8 capable, driver ≥ 570) | Use the driver already on the system. No CUDA toolkit install is required. |
ffmpeg, libportaudio2 |
sudo apt install ffmpeg libportaudio2 |
| Python 3.12 for the engine runtime | Created by first-run setup or scripts/bootstrap.sh with uv (no sudo; several GB for the CUDA stack) |
| Model weights | Offered in-app the first time they are needed, after size, source and license are shown (Qwen ≈ 4.5 GB, whisper ≈ 0.5 GB, Chatterbox ≈ 3 GB optional) |
| Building from source | Rust, Node 22, and libwebkit2gtk-4.1-dev build-essential curl wget file libxdo-dev libssl-dev libayatana-appindicator3-dev librsvg2-dev patchelf |
Build a package with scripts/build.sh (or scripts/build.sh --deb-only), then:
sudo dpkg -i src-tauri/target/release/bundle/deb/*.debThat installs one launcher (com.shadowfetch.voicestudio.desktop). The Python engine runtime is not inside the package; the app creates it under the data directory on first run.
scripts/doctor.sh # what is installed / missing, GPU + CUDA check
scripts/bootstrap.sh # isolated Python environments (asks before optional Chatterbox)
scripts/install-linux.sh # build the .deb, install it, and isolate the runtimescripts/install-linux.sh replaces the installed desktop app and writes a real venv under
$XDG_DATA_HOME/com.shadowfetch.voicestudio/runtime (default ~/.local/share/…/runtime) — not a symlink back to the checkout.
Flags: --skip-build to reuse an existing .deb; --with-chatterbox to include the optional engine env.
scripts/dev.sh # desktop app in development mode
scripts/test.sh # unit tests (add --real for GPU/model integration tests)
scripts/build.sh --deb-only # supported Linux package
scripts/build.sh # also tries an AppImage (best-effort)The .deb is the supported Linux install. AppImage output is optional: linuxdeploy can fail (FUSE or packaging), even when APPIMAGE_EXTRACT_AND_RUN=1 is set.
The app opens on Speak. The first time you clone a voice or press Speak, Voice Studio offers the local model it needs (name, download size, source, license) and downloads it only when you click Download. After that it works offline.
- Speak — voice menu, a large text editor (saved automatically, restored after a restart), the Speak button, the result with Save Audio, and a short Recent list.
- Voices — your cloned voices: play the sample, use, rename, add a recording, edit the sample, delete.
- Settings (gear) — auto-play, Save Audio format, offline mode, microphone. Advanced holds engine selection, the engine's own settings, seed, pauses, pronunciation, export details, model management, storage, the system check and the multi-take project editor.
Keyboard: Ctrl+Enter speaks, Esc stops.
Paths follow XDG. If the variables are unset, the defaults below apply.
| Default | |
|---|---|
| Projects, recordings, voices, masters, SQLite, engine runtime | $XDG_DATA_HOME/com.shadowfetch.voicestudio/ → ~/.local/share/com.shadowfetch.voicestudio/ |
| Settings | $XDG_CONFIG_HOME/com.shadowfetch.voicestudio/settings.json → ~/.config/com.shadowfetch.voicestudio/settings.json |
| Regenerable caches (peaks, engine prompts) | $XDG_CACHE_HOME/com.shadowfetch.voicestudio/ → ~/.cache/com.shadowfetch.voicestudio/ |
| Model weights (Hugging Face cache layout) | $XDG_DATA_HOME/com.shadowfetch.voicestudio/models/hf/ |
Local storage is not encrypted. Treat the data directory like any other private files.
- docs/USER_GUIDE.md — Speak, Clone Voice, Voices, Settings
- docs/ARCHITECTURE.md — process layout and data model
- docs/PROTOCOL.md — JSON-lines worker protocol
- docs/MODEL_LICENSES.md — engine and model licenses
- docs/FINETUNING.md — optional Qwen dataset export (no in-app training)
- docs/THIRD_PARTY_NOTICES.md
- docs/TEST_REPORT.md
- docs/screenshots/ — UI layout (browser preview mock)
License: MIT.