← Writing

Technical report · 18 February 2026 · 14 min

Training a voice for $14, then fitting it inside the browser

The training bill was the obvious constraint. Deployment turned out to be the more interesting one.

A cheap training run does not create a cheap product if every generated sentence still requires a remote GPU. The goal of this project was an end-to-end path from audio and transcripts to a voice that could run in a browser, on a phone, on ARM hardware, or on a desktop without an inference API.

The pipeline combines VITS through Piper’s training wrapper with Sherpa-ONNX for portable inference. The reported experiment used LJSpeech, two A100 GPUs on a GCP spot instance, an effective batch size of 128, and 10,000 training steps. Training took roughly three hours and cost about $14 at the spot price available during the run.

The dataset and preprocessing boundary

LJSpeech contains 13,100 clips—approximately 24 hours—from one English speaker reading public-domain books. Audio is 22,050 Hz, 16-bit. It is not a diverse production dataset; it is useful here because its format and failure modes are familiar enough that pipeline problems are easier to distinguish from dataset problems.

Preprocessing converts text into IPA phonemes with eSpeak-NG, computes mel spectrograms, and caches the results in a training directory. This takes roughly ten minutes in the reported setup. Keeping preprocessing separate means a preempted training machine does not need to repeat deterministic CPU work.

What is actually being trained

VITS combines a conditional variational autoencoder, a normalizing flow, and adversarial training. The generator maps phoneme-conditioned latent representations to a raw waveform; a discriminator learns to distinguish generated audio from real samples. The adversarial component is why the discriminator loss oscillates rather than decreasing smoothly.

The detailed CSV logging covers only the first 1,223 steps, where generator loss fell from 59.98 to the low-to-mid 40s and discriminator loss oscillated between roughly 1.4 and 2.5. The full training log reports a differently scaled final generator loss of about 4.32 at step 10,000. Those two logging scales should not be spliced into one curve, so the explorer below stops exactly where the dense source data stops.

What the recorded losses actually show

Dense logging from steps 49–1,223.Plotted directly from the repository’s generator, discriminator, and validation CSV files.
generatordiscriminatorvalidation
step 49generator 59.98discriminator 2.28validation 51.83

Exporting was not a file-format conversion

piper_train.export_onnx converts the checkpoint into an approximately 61 MB FP32 ONNX model, but two compatibility problems sit between “export succeeded” and usable inference.

First, PyTorch 2.6 and newer default torch.load to weights_only=True. Piper checkpoints contain pathlib.PosixPath objects, so loading fails unless that type is explicitly allow-listed with torch.serialization.add_safe_globals.

Second, Piper’s exported file does not automatically include every metadata field Sherpa-ONNX needs. The pipeline patches sample_rate, add_blank, n_speakers, voice, comment, and language into the ONNX metadata. The most important value is voice = en-us: Sherpa passes it to eSpeak-NG. When it is absent or invalid, native builds throw a useful initialization error while the WASM build may fail silently.

Quantization changes the product equation

Dynamic INT8 quantization reduced the model from roughly 61 MB to about 20 MB. An FP16 build around 30 MB provides another option. The decoder proved tolerant of the small weight perturbations introduced by INT8 in this experiment, but perceptual quality needs to be evaluated for every new voice rather than assumed.

The browser bundles the model with a Sherpa-ONNX WebAssembly runtime. It downloads once, can be cached, and then generates speech without an API key or network round trip. This removes GPU scheduling, inference-server scaling, request metering, and audio-stream transport from normal use.

The report’s measured ten-word benchmark was 307 ms with single-threaded ONNX Runtime (RTF 0.095), 201 ms with four-thread Sherpa-ONNX (RTF 0.074), and 1,360 ms in Chromium/WASM (RTF 0.472). For two, ten, and thirty words, Sherpa’s CPU RTF remained 0.069, 0.074, and 0.077. These are measurements on an Intel Xeon at 2.20 GHz and a Cloudflare Pages browser build, not universal device guarantees.

The browser build still has costs

The prebuilt English runtime’s .data payload is approximately 93 MB before aggressive stripping, considerably larger than the model alone because it includes eSpeak-NG resources. The runtime does not stream VITS output: it generates the entire utterance before playback. Splitting long input reduces waiting time but creates prosody discontinuities at chunk boundaries.

Model swapping is also expensive. The current Emscripten build preloads the ONNX file into the WASM data bundle, so changing the voice requires rebuilding the bundle rather than replacing one URL at runtime. That is acceptable for a demonstration and awkward for a product with many voices.

Dependency failures took longer than training

Piper pins older PyTorch expectations, while the GCP image used PyTorch 2.7.1 with CUDA 12.6/12.8-era libraries. PyTorch Lightning 1.7.7 also disagrees with newer scheduler validation. CUDA libraries existed in multiple paths without the correct LD_LIBRARY_PATH. Python 3.14 broke Piper’s Cython extensions, so the reproducible environment settled on Python 3.10.

This is the part hidden by a one-line “trained for $14” summary. The compute invoice was small; aligning Piper, Lightning, modern PyTorch, CUDA, ONNX metadata, Sherpa, eSpeak-NG, Emscripten, and browser isolation headers was the actual engineering work.

Spot pricing changes reliability

The instance survived the complete run, but spot capacity can be preempted. With an approximately 18-minute checkpoint interval, interruption would lose up to one interval of progress. More frequent checkpoints reduce lost work while increasing storage traffic and synchronization overhead.

What I would change next

I would train longer, checkpoint more deliberately, decouple model files from the WASM bundle, strip unused eSpeak-NG languages, add utterance chunking with overlap-aware prosody handling, and benchmark real-time factor and memory across representative phones, laptops, and ARM boards. The interesting result is not that browser TTS is “free”; it is that deployment costs move from a permanent server dependency into a measurable first-load and device-performance budget.

Read the pipeline ↗Read the full training report (PDF) ↗