← Computing Methods

Text-to-Speech Narration

The spoken track behind every animation, generated fully automatically.

Three Stages

All three animation types (City, Station and Decay; see Physics Animation (Blender)) include a spoken narration track describing the earthquake and its local impact. The narration is generated fully automatically in three stages:

  1. Place names from the gazetteer. The affected city's county and state come from the GeoNames data behind our 2,002 named places, so the narration names them correctly ("Bigfork is located in Flathead County in Montana"). Earlier versions asked a language model for these facts and for a notable characteristic of each city; that step was dropped in October 2026 because the model sometimes invented facts.
  2. Speech-friendly text normalization. Before synthesis, the narration text is rewritten so every element is spelled out the way a person would say it: clock times become words ("two thirty four in the morning"), dates use ordinal words ("December twentieth twenty twenty two"), magnitudes are expanded ("magnitude six point four"), and seismic station codes are spelled letter by letter (e.g. "BK.BARR.HNE" → "B K, B A R R, H N E"). A city/state comma is dropped and a bare state abbreviation expanded ("Redwood Valley, CA" → "Redwood Valley California") because the TTS engine reads a comma as a pause and produces an audible stumble mid-placename. This keeps the spoken track clear and unambiguous.
  3. Text-to-speech synthesis. The normalized narration is converted to an MP3 audio file using Kokoro, an open-weight 82-million-parameter neural text-to-speech model running locally via the ONNX runtime: no third-party API calls are made and no audio leaves the server. The synthesised audio is muxed into the MP4 with ffmpeg so it plays from the start of the animation; if the narration outlasts a short animation, the final video frame is held until the narration completes.

Voice Options

Kokoro ships a set of built-in voices; Intensity Lab currently narrates with a single configured voice (am_michael by default), set via an environment variable rather than hardcoded, so a different voice can be selected without a code change if narration quality needs adjusting for a particular audience.

Kokoro replaced an earlier Coqui TTS engine, which required an older Python version; Kokoro has no such dependency and its model files are baked directly into the deployed image, so narration never depends on a runtime model download.