Guide

Articles read aloud

ssg can make an article listenable in two ways. They work alone or together.

Browser button (listen:) Build-time MP3 (tts:)
Where the voice comes from The visitor's device (Web Speech API) A TTS API, called while the site builds
Cost Nothing The API's price, once per changed article
Same voice for everyone No: whatever the device has Yes
Works offline, in a podcast app No Yes, and there is a podcast feed
Needs JavaScript Yes No: a plain <audio> element
Leaves the device Nothing The article text goes to the API at build time

Both appear as the same small Listen button with a speaker icon, next to the reading time in the bundled themes. On a page with an MP3 it plays the file; on a page without one it reads the text with the browser's voice. That includes a page whose MP3 could not be made because the API was down, so a failed build-time call still leaves the reader a way to listen.

The browser button

1listen:
2  enabled: true
3  label: "Listen"        # button text
4  voice: ""              # preferred browser voice, by name: "Natural", "Google UK English Female"
5  player: compact        # compact (default): the button plays the MP3; full: the browser's <audio> bar
6  sections: [posts]      # posts (default), pages, or both
7  auto: true             # place it where the theme did not; default true

The button reads the article with speechSynthesis in the page's language (<html lang>), one paragraph at a time. Chrome stops a single long utterance after about 15 seconds, so one utterance per paragraph is what lets a long article finish. Pressing it again stops. It reports its state with aria-pressed, and it stays hidden in browsers without the API and when JavaScript is off, so nobody sees a button that does nothing.

The voice is chosen, not left to the browser's default, which is usually the oldest one installed. Among the voices for the page's language the button prefers one whose name contains listen.voice, then neural voices (Edge "Natural"/"Online", Chrome "Google", Apple "Premium"/"Enhanced"/"Siri"), then an exact language-and-region match.

Voice quality still depends on the device. Recent macOS, iOS, Android and Windows ship good neural voices for major languages; a minimal Linux desktop may have only a robotic one, or none.

The MP3

 1tts:
 2  enabled: true
 3  provider: generic                       # generic | openai | elevenlabs | google
 4  api_url: http://localhost:8080/v1/speech
 5  api_key: $TTS_API_KEY                   # from the environment; never a literal
 6  voice: en_US-lessac-medium
 7  lang: ""                                # default: each page's language
 8  speed: 1.0
 9  jingle_url: https://example.com/intro.mp3   # optional, played before every article
10  sections: [posts]
11  dir: audio                              # → /audio/<page>.mp3
12  timeout: 60s                            # per request
13  retries: 2                              # extra attempts on 429, 5xx, network errors
14  breaker: 3                              # failed articles in a row → stop calling
15  on_failure: stale                       # stale | skip | fail
16  feed: true                              # podcast feed
17  feed_path: podcast.xml
18  feed_title: "Example, read aloud"
19  feed_author: "Example"
20  feed_image: /images/podcast.png
21  feed_limit: 50

What is read: the title, then the article text with code blocks removed and markup stripped. A listener does not need func main() { spelled out.

Cache

Each MP3 is cached in .ssg-cache/tts/ under a key made from everything that changes the sound: provider, endpoint, model, voice, speed, instructions, language, the jingle and the text. A rebuild calls the API only for articles whose words changed. ssg cache clean --namespace=tts forgets them all. Keep .ssg-cache/ between CI runs (see DEPLOYMENT), or every build pays for every article again.

The jingle is downloaded once per URL and cached next to the audio. It is joined in front of each article's MP3, after the ID3 tags between the two are removed. Both should use the same encoding, for example 44.1 kHz stereo MP3.

When the API does not answer

A TTS API is a network dependency, and a build should not wait on it once it is plainly down:

  1. Timeout and retries. Each request has timeout. A network error, a timeout, 429 or a 5xx is retried retries times, waiting 1 s, 2 s, 4 s… or what Retry-After asks for, capped at 30 s. A 4xx other than 429 (bad key, unknown voice) is not retried, because asking again changes nothing.

  2. Breaker. After breaker articles fail in a row, the API is not called again for the rest of the build. One line says so. Without it, a down API costs timeout × retries × articles.

  3. on_failure decides what the page gets:

    • stale (default): the last MP3 that page ever had, from .ssg-cache/tts/last/, even if the text has changed since. If there is none, the page has no audio.
    • skip: no audio for that page.
    • fail: the build stops.

    strict: true makes stale and skip behave as fail.

A page without audio still gets the browser button when listen is on.

Providers

generic is ssg's own contract: POST api_url with JSON {"text", "voice", "lang", "speed", "format": "mp3"} and Authorization: Bearer <api_key>, answered with audio/mpeg. The self-hosted server in services/tts-server/ implements it, and so can any small adapter you write in front of another engine.

Provider Endpoint (default) Auth Per request Notes
generic your api_url Authorization: Bearer 20,000 chars the contract above
openai https://api.openai.com/v1/audio/speech Authorization: Bearer 4,096 chars model default gpt-4o-mini-tts, voice default alloy; instructions steer delivery
elevenlabs https://api.elevenlabs.io/v1/text-to-speech/{voice} xi-api-key split at 4,500 voice (the voice id) is required; model default eleven_multilingual_v2; MP3 44.1 kHz 128 kbps
google https://texttospeech.googleapis.com/v1/text:synthesize X-Goog-Api-Key 5,000 bytes voice is a voice name such as pl-PL-Wavenet-A; audio arrives base64 in JSON

Longer articles are split at sentence ends under the limit, and the parts are joined into one file. max_chars overrides the split size.

Services without a simple API-key request (Amazon Polly signs requests with SigV4, Azure Speech takes SSML) are reached through generic with a small adapter in front.

Templates

The block is placed automatically:

  1. into <span data-ssg-listen-slot></span>, if the theme has one. The bundled themes put it next to the date or the reading time;
  2. otherwise after the first </h1> of each selected page.

listen.auto: false turns automatic placement off. A slot that receives nothing is removed, so a site with both features off is byte-for-byte unchanged. An ssg older than 1.8.66 leaves the empty <span> in place, which is harmless. That is why the bundled themes use a slot rather than a template function.

{{ listen .Post }} (or .Page) renders the same block where a theme calls it: the button, or the <audio> bar with player: full, or nothing. The page also exposes .AudioURL (site-relative, e.g. /audio/2026-10-01-hello.mp3) and .AudioLength (bytes), for a theme that builds its own player.

Without JavaScript the button stays hidden. On a page with an MP3, a plain "Listen (MP3)" link takes its place.

The podcast feed

With tts.feed: true, podcast.xml lists every page that has an MP3, newest first, as RSS 2.0 with <enclosure> elements and the iTunes namespace. Podcast apps and RSS readers subscribe to it directly. URLs are absolute on domain.