Guide
Articles read aloud
ssg can make an article listenable in two ways. They work alone or together.
Browser button (listen:) |
Build-time MP3 (tts:) |
|
|---|---|---|
| Where the voice comes from | The visitor's device (Web Speech API) | A TTS API, called while the site builds |
| Cost | Nothing | The API's price, once per changed article |
| Same voice for everyone | No: whatever the device has | Yes |
| Works offline, in a podcast app | No | Yes, and there is a podcast feed |
| Needs JavaScript | Yes | No: a plain <audio> element |
| Leaves the device | Nothing | The article text goes to the API at build time |
Both appear as the same small Listen button with a speaker icon, next to the reading time in the bundled themes. On a page with an MP3 it plays the file; on a page without one it reads the text with the browser's voice. That includes a page whose MP3 could not be made because the API was down, so a failed build-time call still leaves the reader a way to listen.
The browser button
1listen:
2 enabled: true
3 label: "Listen" # button text
4 voice: "" # preferred browser voice, by name: "Natural", "Google UK English Female"
5 player: compact # compact (default): the button plays the MP3; full: the browser's <audio> bar
6 sections: [posts] # posts (default), pages, or both
7 auto: true # place it where the theme did not; default true
The button reads the article with speechSynthesis in the page's language
(<html lang>), one paragraph at a time. Chrome stops a single long utterance
after about 15 seconds, so one utterance per paragraph is what lets a long
article finish. Pressing it again stops. It reports its state with
aria-pressed, and it stays hidden in browsers without the API and when
JavaScript is off, so nobody sees a button that does nothing.
The voice is chosen, not left to the browser's default, which is usually the
oldest one installed. Among the voices for the page's language the button
prefers one whose name contains listen.voice, then neural voices (Edge
"Natural"/"Online", Chrome "Google", Apple "Premium"/"Enhanced"/"Siri"), then an
exact language-and-region match.
Voice quality still depends on the device. Recent macOS, iOS, Android and Windows ship good neural voices for major languages; a minimal Linux desktop may have only a robotic one, or none.
The MP3
1tts:
2 enabled: true
3 provider: generic # generic | openai | elevenlabs | google
4 api_url: http://localhost:8080/v1/speech
5 api_key: $TTS_API_KEY # from the environment; never a literal
6 voice: en_US-lessac-medium
7 lang: "" # default: each page's language
8 speed: 1.0
9 jingle_url: https://example.com/intro.mp3 # optional, played before every article
10 sections: [posts]
11 dir: audio # → /audio/<page>.mp3
12 timeout: 60s # per request
13 retries: 2 # extra attempts on 429, 5xx, network errors
14 breaker: 3 # failed articles in a row → stop calling
15 on_failure: stale # stale | skip | fail
16 feed: true # podcast feed
17 feed_path: podcast.xml
18 feed_title: "Example, read aloud"
19 feed_author: "Example"
20 feed_image: /images/podcast.png
21 feed_limit: 50
What is read: the title, then the article text with code blocks removed and
markup stripped. A listener does not need func main() { spelled out.
Cache
Each MP3 is cached in .ssg-cache/tts/ under a key made from everything that
changes the sound: provider, endpoint, model, voice, speed, instructions,
language, the jingle and the text. A rebuild calls the API only for articles
whose words changed. ssg cache clean --namespace=tts forgets them all. Keep
.ssg-cache/ between CI runs (see DEPLOYMENT),
or every build pays for every article again.
The jingle is downloaded once per URL and cached next to the audio. It is joined in front of each article's MP3, after the ID3 tags between the two are removed. Both should use the same encoding, for example 44.1 kHz stereo MP3.
When the API does not answer
A TTS API is a network dependency, and a build should not wait on it once it is plainly down:
-
Timeout and retries. Each request has
timeout. A network error, a timeout,429or a5xxis retriedretriestimes, waiting 1 s, 2 s, 4 s… or whatRetry-Afterasks for, capped at 30 s. A4xxother than429(bad key, unknown voice) is not retried, because asking again changes nothing. -
Breaker. After
breakerarticles fail in a row, the API is not called again for the rest of the build. One line says so. Without it, a down API coststimeout × retries × articles. -
on_failuredecides what the page gets:stale(default): the last MP3 that page ever had, from.ssg-cache/tts/last/, even if the text has changed since. If there is none, the page has no audio.skip: no audio for that page.fail: the build stops.
strict: truemakesstaleandskipbehave asfail.
A page without audio still gets the browser button when listen is on.
Providers
generic is ssg's own contract: POST api_url with JSON
{"text", "voice", "lang", "speed", "format": "mp3"} and
Authorization: Bearer <api_key>, answered with audio/mpeg. The self-hosted
server in services/tts-server/ implements
it, and so can any small adapter you write in front of another engine.
| Provider | Endpoint (default) | Auth | Per request | Notes |
|---|---|---|---|---|
generic |
your api_url |
Authorization: Bearer |
20,000 chars | the contract above |
openai |
https://api.openai.com/v1/audio/speech |
Authorization: Bearer |
4,096 chars | model default gpt-4o-mini-tts, voice default alloy; instructions steer delivery |
elevenlabs |
https://api.elevenlabs.io/v1/text-to-speech/{voice} |
xi-api-key |
split at 4,500 | voice (the voice id) is required; model default eleven_multilingual_v2; MP3 44.1 kHz 128 kbps |
google |
https://texttospeech.googleapis.com/v1/text:synthesize |
X-Goog-Api-Key |
5,000 bytes | voice is a voice name such as pl-PL-Wavenet-A; audio arrives base64 in JSON |
Longer articles are split at sentence ends under the limit, and the parts are
joined into one file. max_chars overrides the split size.
Services without a simple API-key request (Amazon Polly signs requests with
SigV4, Azure Speech takes SSML) are reached through generic with a small
adapter in front.
Templates
The block is placed automatically:
- into
<span data-ssg-listen-slot></span>, if the theme has one. The bundled themes put it next to the date or the reading time; - otherwise after the first
</h1>of each selected page.
listen.auto: false turns automatic placement off. A slot that receives nothing
is removed, so a site with both features off is byte-for-byte unchanged. An ssg
older than 1.8.66 leaves the empty <span> in place, which is harmless. That is
why the bundled themes use a slot rather than a template function.
{{ listen .Post }} (or .Page) renders the same block where a theme calls
it: the button, or the <audio> bar with player: full, or nothing. The page
also exposes .AudioURL (site-relative, e.g. /audio/2026-10-01-hello.mp3)
and .AudioLength (bytes), for a theme that builds its own player.
Without JavaScript the button stays hidden. On a page with an MP3, a plain "Listen (MP3)" link takes its place.
The podcast feed
With tts.feed: true, podcast.xml lists every page that has an MP3, newest
first, as RSS 2.0 with <enclosure> elements and the iTunes namespace. Podcast
apps and RSS readers subscribe to it directly. URLs are absolute on domain.