Post
Two Ways to Read an Article Aloud
People listen to articles while they cook, drive or rest their eyes. A site can help in two ways, and they cost very different amounts. ssg 1.8.65 does both, and you can turn on either one or both.
The browser does it
Every current desktop and phone browser ships a speech engine, exposed to pages as the Web Speech API. A button that hands it the article costs the site nothing. There is no service to run, no audio to store, and the text never leaves the reader's device.
1listen:
2 enabled: true
3 label: "Listen"
The button appears under the post title. It reads one paragraph at a time.
That is deliberate: Chrome stops a single long utterance after about fifteen
seconds, so one utterance per paragraph is what lets a long article finish.
Press it again to stop. It reports its state with aria-pressed, so a screen
reader announces it, and it stays hidden where the API is missing or JavaScript
is off. Nobody sees a button that does nothing.
What this cannot control is the voice. A recent Mac, iPhone or Android phone has good neural voices for the major languages. A minimal Linux desktop may have a robotic one, or none. The reader gets whatever their device has.
The build does it
The second way makes an MP3 of each article at build time, through a text-to-speech API, and puts a player on the page instead of the button. Everyone hears the same voice, the file plays offline, and it can go into a podcast feed.
1tts:
2 enabled: true
3 provider: openai # or generic, elevenlabs, google
4 api_key: $OPENAI_API_KEY
5 voice: alloy
6 jingle_url: https://example.com/intro.mp3
7 feed: true
Four providers are built in:
| Provider | Per request | What you set |
|---|---|---|
openai |
4,096 characters | voice, model, and instructions such as "calm, unhurried" |
elevenlabs |
split at 4,500 | voice (the voice id), model |
google |
5,000 bytes | voice such as pl-PL-Wavenet-A |
generic |
20,000 characters | api_url; any server that takes the JSON below |
Longer articles are split at sentence ends and the parts are joined into one file. Code blocks are left out, because "func main, open brace" helps no listener. The jingle is downloaded once and played before every article.
generic is a small contract you can implement in an afternoon:
POST {"text", "voice", "lang", "speed", "format": "mp3"} with a bearer key,
answered with audio/mpeg. The repository ships a server that implements it.
A TTS server on your own machine
services/tts-server is a Go API in a Docker image. It uses espeak-ng, which
lists about 140 voices across its languages, and Piper for neural voices. You can bake voices into the
image at build time or upload them to a running server.
1cd services/tts-server
2cp .env.example .env # set TTS_API_KEYS and TTS_ADMIN_KEYS
3WITH_VOICES=en_US-lessac-medium,pl_PL-gosia-medium docker compose up -d --build
4open http://127.0.0.1:8080/docs # Swagger UI
Then point the site at it:
1tts:
2 enabled: true
3 provider: generic
4 api_url: http://127.0.0.1:8080/v1/speech
5 api_key: $TTS_API_KEY
On a 16-core desktop, the Polish Piper voice read 5,433 characters, about five and a half minutes of speech, in 13.1 seconds. The result was a 2.6 MB MP3. espeak-ng is faster still and sounds like the machine it is. Check the voice's licence before you publish what it says: Piper voices are trained on datasets with their own terms, and the server's README links them.
Paying once
A TTS API charges per character, so the build must not send the same article twice. Each MP3 is cached under a key made from everything that changes the sound: provider, voice, model, speed, language, jingle and text. A rebuild calls the API only for articles whose words changed. In our test, the second build of a two-article site printed:
1🎧 audio: 0 made, 2 from cache, 0 stale, 0 without
Keep .ssg-cache/ between CI runs. Without it, every deploy pays for every
article again.
When the API is down
A build that depends on someone else's service needs a plan for when that service does not answer. This one has three parts.
- Retry what is worth retrying. A timeout, a network error,
429or a5xxis tried again after 1 s, 2 s, 4 s, or after whatRetry-Afterasks for. A401or an unknown voice is not, because asking again changes nothing. - Stop asking. After three articles fail in a row, the API is not called for the rest of the build. A down API would otherwise cost the timeout times the retries times every article on the site.
- Decide what the page gets.
on_failure: stale, the default, publishes that page's last good MP3, even if the text has changed since.skippublishes no audio.failstops the build, andstrict: truedoes the same. A page left without audio still gets the browser button.
We stopped the server, edited one article and rebuilt. The build said why once, used the old reading for the edited page and the cache for the other, and finished in three seconds:
1⚠️ audio for /2026/10/01/hello/: … connect: connection refused
2🎧 audio: 0 made, 1 from cache, 1 stale, 0 without
The podcast feed
tts.feed: true writes podcast.xml: RSS 2.0 with an <enclosure> for every
MP3 and the iTunes tags podcast apps read. Subscribe to it in any podcast app
and the blog arrives as episodes, with the jingle at the start of each.
Which one to use
Use the button if you want listening at no cost and do not mind each reader hearing their own device. Use the MP3 if the voice matters, if people listen away from a screen, or if you want a podcast. Use both if you want the button to cover the pages the API could not do.
Configuration reference: docs/AUDIO.