Post

Two Ways to Read an Article Aloud

People listen to articles while they cook, drive or rest their eyes. A site can help in two ways, and they cost very different amounts. ssg 1.8.65 does both, and you can turn on either one or both.

The browser does it

Every current desktop and phone browser ships a speech engine, exposed to pages as the Web Speech API. A button that hands it the article costs the site nothing. There is no service to run, no audio to store, and the text never leaves the reader's device.

1listen:
2  enabled: true
3  label: "Listen"

The button appears under the post title. It reads one paragraph at a time. That is deliberate: Chrome stops a single long utterance after about fifteen seconds, so one utterance per paragraph is what lets a long article finish. Press it again to stop. It reports its state with aria-pressed, so a screen reader announces it, and it stays hidden where the API is missing or JavaScript is off. Nobody sees a button that does nothing.

What this cannot control is the voice. A recent Mac, iPhone or Android phone has good neural voices for the major languages. A minimal Linux desktop may have a robotic one, or none. The reader gets whatever their device has.

The build does it

The second way makes an MP3 of each article at build time, through a text-to-speech API, and puts a player on the page instead of the button. Everyone hears the same voice, the file plays offline, and it can go into a podcast feed.

1tts:
2  enabled: true
3  provider: openai            # or generic, elevenlabs, google
4  api_key: $OPENAI_API_KEY
5  voice: alloy
6  jingle_url: https://example.com/intro.mp3
7  feed: true

Four providers are built in:

Provider Per request What you set
openai 4,096 characters voice, model, and instructions such as "calm, unhurried"
elevenlabs split at 4,500 voice (the voice id), model
google 5,000 bytes voice such as pl-PL-Wavenet-A
generic 20,000 characters api_url; any server that takes the JSON below

Longer articles are split at sentence ends and the parts are joined into one file. Code blocks are left out, because "func main, open brace" helps no listener. The jingle is downloaded once and played before every article.

generic is a small contract you can implement in an afternoon: POST {"text", "voice", "lang", "speed", "format": "mp3"} with a bearer key, answered with audio/mpeg. The repository ships a server that implements it.

A TTS server on your own machine

services/tts-server is a Go API in a Docker image. It uses espeak-ng, which lists about 140 voices across its languages, and Piper for neural voices. You can bake voices into the image at build time or upload them to a running server.

1cd services/tts-server
2cp .env.example .env          # set TTS_API_KEYS and TTS_ADMIN_KEYS
3WITH_VOICES=en_US-lessac-medium,pl_PL-gosia-medium docker compose up -d --build
4open http://127.0.0.1:8080/docs   # Swagger UI

Then point the site at it:

1tts:
2  enabled: true
3  provider: generic
4  api_url: http://127.0.0.1:8080/v1/speech
5  api_key: $TTS_API_KEY

On a 16-core desktop, the Polish Piper voice read 5,433 characters, about five and a half minutes of speech, in 13.1 seconds. The result was a 2.6 MB MP3. espeak-ng is faster still and sounds like the machine it is. Check the voice's licence before you publish what it says: Piper voices are trained on datasets with their own terms, and the server's README links them.

Paying once

A TTS API charges per character, so the build must not send the same article twice. Each MP3 is cached under a key made from everything that changes the sound: provider, voice, model, speed, language, jingle and text. A rebuild calls the API only for articles whose words changed. In our test, the second build of a two-article site printed:

1🎧 audio: 0 made, 2 from cache, 0 stale, 0 without

Keep .ssg-cache/ between CI runs. Without it, every deploy pays for every article again.

When the API is down

A build that depends on someone else's service needs a plan for when that service does not answer. This one has three parts.

  1. Retry what is worth retrying. A timeout, a network error, 429 or a 5xx is tried again after 1 s, 2 s, 4 s, or after what Retry-After asks for. A 401 or an unknown voice is not, because asking again changes nothing.
  2. Stop asking. After three articles fail in a row, the API is not called for the rest of the build. A down API would otherwise cost the timeout times the retries times every article on the site.
  3. Decide what the page gets. on_failure: stale, the default, publishes that page's last good MP3, even if the text has changed since. skip publishes no audio. fail stops the build, and strict: true does the same. A page left without audio still gets the browser button.

We stopped the server, edited one article and rebuilt. The build said why once, used the old reading for the edited page and the cache for the other, and finished in three seconds:

1⚠️  audio for /2026/10/01/hello/: … connect: connection refused
2🎧 audio: 0 made, 1 from cache, 1 stale, 0 without

The podcast feed

tts.feed: true writes podcast.xml: RSS 2.0 with an <enclosure> for every MP3 and the iTunes tags podcast apps read. Subscribe to it in any podcast app and the blog arrives as episodes, with the jingle at the start of each.

Which one to use

Use the button if you want listening at no cost and do not mind each reader hearing their own device. Use the MP3 if the voice matters, if people listen away from a screen, or if you want a podcast. Use both if you want the button to cover the pages the API could not do.

Configuration reference: docs/AUDIO.