Jerhemy Waldon
Aria

Aria's voice, Part 2: When its voice leads

Reading a finished reply aloud feels like a text-to-speech button. Speaking each sentence as soon as it's written, with the text following the voice as captions, feels like talking. How voice-led replies work with streaming, interruptions, pictures and the next reply.

Aria's voice, Part 2: When its voice leads

Part 1 built the voice server: Whisper to listen, Chatterbox, Kokoro and Orpheus to speak, cloned voices and pronunciations. This part is about the chat side, where a technically working feature became one that feels right.

Read-aloud, version one

The first version of “Read replies aloud” did the obvious thing. The reply streamed in as text, as always. When it finished, the text was split into sentences, and they were spoken one after another. To keep the gaps short, the next sentence was synthesised while the current one played, so speech started after the first sentence was ready.

It worked, and it felt exactly like what it was: a text-to-speech button pressed automatically. You read the reply in a few seconds, then listened to it being read to you for thirty. The voice was always behind.

I wanted it to feel like Aria was talking.

Voice-led replies

So the voice now leads, and the text follows.

With “Read replies aloud” on:

  • Each sentence is spoken as soon as it’s written. The reply streams from the model as usual, but instead of showing text immediately, the chat cuts it into sentences as they complete and sends each one to be spoken right away.
  • The text appears as it’s said, word by word, timed to the audio, like captions on the speech. You never read ahead of the voice.
  • The next sentence is prepared while one plays, so there’s no gap between sentences unless the model itself is slower than speech.
  • Code is shown, not spoken.
  • If speech fails or you stop it, the rest of the text appears at once. The voice is never a gate on reading the reply.

The difference is hard to describe and immediately obvious in use. A reply that starts speaking as soon as its first sentence exists, with the words appearing as they’re said, reads as a voice message from someone. The same audio after the text has finished reads as a screen reader.

A voice-led reply: the text appears as it is spoken

Letting it finish its sentence

Voice made some old behaviours feel wrong, because speaking takes much longer than reading.

Messages you send while Aria is talking. Conversations, Part 2 described how messages you send while Aria is writing are kept and answered together afterwards. With voice, there’s a second phase: the reply is written but still being said. Aria now finishes saying its current reply before it answers, the way a person finishes their sentence. (A short “stop” or “wait” still cuts it off mid-sentence.)

One reply at a time. Aria sometimes sends a follow-on message a few seconds after a reply. With voice on, the next reply waits until the current one has been spoken. Two voices talking over each other is worse than a short pause.

Auto-scroll. Captions revealed word by word, a picture arriving and formatting applied at the end all change the height of a reply after it starts. The chat follows the conversation’s real height, unless you’ve scrolled up to read something earlier.

Pictures and the voice

Pictures caused the trickiest problem. Aria can send pictures in the chat (Her looks, Part 2). When a reply leads to a picture, the picture is drawn from a description, and the final message replaces that description with the picture itself, so the scene isn’t shown twice.

With voice-led reading, the description had already been spoken by the time the shortened message arrived. And the shortened text was no longer a continuation of what streamed, so its new ending wasn’t spoken at all.

The fix has two halves:

  • Speech holds when a picture is likely. As soon as the reply looks like it will get a picture, the server sends a speech-hold event. The voice waits for the completed message instead of speaking the stream.
  • A changed ending is spoken from where it changed. If the completed text differs from what streamed (the rare unasked picture, with no hold), the voice continues from the first point of difference.

On top of that, a reply that gets a picture is delivered whole: the turn waits for the picture (up to a limit) before completing, so you see and hear the reply and its picture together.

Voice in, voice out

The microphone completes the loop. Press the microphone, speak, and your words are transcribed into the message box by Whisper, where you can check them before sending. The recording uses the browser’s own recorder (Opus) and goes to Aria’s API, never directly to the voice server.

With both on, Aria can be a conversation you mostly listen to and speak to, with the chat as a transcript.

The settings that stuck

A few settings ended up mattering:

  • “Read replies aloud” in the chat, remembered per browser.
  • The voice, per persona (Part 1).
  • Pronunciations, per user.

Everything else (sentence splitting, preparation ahead, holds) has no setting. It either works or it’s a bug.

What I learned

  • Order matters more than speed. The audio didn’t get faster. Letting the voice lead and the text follow is what made it feel like talking.
  • Never let the voice block reading. Every failure path shows the rest of the text at once. A voice feature that can hide a reply is worse than none.
  • Voice changes timing everywhere. Pending messages, follow-ons, auto-scroll and pictures all had rules written for text that broke once replies took thirty seconds to “finish”.
  • Hold, then reconcile. Holding speech when a change is likely, and speaking from the point of difference when it isn’t, kept the voice and the text telling the same story.

Next, Tools and the web, Part 1: what Aria can do in a reply, how it’s offered only the tools a reply needs, and what happens when it says it did something and didn’t.