Jerhemy Waldon
Aria

Aria's talking picture, Part 2: Tuning how it moves

A talking head with default settings nods too much and snaps back between sentences. How Ditto's motion settings became a tuning page on the avatar server, with a test bench that renders settings side by side before saving them, and the two Ditto quirks that had to be worked around.

Aria's talking picture, Part 2: Tuning how it moves

Part 1 got Aria’s portrait speaking in about real time. Real time was the hard requirement. Once it was met, the next problem was obvious to anyone watching: how it moved.

Ditto’s default motion is lively for a calm portrait, and in sentence-by-sentence mode each sentence is its own take, so the head reset between them. A talking picture that nods through every sentence and snaps back between them looks more like a puppet than a person.

Ditto has the controls to fix that. They just weren’t exposed anywhere.

Settings live with the model

The first decision was where the settings belong. They’re Ditto’s run options, so they live with Ditto, on the avatar server:

  • Saved on the server (/config/talking.json). Environment variables only provide the first-run defaults.
  • Used for every video Aria asks for. Aria still sends only the picture and the audio, and doesn’t need to know anything about motion.
  • Overridable per video, which is what makes a test bench possible: try settings on one render without saving them.

Aria’s Settings → Talking picture links to the avatar server’s page. If the avatar server runs on another machine, its settings travel with it.

What can be tuned

The Talking picture tab on the avatar server’s page exposes the settings that turned out to matter:

  • Motion detail: the number of diffusion steps for the motion, and smoothing. Fewer steps are faster. Smoothing takes the jitter out.
  • Expression amount: how much the face moves with the speech. A gentle setting (0.6) gives roughly 10–15% less mouth movement, which looks calmer on a portrait.
  • Emotion: Ditto’s eight emotion classes (happy, sad, surprised and so on), blended from its mostly neutral default by a strength.
  • Head motion: how much the head turns, tilts and nods, and angle offsets for its resting pose.
  • Easing: a gentle ease from the resting face at the start of a video and back to it at the end, optionally including the head pose.
  • Framing: the face crop.

Easing is the setting that fixed sentence-by-sentence mode. When every video starts and ends at the persona’s resting face, the cut between sentences lands on the same frame, and it disappears. I measured it on the test portrait: with easing, the last frame of a video comes back to the still picture.

Two Ditto quirks

Exposing the settings uncovered two behaviours in Ditto that had to be worked around rather than just passed through:

  • The head-motion amounts only work with an offset. Ditto’s alpha_* amounts only scale the angles when the matching delta_* offset is also given. So Aria’s avatar server always sends the offsets, even when they’re zero.
  • They scale the absolute angle, not the motion. Lowering the head motion would scale the whole angle, which turns a calmer head toward facing straight ahead, even if the portrait looks slightly to the side. So after the picture is registered, the offsets are set so the motion is scaled around the portrait’s own resting pose.

With that fixed, 0.4 head motion halves the overall motion, and the first frame still matches the still picture. Both of those were measured on the test portrait, because “looks calmer” is easy to imagine and hard to see.

The test bench

Tuning by saving, chatting and waiting for a reply would take forever. So the tab has a test bench: drop in a picture and a short recording, choose settings, and render. Each render shows what was changed for it.

The rendered videos have their own column with a large preview. The newest render plays big, and earlier ones sit underneath as small thumbnails, to switch between or to play all together for a side-by-side comparison. When a combination looks right, save it, and every video of the persona uses it from then on.

The page tries to make comparing easy:

  • Settings show what changed from the saved values, and each can be typed or slid.
  • The settings column scrolls on its own, while the test bench and the preview stay in view beside it.
  • Tabs stay mounted, so a test keeps running if you switch to the pictures tab and back, and the tab is in the address.
  • A bar at the top shows whether the pictures and the talking picture are loaded, and how much GPU memory is in use.

Tuning the talking picture: settings, the test bench and a large preview of renders

A React app on a GPU server

The avatar server’s page started as one HTML file with inline scripts. Two tabs, a settings form with highlighting and per-group counts, and a test bench with results outgrew that quickly.

It’s now a small React app, with the same stack and tooling as Aria’s own web app (React 19, TypeScript, Vite, oxlint and Vitest). It’s built in a Node stage of the avatar server’s Dockerfile and served by the server itself, so the server stays self-contained when it runs on another machine. Its API didn’t change. The setting definitions (ranges, help text, formatting, what counts as changed) are one module with unit tests.

The voice server’s page got the same treatment afterwards, with the same look, so the two GPU servers feel like parts of one project (Voice, Part 1). The shared pieces are copied rather than shared as a package: each server’s image is built from its own folder, and a shared package would mean a wider build context for both, for a few small components.

What I learned

  • Settings belong with the thing they configure. Keeping motion settings on the avatar server meant Aria never had to learn what a diffusion step is.
  • A per-request override makes a test bench free. Saving is one action. Trying is another. Separating them is what makes tuning fast.
  • Measure “looks better”. Whether the last frame matches the still picture, and how much head motion actually changed, are checkable. Impressions aren’t.
  • Small tools deserve the same stack. Moving a one-file settings page to the project’s own React setup made it easy to grow, and the voice server’s page followed for free.

That’s the talking picture. Next, the lightest part of the series: Playing together, Part 1, where Aria plays games.