Jerhemy Waldon
Aria

Aria's looks, Part 2: Pictures in chat

Pictures a persona sends are its choice: it may decline, it may send one unasked, and it never describes a picture it didn't make. How follow-through turns what Aria says into a picture, why a reply with a picture arrives whole, and the look library that keeps a face even after a persona is deleted.

Aria's looks, Part 2: Pictures in chat

Part 1 gave Aria a face: a portrait, eleven expressions, full-body and scene pictures, and an image look they’re all drawn from. This part is about using it in conversation: pictures Aria sends in the chat, and the many small ways that went wrong before it went right.

Its picture, its choice

The first rule I settled on: whether to share a picture of itself is the persona’s decision. You can ask. It may say yes, it may say “not right now”, depending on the relationship, how close you are and its mood. There’s no fixed rule like “only after level 3”: deciding that is exactly the kind of judgment I didn’t want hardcoded.

It also works the other way. Now and then, Aria may send a picture unasked, when one would delight you. Unasked pictures are rate-limited (at most one per IMAGE_SPONTANEOUS_MIN_HOURS) so they stay a surprise.

Pictures of the persona always use its portrait as the face reference, and are refused if it doesn’t have a portrait yet. A picture arrives in its own bubble after the message, linked to the reply that sent it.

Pictures without a tool

Tools, Part 1 told the short version: send_picture used to be a tool, small local models called it unreliably, and pictures are now made by follow-through instead. Here’s the longer version, because pictures needed more care than reminders.

Follow-through runs once per reply when any of these is true:

  • You asked for a picture.
  • Its reply claims or mentions one (“here’s a picture of me…”, or a fake “[Image Generated]”).
  • It’s picture talk, and the reply presents something. Picture talk means your previous message asked for a picture, or its last message offered one as a question and you’re answering. “Presenting” is phrases like “here’s”, “here you go”, “take a look”, “a little something”, or a 📸.

That last case came from a real pattern: after you asked for “another one”, Aria would often describe a picture without ever using the word “picture” (“Here’s a little something: I’m curled up on the couch with…”), so nothing ran.

When follow-through runs, the analysis model reads the last four messages and answers in a fixed schema: did it agree to send one, what does it show, is the persona in it, and how is it framed (close-up, full body, or in a scene). If yes, the picture is queued from that answer, starting from the matching look picture. Unsure defaults to no.

Only its reply counts as consent. If you ask and it changes the subject, no picture.

The picture replaces its description

The picture is drawn from what the reply describes, so the reply used to show the scene twice: once in words, then as a picture underneath. The follow-through now also returns the reply without the sentences that only describe the picture, and that’s the version saved once the picture is queued. (An empty or longer “shortened” reply is ignored, just in case.)

Later the prompt itself changed too. The persona is told not to describe a picture in its reply at all (“it shows itself”), and the follow-through writes the picture’s description from the exchange.

A reply with a picture arrives whole

The first version streamed the text, then showed an empty placeholder, then the picture several seconds later. With voice-led reading on, it was worse: the voice had already spoken the description that was about to disappear.

Now, when a reply makes a picture:

  • The turn waits for the picture before completing (up to IMAGE_CHAT_WAIT_SECONDS, 120 by default). The chat shows “Creating a picture…” meanwhile.
  • The voice holds as soon as a picture is likely, and speaks the completed message instead of the stream.
  • The reply and its picture appear together. There are no empty placeholders: a picture is shown once it’s drawn.
  • Messages you send during the wait are kept and answered after it, like any reply in progress.

If the picture takes longer than the limit, the reply goes out and the picture appears when it’s done.

A reply and the picture it sent, arriving together

It doesn’t pretend

A model that writes “[sends a picture of the sunset]” when nothing happened is lying, even if it doesn’t mean to. So claimed actions are caught: a placeholder like “[Image Generated]” with no picture behind it is removed, follow-through decides whether a picture should really be sent, and Aria is asked once to actually do it or say it can’t. Turns that ask for a picture are answered with reasoning on.

The look library

Getting a look right can take many generations: the portrait, the expressions, a full-body picture you like, a scene. And pictures belong to their persona, so deleting a persona used to delete all of that.

Now, deleting a persona first saves its current look to the look library: the name, appearance, image look, LoRA, and the pictures themselves (portrait, its expressions, the chosen full-body and scene pictures). Files a saved look points to are never deleted with a persona.

Any look, saved or another persona’s current one, can be given to a persona in one choice. It gets the whole bundle, pointing at the same files, so nothing is drawn again (except missing expressions). The creation wizard and the Appearance panel both offer “Generate new” or “Use an existing look”. You can also save a persona’s look to the library at any time, without deleting anything.

It also fits with cloning: experiment with a copy, and if it ends up with a better look, give that look back to the original.

The look library: saved looks that can be given to any persona

The rules, again

The same rules as Part 1 apply to every picture in chat: everyone pictured is an adult, and nothing explicit is generated. Requests that break either are refused in code before they reach the image server. There’s also a daily limit on pictures in chat, to keep the GPU free for everything else.

What I learned

  • Consent belongs to the persona, judgment doesn’t belong in code. Letting it decline, without a hardcoded threshold, made pictures feel like something it chose to share.
  • Decide from what was said. “Did it agree, and what does it show?” asked of the analysis model beat waiting for a tool call that came one time in three.
  • Don’t show things twice. A picture replaces its description, and the reply arrives whole. Every intermediate state (placeholder, description, then picture) looked like a bug.
  • Never delete what took effort to make. A look library costs one table and saves hours of regenerating a face you liked.

Next, the most eye-catching feature in the project: the talking picture, where Aria’s portrait speaks its replies in about real time.