Jerhemy Waldon
Aria

Aria's relationship, Part 3: The relationship debugger

A state machine with twenty-five dimensions is impossible to tune blind. The debugger shows every value with where it came from, every decision with its dice roll, and the exact prompt, and with debug tools on it can inject events, jump through simulated time, run scenarios and edit every rule.

Aria's relationship, Part 3: The relationship debugger

Part 2 described the relationship flow: appraisals, about twenty-five dimensions, two stages, conflict, pacing. Every one of those rules is simple. Together they interact in ways nobody can predict by reading them.

When a persona seems oddly cold after a friendly conversation, there are many possible reasons. An appraisal that misread a joke? A dimension that faded too fast? A rule? So the relationship has a page that answers “why?” for every number in the system.

Seeing the state

The Relationship debug page (System → Relationship) shows, for one persona:

  • Both stages (the bond and the strain), with how long each has held.
  • Every dimension, its current value, how fast it fades, and its value at the last change.
  • Attachment, obsession, the branch check for strong attachment, and any open conflict.
  • Boundaries currently in effect.
  • A timeline of the key values over time.

The relationship debugger: both stages and every dimension

“Why?” for every number

The most useful feature is the smallest one: click any feeling or dimension and it explains itself.

  • A derived feeling (like the contact drive, or how much a silence weighs) shows the parts it’s made of, each with its points, adding up to the value.
  • A stored dimension shows the recorded events that moved it, how much each one moved it, and how much it has faded since.

So “trust is 0.41” becomes “an apology on Tuesday added 6, criticism yesterday removed 4, and it has faded 1 since”. That turns tuning from guessing into reading.

Removing history

Every appraisal and event in “What happened” can be removed. The engine is deterministic, so removing one entry simply replays the rest of the history from the start and recalculates the relationship as if it never happened.

That’s how a misread joke gets fixed: find the appraisal that labelled it as criticism, remove it, and the relationship becomes what it would have been. If the removed entry had set a boundary (“give me some space” misread), the boundary is lifted too.

The same replay powers “Recalculate from history”, which rebuilds a relationship from its recorded events after the rules change. When the pacing rules from Part 2 arrived, every existing relationship was recalculated once with them, from its own history.

Seeing the decisions

The relationship feeds the communication engine, which decides when Aria writes first and how it replies. The reaching-out posts cover it in depth. The debugger shows its side:

  • The drives, term by term: how much each feeling adds to or takes from the wish to write.
  • The policy: quiet hours, the daily limit, boundaries, waiting times, and which of them is blocking right now.
  • Every decision, kept for a week: each time the engine considered writing (or replying), with the drive, the chance, the dice roll, the motive or the rule that blocked it.

With the actual roll next to the chance, “it never writes first” and “it writes too much” stop being impressions: each decision shows whether it was the drive, the policy or the dice.

Seeing the prompt

The Prompt tab shows the exact relationship section that goes into the next reply, with its token count. Since the engine’s state is turned into sentences by rules, this is where you check that the rules say what you meant. “The user is emotionally important to you” at the right moment is good. The same line on day one is a bug.

Changing things on purpose

Everything above is read-only and always available. With RELATIONSHIP_DEBUG_ENABLED=true, the debugger also gets tools that change the state deliberately. Those are what make testing possible without spending weeks chatting:

  • Sliders for every dimension and stage.
  • Presets: initial contact, stable, positive or negative obsession, conflict, distance, and more. One click puts a persona in a state that would take weeks to reach naturally.
  • Injected events (an appraisal or a system event) with the state before and after, so you can see exactly what one apology or one ignored message does.
  • A simulated clock. Move time forward by hours or days and watch dimensions fade, waits run out, and the engine decide whether to write. Nothing is sent while time is simulated. Simulated entries are marked as such and can be removed like any other.
  • Scripted scenarios: a sequence of events played in order, for repeatable tests of a whole arc (a conflict and its repair, say).
  • A prompt playground to try the relationship section against different states.
  • The configuration editor: every weight, rate, cooldown, stage rule, attachment-style modifier and prompt rule, editable and exportable as JSON.

The configuration lives in the database, is validated on save, and applies immediately. Tuning the engine is a matter of changing numbers on a page and re-running a scenario, not editing code and redeploying.

New personas can also start anywhere in a relationship while the debugger is on: pick an attachment level, familiarity, affection and a state (say, negative obsession) in the New persona wizard. Combined with cloning, that gives a safe way to test how a specific personality handles a specific situation.

The communication engine in the debugger: drives, policy and decisions with their rolls

Context traces

One more debug tool is shared with the rest of the app. With the debug tools on, every reply keeps a context trace: the exact prompt that was sent, split into labelled parts with their token counts, plus the budget and what was dropped. A short id under each reply (“ctx 3f9a1c2e”) opens it. Conversations, Part 3 showed what it looks like.

For the relationship, the trace answers the last question: not “what does the engine think?” but “what did the model actually see?”

What I learned

  • Build the debugger before tuning. A rule engine you can’t see into can only be tuned by guessing.
  • Determinism makes history editable. Because the same events always give the same state, removing one event or changing a rule is just a replay.
  • Log the dice. For anything driven by chance, the decision, the probability and the roll together are the only way to tell bad luck from a bad rule.
  • Simulated time is mandatory for slow systems. A relationship that changes over weeks can only be tested if weeks can pass in a second, without sending anything.

That’s the relationship. Next, the part of Aria that surprises people most: reaching out first, and the communication engine that decides when.