Jerhemy Waldon
Aria

Aria's knowledge, Part 2: Learning what you care about

Interests learned from your own words (not from what Aria brought up), signals that decay at different speeds, an adaptive reader that spends most of its budget on what you like and some on exploring, and the stories, entities and trends that come out of reading many sources.

Aria's knowledge, Part 2: Learning what you care about

Part 1 covered how Aria reads the web safely and cites what it read. A busy feed publishes hundreds of articles a day, though, and reading all of them would waste the model server on things nobody will ever ask about. So Aria needs to know what’s worth reading. That’s the third of the original “three memories”: what you care about.

Interests from conversation

You never fill in an interests form. After each reply, a background job reads your new messages and extracts the topics in them. The interests it builds up are shown on the Interests page.

The rule that matters most here is anti-self-reinforcement: every topic must be grounded in your own words. If Aria brought up astronomy and you politely answered, that counts as a weak signal at most. Without this rule, a recommender built on conversation turns into an echo chamber of the persona’s own interests in no time: it mentions a topic, you reply, the topic gets stronger, it mentions it more.

Some statements are explicit:

  • “Keep me updated on Formula 1” follows Formula 1.
  • “Don’t show me crypto stuff” mutes it.

Signals that fade at different speeds

An interest isn’t a single number that’s set once. It’s computed from evidence, each signal with a weight and a date:

  • Explicit evidence (statements, follow and mute, “more or less like this” on an article, your manual adjustments) has a half-life of a year.
  • Implicit evidence (topics you raised, or asked follow-up questions about) has a half-life of 30 days.

The affinity is the tanh of the sum of decayed, weighted signals, from −1 to 1, with a confidence that grows with the amount of evidence. Followed topics never drop below 0.8, muted ones never rise above −0.9. A periodic job recomputes every interest from its signals, so interests drift with you: the thing you talked about all summer fades by winter unless you keep bringing it up.

On the Interests page you can see the evidence behind each interest, follow or mute it, nudge it up or down, reset it (which deletes its evidence) or delete it.

The Interests page: what Aria thinks you care about, and why

Two kinds of interest

The glossary is careful about one distinction, and so is the code. Topic interests are yours: what you care about, which guides reading. Persona interests are Aria’s: how curious it is about an open thread, for example (Reaching out, Part 2). They’re different numbers with different purposes, and they’re easy to confuse, which is why the glossary gives them separate names.

The adaptive reader

Sources can be read in two modes. Read everything archives every new article (within the rate limits). Interest-driven sources go through the adaptive reader, which picks what to read.

Every discovered article from an interest-driven source gets a read priority:

priority = interest × 0.40 + topic similarity × 0.25 + freshness × 0.15
         + source priority × 0.10 + novelty × 0.10

Novelty is the distance to the nearest excerpt already in the library: an article that’s very similar to things Aria has already read is worth less.

Scores are stored on the article and refreshed every six hours. Each reading run takes the best candidates above a minimum priority, with two rules:

  • Exploration is reserved. A share of each run (KNOWLEDGE_EXPLORATION_RATE) goes to novel articles outside your known interests. Otherwise the reader would only ever reinforce what it already knows you like, and Aria would never find something new to talk about.
  • Muted topics are never read, exploration or not.

Unread candidates are archived after a while. Every article records why it was read (archive, interest, exploration, manual or watched) and its priority, so the library can answer “why did it read this?”

Interests reach the conversation

Interests show up in more places than reading:

  • Retrieval: your interest in an article’s topics adds a little to its score (0.05). Muted topics count as −1, so they only reorder material that’s already relevant. They never make an irrelevant article appear.
  • Reaching out: new articles on followed topics become a reason to write first.
  • Understanding: interests Aria learns as memories, with a topic, become explicit interest signals (Relationship, Part 1).

Reading many sources creates something more useful than a pile of articles. After indexing, three things are derived from the documents (all from the database, so they can be rebuilt):

  • Stories. Each new article’s title and summary are compared with recent articles (within three days) from other documents. A close match that shares a person, organisation or topic joins the same story. A story page shows its articles in time order, each source’s summary, the entities all sources mention, and the ones only a single source mentions. That last part is often the most interesting.
  • Entities. People and organisations are tracked across the library: how many documents and sources mention them, the recent mentions, when they were first and last seen, and a timeline.
  • Trends. Topic counts this week compared with the previous one: rising, falling, new or steady.

In replies, excerpts from different articles in the same story carry a story label, so Aria can say “two sources agree on this, one adds that” instead of blending them into one confident summary.

Insights: a story told by several sources, and trending topics

What I learned

  • Only count the user’s own words. Anti-self-reinforcement is the single rule that keeps a conversational recommender from becoming the persona’s echo chamber.
  • Keep the evidence, not the score. Computing interests from dated signals made decay, explanations, reset and “why?” all trivial.
  • Reserve room for exploration. A reader that only follows interests narrows over time. A fixed share for novelty keeps it surprising.
  • Compare sources, don’t merge them. Stories, and the entities only one source mentions, turned a feed reader into something closer to a briefing.

That’s the knowledge side. Next is a short one about a persona that knows itself: what it can do, what it wants, and when it isn’t well.