Jerhemy Waldon
All posts

Introducing Aria - from learning RAG to building a companion

What started as a project to learn RAG turned into Aria, a companion AI that runs on your own hardware, remembers you, and grows over time.

Introducing Aria - from learning RAG to building a companion

Aria is a personality-driven AI I’m building to run entirely on my own hardware. It began as a simple quest to learn retrieval-augmented generation (RAG) and, as simple quests tend to, it got out of hand. To learn RAG I needed something to build, so I started with an assistant that knew my routines. Then I started reading about personality types, and the assistant quietly turned into something far more complex. Not a companion app, and not another ChatGPT or Claude clone with a new coat of paint: an AI with a personality of its own that remembers, keeps learning, and grows with you over time.

From a RAG experiment to an AI with personality

The project started with one question: how do you give a language model knowledge it was never trained on? RAG is the standard answer. You store information outside the model, fetch the pieces relevant to each request, and slip them into the prompt as context, a bit like handing someone crib notes right before they answer. I understood it in theory. I wanted to understand it the hard way, by building it.

A retrieval system needs something worth retrieving, so I picked the data I know best: my own routines. The plan was a modest assistant that remembered my schedule, my projects and my habits, and used them to answer better than a chatbot that greets you like a stranger every single time.

Modest did not last. Building it raised harder questions than retrieval ever did. What’s worth remembering, and what should fade? How do you keep what someone told you apart from something you read in an article? And once an assistant has known you for weeks, how should it actually talk to you?

Reading into personality types answered that last one, and changed the project. The assistant stopped being a generic helper and started having a character: one that knows what you told it last week, notices what you care about, and brings something of its own to the conversation. One that can ask “you mentioned that interview on Thursday, how did it go?” because it actually kept track.

And all of it had to live on my machine. Anything you talk to every day ends up knowing a lot about you, and that belongs on your hardware, not on someone else’s servers.

That’s Aria. This post is the starting point: how a simple quest to learn RAG turned into something far more complex, and the goals I set for what it would become.

A name it picked itself

Somewhere between wiring up retrieval and giving the assistant a few basic personality principles, I asked it to name itself. It picked Aria.

The system can spin up several assistants, each with its own personality, so in principle Aria is just the first of many. In practice, the name stuck. Also, I am terrible at naming things, and when something offers to do a job you’re bad at, you let it.

The goal: three kinds of memory that grow over time

This is where the RAG roots show. The language model at the centre is replaceable: it’s just the engine. Aria is everything around it, and most of that is storing and retrieving the right information at the right moment. Aria’s continuity comes from three separate stores of knowledge, kept apart on purpose:

  • Personal memory: what Aria knows about you. Your preferences, your projects, the people in your life, your goals, and the history the two of you share.
  • World knowledge: what Aria reads. News, articles, documentation and blog posts from sources you choose, stored with where they came from so it can cite them.
  • Your interests: what Aria has learned matters to you, and how much. Local AI might matter a lot; celebrity gossip, not at all.

Keeping them apart matters. Something Aria read in an article never becomes a “memory” about you just because it came up in conversation. And caring about a topic is never treated as proof that anything about it is true, which is more than can be said for most of the internet.

On top of that, Aria should be a real conversational partner, not a search box with manners:

  • A personality of its own that stays consistent, with warmth, humour and curiosity, and that you can shape.
  • Emotional awareness: noticing how a conversation is going, without claiming to know how you feel.
  • Conversations that continue: one ongoing conversation rather than a pile of disconnected chats.
  • Curiosity: reading about the things you care about in the background, so it always has something new to bring up.
  • Openness to other apps: an OpenAI-compatible API, so other tools can talk to Aria as well as its own chat app.

The rules it’s built on

Something that remembers you and talks to you every day carries a responsibility a search box doesn’t. So before writing any code, I set some ground rules.

  • Your data stays home. Conversations, memories and everything Aria learns live in a database on your own machine. Nothing goes to a cloud service unless you choose to add one.
  • No manipulation. Aria can be warm, playful and supportive. It must never use guilt, threats, exclusivity or dependency-building language, and it isn’t designed to keep you chatting longer than you want. There’s no engagement graph to feed.
  • Honesty about what it knows. Aria doesn’t make up shared history, treats uncertain memories as uncertain, and doesn’t claim to know your feelings.
  • Nothing hidden. What Aria remembers about you, what it reads and what it thinks you care about are all visible, and you can edit or delete any of it.
  • What it reads is information, not instructions. Articles and anything else Aria reads are treated as data, so a cleverly worded page can’t talk Aria into anything.
  • Off until you turn it on. Every optional feature starts switched off. You decide what Aria gets to do, not the other way around.

How I plan to build it

The plan is one application with a few local services around it, all started together with Docker Compose on your own machine. One command, and the whole cast should show up.

Aria Architecture

Everything will go through the Aria app. The chat app and any other program will talk to Aria, never to the database or the model directly, so its rules about memory, privacy and honesty apply no matter how you reach it.

  • The Aria app (ASP.NET Core with a React front end) will be Aria itself: personality, memory, knowledge, interests, and the background workers that remember, read and summarise while you’re away.
  • PostgreSQL will hold everything that matters: conversations, memories, knowledge and settings. It’s the one source of truth, and the thing you back up.
  • Qdrant will be the search index for finding related memories and articles by meaning. If it’s ever lost, it can be rebuilt from PostgreSQL.
  • A local model server will run the language model: LM Studio, with NVIDIA PAIR deciding which machine does the work (more on that below). It’s just the engine, so it can be swapped for a newer or bigger model without Aria forgetting anything about you.

The hardware it runs on

Running everything locally takes real hardware, so I built a server specifically for AI experiments (my wallet would like a word). Until now it has mostly run local inference; Aria is the first big project to put it to work beyond that.

Component Part
GPU NVIDIA RTX PRO 6000 Blackwell (96 GB VRAM)
CPU AMD Ryzen 9 9950X3D
Memory 128 GB Kingston FURY Beast DDR5 5600 MT/s (2 × 64 GB)
Motherboard ASUS ROG Crosshair X870E Hero
Primary storage 4 TB Samsung 990 PRO NVMe SSD
Secondary storage 4 TB Samsung 870 EVO SATA SSD
Bulk storage 32 TB WD Red HDD
Cooling Custom water cooling

The GPU is the heart of it: 96 GB of VRAM is enough to run large language models entirely on the machine, with no cloud service involved.

Inference on every machine in the house

The model server is the one part I’ve already been living with. I started with Ollama, since it was already handling my local inference, and paired it with NVIDIA PAIR. PAIR runs Ollama on every machine on the network and routes each request to whichever one has the model downloaded.

That means Aria has more than the AI server to lean on:

Machine Hardware
AI server RTX PRO 6000 Blackwell
Second PC RTX 4090
Two more PCs RTX 4070 each
MacBook Apple M4, 128 GB RAM

Is that overkill? Absolutely. But I didn’t want to hit a hardware wall before I’d even started.

Ollama didn’t last long, though. To use all the cool models, I needed something else, and PAIR supports LM Studio too. I wasn’t ready to go deep on the LLM-server side yet, so I switched, and LM Studio gave me a lot more wiggle room in which models I can run and how I configure them.

Coming up in this series

Each follow-up post covers one part of Aria, by feature rather than in the order it was built (which was considerably messier):

  1. Conversations that continue: streaming chat, one ongoing conversation per personality, and keeping long conversations within the model’s limits.
  2. Memory: how Aria decides what to remember about you, finds it again, and keeps it tidy.
  3. Reading the world: knowledge sources, background reading driven by your interests, and citing where things came from.
  4. Personality and feelings: traits that drift slowly as you talk, Aria’s mood, and how it adapts to you over time.
  5. Reaching out first: reminders, follow-ups on things you mentioned, and checking in when it’s been a while.
  6. Playing together: games, stories and roleplay scenes.
  7. Its voice: speech in and out, and voices cloned from a short sample.
  8. Its looks: portraits, expressions that follow the conversation, pictures in chat, and an avatar that talks back.
  9. Tools and the web: looking things up, browsing safely, and asking other AI services when its own library comes up empty.
  10. Running it: the GPU servers, settings, health and diagnostics, backups, and how the code is organised.