Jerhemy Waldon
Aria

Running Aria locally, Part 3: Moving, cloning and many personas

One backup file that restores on another computer and checks itself, cloning a persona with everything it knows (or a fresh start), several personas side by side, and one inference budget that keeps replies fast while background work fills the gaps.

Running Aria locally, Part 3: Moving, cloning and many personas

Part 2 was about surviving the stack’s bad days. This part is about the longer run: moving Aria to a new computer, making copies of a persona to experiment with, and running more than one persona on the same model server.

All three come down to the same thing as everything else in the stack: the database is the persona. Get that right and the rest is plumbing.

One backup file

The goal was simple to state: take Aria from one computer to another without losing anything, with nothing installed on either side except Docker.

A backup is one .tar file:

In the backup Why
The database (pg_dump) Everything Aria knows and remembers: conversations, memories, personality, relationship, settings. The source of truth.
The master key Credentials (API keys for outside services) are encrypted in the database with a key kept in its own volume. Without the key they can’t be decrypted.
Generated images Portraits, expressions, pictures sent in chat.
./data LoRA files for the image model.
The Qdrant index Optional. Aria can rebuild it from the database, which takes a while with a large knowledge library, so it’s included unless you skip it.
.env The configuration, used when the new computer doesn’t have one yet.

Models aren’t included. Ollama, LM Studio and the voice server download them again, and they’re far bigger than everything else put together.

Taking a backup stops Aria and the index for about a minute, so the database and the index are captured at the same moment:

.\scripts\aria-backup.ps1          # -> backups\aria-backup-<UTC time>.tar

On the new computer: clone the repository, install Docker, copy the file over, and restore:

.\scripts\aria-restore.ps1 D:\transfer\aria-backup-20261004-154405.tar

A restore that checks before it changes anything

Restoring replaces Aria’s data, so the restore is careful about it:

  • Every part is checksummed, and the checksums are verified before anything is touched.
  • A backup from a newer version of Aria is refused. Older databases are fine: migrations bring them up to date when Aria starts. A newer one would have tables the code doesn’t know about.
  • After restoring, every table’s row count is compared with the counts recorded in the backup. If anything arrived short, you know right away.
  • An existing .env is kept, with a warning if the backup used a master key that the new .env doesn’t have.

The scripts are thin wrappers around a maintenance container defined in the Compose file, so the same backup and restore logic runs on Windows, Linux and macOS.

The backup contains the master key and the configuration, so it’s private by definition. The backups/ folder is excluded from git.

Cloning a persona

The thing I wanted most once Aria had a real history was a way to experiment without risking it. Try a different personality setting, a new prompt section, a relationship preset in the debugger, and see what happens, without touching the persona I actually talk to.

So a persona can be cloned with everything it has: settings, personality, conversations and messages, summaries, memories (with their history and sources), the relationship and its flow, understanding, goals, open threads, personality changes and pictures. The copy is independent. Change it freely and the original never notices.

Doing that in a database with dozens of related tables is less trivial than it sounds:

  • One transaction. The whole copy runs in a single repeatable-read transaction, so it’s a consistent snapshot even if the original is mid-conversation.
  • Ids are remapped. Every copied row gets a new id, and every id column that points at a copied row is translated through an old-to-new map, so a memory’s source still points at the right message in the copy.
  • Columns come from the database schema, not from a list in the code. A new column added in a later migration is copied without anyone remembering to update the clone.
  • Memories get their own vectors. The copied memories are marked pending and the index sweep embeds them again for the clone. Picture files are shared, not duplicated.
  • Some things are deliberately not copied: reminders, event follow-ups and inbox items. Otherwise both personas would send you the same reminder.

Later I added the opposite option: start fresh. A fresh clone keeps the personality, settings and look, but has no conversations, memories or relationship. It meets you for the first time. That turned out to be the better way to test how a personality starts, which is very different from how it behaves after weeks.

Deleting a persona removes everything that belongs to it, including the vectors in the index and pictures no other persona uses. The default persona can’t be deleted. Clone it if you want to experiment.

Cloning a persona, keeping its history or starting fresh

Several personas

Aria is the name of the project and of the first persona, but the app holds as many as you like. Each one has its own personality, conversation, memories, relationship and look. They don’t share memories: what you tell one, the others don’t know.

The sidebar switches between them, each with its own unread count (personas can write first, which a later post covers), and the chat opens the one you last talked to.

Memories live in one shared Qdrant collection, filtered by persona and user on every search, and each result is checked again against PostgreSQL when it’s loaded. A bug in a filter can’t leak one persona’s memories into another’s reply.

One budget for every model call

Several personas on one GPU, each with its own background jobs, raise an obvious question: who gets the model?

Every model call in Aria (chat, embeddings, structured analysis, seeing pictures) goes through one inference budget: a limit on how many calls are in flight at once (INFERENCE_MAX_CONCURRENCY). Within it:

  • Replies come first. Calls are marked interactive or background. A reply you’re waiting for is served before any background job, and a few slots (INFERENCE_RESERVED_INTERACTIVE) are never used by background work at all.
  • Background work adapts. Behind a model server that queues requests, being overloaded doesn’t show up as errors. It shows up as replies getting slow. So Aria watches the time to the first token of each reply. When it’s far above normal, the background share is halved. Then it grows back by one slot every 30 seconds. It’s the same additive-increase, multiplicative-decrease idea TCP uses for congestion.
  • Identical embeddings are shared. A reply embeds its query once and uses it for both memory and knowledge search. The embedding starts early, alongside the mood estimate, so it’s ready when retrieval needs it.

The background worker can run several jobs in parallel (4 by default). Jobs that don’t need a model, like fetching or indexing articles, keep moving even when the model slots are taken.

This setup came from running Aria against a small cluster of Ollama instances behind a router. Point Aria at the router, size the budget to the cluster’s real parallel capacity, and it fills the gaps with background work without slowing down replies. The System page shows the budget, what’s in use, and whether background work is currently backing off.

The inference budget on the System page

What I learned

  • If the database is the persona, a backup is a pg_dump plus the things the database points to. Key, pictures, configuration. Write them down once and the backup script writes itself.
  • Check a restore before and after. Checksums before changing anything, row counts after. Restores fail rarely, which is exactly why they need to say so loudly.
  • Copy by schema, not by list. A clone that reads the columns from the database never falls behind the migrations.
  • Measure what the user feels. Time to first token was a better signal for “the GPU is overloaded” than anything the model server reported.

That’s the end of the three foundations: memory, conversations, and running it all. Next, the series moves to what makes Aria feel like someone: Personality, Part 1, a personality defined by settings rather than prose.