Jerhemy Waldon
Aria

Running Aria locally, Part 1: The local stack

One port, three containers and a model server you already have. Why PostgreSQL is the source of truth, why Qdrant is only an index, and how Aria switches between Ollama and LM Studio without a restart.

Running Aria locally, Part 1: The local stack

The conversation posts ended with fitting a long conversation into a small model. All of that has to run somewhere, and one of the goals from the introduction was that it runs on my own hardware. No cloud account, no subscription, and nothing about my conversations leaving the house unless I ask for it.

This post covers the stack: what runs, why each piece is there, and the one rule that keeps it all recoverable.

Three containers and one port

The default stack is small:

cp .env.example .env      # set COMPANION_API_KEY and POSTGRES_PASSWORD
docker compose up -d      # aria + postgres + qdrant
  • aria is the application: a .NET service that serves the React chat app, the REST API, the OpenAI-compatible API and the health endpoints, all on one published port (8080 by default).
  • postgres holds everything Aria knows.
  • qdrant is the vector index used for searching memories, knowledge and conversation examples.

PostgreSQL and Qdrant are only reachable on the Compose network. The only thing exposed is Aria itself, behind an API key. The chat app never talks to the database, the vector store or the model server directly. Everything goes through the application, so there’s one place where access is checked.

Everything else is optional and switched on with a Compose profile: the voice server, the avatar server for pictures and the talking picture, and a browser container for reading JavaScript-heavy pages. Those get their own posts later. Optional features are off by default, so the basic stack is just a persona you can talk to.

PostgreSQL is the source of truth

This is the rule everything else depends on: PostgreSQL is canonical, Qdrant is only an index.

Every memory, message, article, relationship value and setting is stored in PostgreSQL first. Qdrant gets a copy of the vectors needed for search, plus a few fields for filtering, and nothing that exists only there.

That sounds like duplication, and it is, on purpose:

  • Qdrant can be thrown away. If the index is lost, corrupted, or built with the wrong embedding model, Aria rebuilds it from the database. Each record carries an index status, and a reconciler picks up anything marked pending.
  • Changes are transactional. A memory and the job that indexes it are saved in the same database transaction. If Aria crashes in between, the job is still there when it starts again. Nothing waits on Qdrant being up to be saved.
  • Deleting means deleting. Removing a memory removes the row and queues the deletion of its vector. Search results are loaded from PostgreSQL before they’re used, and deleted, archived or foreign ones are dropped, so a vector that hasn’t been removed yet can’t bring a forgotten memory back.

The same pattern runs the background jobs. There’s no separate message broker: the queue is a PostgreSQL table, and workers claim jobs with FOR UPDATE SKIP LOCKED, so two workers never run the same job. A job is added in the same transaction as the change that caused it. One fewer service to run, and one fewer thing that can disagree with the database.

The model server is yours

Aria doesn’t ship a model. It expects a model server you already run, and it uses three kinds of model:

Capability What it’s for
Chat Writing replies. The biggest model you can run comfortably.
Analysis Background work: extracting memories, reading the mood, matching replies, reviewing the relationship. Always structured JSON, temperature 0, no thinking. Often a smaller model.
Embedding Turning text into vectors for memory, knowledge and example search.

The application only knows these as interfaces: a chat model that streams, an embedding model, a structured-output model and a model catalog that says what’s available. Which server answers is decided by a router underneath. The rest of the code never knows whether it’s talking to Ollama or something else.

Ollama, LM Studio, or anything OpenAI-compatible

Three providers are built in:

  • Ollama, the original one. By default Aria looks for it on the Docker host, and the stack doesn’t start one (there’s a profile for that if you want it).
  • LM Studio, which I use now.
  • Any OpenAI-compatible server: vLLM, llama.cpp’s llama-server, LocalAI, or a hosted API. These can also be mixed per capability, for example chat on a vLLM server while embeddings and analysis stay on Ollama.

Each provider registers itself with one entry: its name, whether it can be chosen as the engine in Settings, and how its models are built. The router, the Settings page and the startup log all read that registry, so adding a provider means adding a folder and one registration.

Switching engines without a restart

The Settings page has an Engine tab. It picks the model server (one engine answers everything at a time), sets its address and lets you pick models from the server’s own list. Nothing needs a restart: the settings are stored in the database, and the next request uses the new server.

The interesting part is what a switch does to the index. Vectors from one embedding model mean nothing to another, and they often don’t even have the same size. So Aria records which embedding model built the index. When the engine or the embedding model changes, it compares, and if they differ it rebuilds the memory, knowledge and example collections in the background: drop, recreate at the new vector size, mark every record pending, and let the reconcilers re-embed everything from PostgreSQL.

That’s the source-of-truth rule paying off. Changing the embedding model would be a data migration if the vectors were the only copy. Here it’s a progress bar. While the memories are re-indexed, the chat header says so (“Re-indexing her memories · 120 of 333”), the message box waits, and anything sent meanwhile is answered once they’re complete. A reply built on half an index would quietly forget things, which is worse than a short wait.

Choosing the engine and its models in Settings

Context windows

LM Studio taught me something about context windows. It loads models with a small default context, and if a prompt is bigger than that it doesn’t fail, it just works with less. Aria now asks LM Studio for the context length of each loaded model on every availability check and budgets its prompts to that (Part 3 of the conversation posts explains the budget). Load the model with a bigger context in LM Studio and Aria gets more room without changing anything.

Thinking is off for replies by default and never on for analysis. A reasoning model that thinks before every background job is slow and expensive for no benefit: extracting a memory doesn’t need a chain of thought.

Starting up without everything ready

Local stacks start in whatever order Docker feels like. Aria’s startup is built so nothing has to be ready first:

  1. Configuration is validated. Invalid settings stop the app immediately, with a message saying which setting is wrong.
  2. The web server starts at once: the chat app, liveness and diagnostics are available right away.
  3. In the background, Aria connects to PostgreSQL with backoff (1 second up to 30) and applies migrations. Until that succeeds, readiness is false and the status page says why.
  4. The model server is never contacted at startup. Its state is checked when it’s needed.

If the model server is off, Aria still starts, the UI works, and the status page shows the inference components as unhealthy. That matters more than it sounds: on a home machine, the model server is often the thing that isn’t running yet.

One port, one key

Every request needs the API key from .env. The chat app asks for it once and keeps it for the browser tab. The OpenAI-compatible API takes the same key as a bearer token, so other apps can talk to Aria the way they’d talk to any model.

Nothing secret goes into logs or diagnostics. Message content isn’t logged unless LOG_CONTENT=true, error messages from probes are reduced to their type, and model-server errors are never shown in the chat as raw text. The next part is about what you see instead.

What I learned

  • Keep one source of truth and make everything else rebuildable. It turned “we changed embedding models” from a migration into a background job.
  • A database table is a fine queue. For one person’s persona, SKIP LOCKED on PostgreSQL is all the job system I need, and it commits with the data.
  • Don’t depend on start order. Start the parts that can start, and report the rest honestly.
  • Treat the model server as replaceable. Aria started on Ollama and runs on LM Studio now. Moving it was a setting, not a code change.

Next, in Part 2: what happens when parts of this stack break, and how Aria tells you about it without breaking the conversation.