Memory¶
Sessions give an agent history within a conversation; nothing carries
across conversations — the user re-explains their preferences every chat.
The Memory plugin adds long-term memory built from two tiers and three
verbs the model already understands:
- Notes (hot) — a small, char-budgeted block of durable facts,
always injected into the system prompt. Curated with
remember(fact)/forget(fact). - Archive (cold) — a full-text-searchable store of past conversations,
pulled in only on demand with
recall(query).
from lovia import Agent, Memory
agent = Agent(
name="assistant",
model="<model>",
plugins=[Memory("./.lovia/memory")],
)
That's the whole zero-config setup: stdlib SQLite full-text search, no services, no embeddings — and it recalls surprisingly well, because the LLM covers the lexical gaps (below).
What lands on disk¶
.lovia/memory/
├── MEMORY.md # hot tier: one `- [YYYY-MM] fact` per line — human-editable
├── MEMORY.md.bak # the previous Notes, written before each dream (one-level undo)
├── .dreamed # mtime marker: when the last dream ran
├── archive.db # cold tier: keyword index of past conversations
└── vectors.db # cold tier: vector arm (only with embedder=)
MEMORY.md is deliberately plain markdown: open it in an editor, delete a
line, done. Writes are atomic, and non-bullet lines are ignored on load.
Privacy. The archive persists user and assistant message text to disk. Keep the memory directory under appropriate access control, and pass
index=Noneto keep no searchable record of conversations (Notes only —recalldisappears).
How memories get written¶
Three paths, from automatic to manual:
- Auto-curation (
auto_curate=True, default). At each run's end (RunCompleted), one digest call over the complete transcript promotes the few durable facts into Notes — date-stamped[YYYY-MM], admission is deliberately strict (no re-derivable state, no one-off task details, most runs yield nothing) — retires existing notes the conversation disproved, and writes a self-contained episode summary into the archive, where it searches far better than raw chat fragments. Raw user/assistant messages are indexed too; document ids are deterministic (run_id:seq), so a replayed run upserts instead of duplicating. - The model, mid-run —
remember/forgettools, reserved for explicit asks and corrections ("remember that I…"); everything else rides the run-end digest. Detail lives in the archive, not in Notes. - Your code — the same verbs are public methods, no model in the loop:
mem = Memory("./memory")
await mem.remember("Prefers concise answers in Chinese.")
await mem.forget("old preference")
body = await mem.notes_body() # editor read
await mem.replace_notes(edited_body) # editor write (normalizes + dedups)
The web UI's sidebar Memory editor (GET/PUT /api/memory) is built on
that last pair.
Dreaming¶
Any always-append memory bloats: near-duplicates pile up, corrections
coexist with what they corrected, one-off details fossilize. So the plugin
periodically dreams — one model call rewrites the whole list: merging
near-duplicates into one well-phrased note, resolving contradictions
(newer wins, by the [YYYY-MM] stamps), dropping event-like or expired
entries — while never inventing a fact and never dropping your explicit
rules and preferences.
Three triggers:
- Cadence — every
dream_every_days(default 3), checked at run end and skipped while Notes haven't changed since the last dream. The state is just the mtime of the.dreamedmarker. - Budget overflow — Notes past
notes_budget(default 8000 chars; the model sees a meter) dream immediately. - On demand —
await mem.dream()returns(before, after)note counts. Wire it to a cleanup button, or run it once to tidy a bloated legacy file.
Before each rewrite the previous body lands in MEMORY.md.bak — a
one-level undo. Deeper history is what git and the archive are for: the
archive keeps the conversations every note came from.
By default curation runs inline — when Runner.run returns, memory is
settled. A long-lived host passes curate_in_background=True so the run's
final event isn't held back by curation's model calls, and awaits
mem.drain() on shutdown to settle anything in flight (the bundled web
server does exactly this, with a 15s bound).
Recall quality, one argument at a time¶
Memory("./memory") # stdlib keyword search (FTS5 bm25)
Memory("./memory", embedder=OpenAIEmbedder()) # + semantic arm → hybrid recall
Memory("./memory", index=my_index) # bring your own retrieval engine
Zero-config is SQLite FTS5 — bm25 over a CJK-aware bigram index — and
the LLM covers what keywords miss: recall queries are expanded with
synonyms and translations before searching (expand_query="auto" turns
this on exactly when the index is lexical-only), and hits come back as a
model-written summary rather than raw excerpts (summarize_recall=True).
Both LLM assists fail open: an expansion or summary error degrades to the
raw query / raw hits.
embedder= upgrades the default index to a keyword | vector hybrid
fused by Reciprocal Rank Fusion — semantic and cross-lingual recall with no
new services (vectors live in SQLite). OpenAIEmbedder speaks any
OpenAI-compatible /embeddings endpoint:
Chat and embeddings often live on different hosts, so the embedder reads
LOVIA_EMBEDDING_BASE_URL / LOVIA_EMBEDDING_API_KEY first, falling
back to the chat endpoint's OPENAI_BASE_URL / OPENAI_API_KEY. Changing
embedders is safe: vectors are a recall cache keyed by embedder id — a
mismatch wipes and re-accumulates rather than mixing spaces.
index= replaces retrieval outright. An Index is three methods over
plain docs — add / remove / search, upsert by Doc.id — implement it
over Elasticsearch, pgvector, whatever:
class Index(Protocol):
async def add(self, docs: list[Doc]) -> None: ...
async def remove(self, ids: list[str]) -> None: ...
async def search(self, query: str, k: int = 5) -> list[Hit]: ...
Compose arms with | — KeywordIndex(...) | VectorIndex(...) | my_arm is
one RRF-fused HybridIndex whose reads fail open (a broken arm is skipped)
and whose writes go to every arm; any index gains the | operator by
mixing in Fusable. (embedder= and index= are mutually exclusive — the
embedder is sugar for building the hybrid.)
Likewise the hot tier: NotesStore is two methods (load/save a fact
list); all normalization, dedup, and budget policy stays in the plugin, so
a Redis- or DB-backed store is a dozen lines (FileNotesStore — the
MEMORY.md writer — is the reference one).
Configuration reference¶
| Field | Default | Effect |
|---|---|---|
root |
./.lovia/memory |
where the default stores live (ignored for a tier you pass explicitly) |
notes |
None → MEMORY.md file store |
hot-tier backend |
index |
default keyword index | cold-tier backend; None disables the tier and the recall tool |
embedder |
None |
adds the vector arm to the default index |
auto_curate |
True |
run-end digest: facts → Notes (disproved notes retired), episode summary → archive |
curate_in_background |
False |
don't hold the run's completion for curation; pair with drain() |
expand_query |
"auto" |
LLM query expansion; auto = only for the lexical-only default index |
summarize_recall |
True |
recall returns a model-written summary of the hits |
recall_k |
5 |
hits retrieved per recall |
notes_budget |
8000 |
char budget for Notes — the prompt meter and the dream's overflow trigger |
dream_every_days |
3 |
cadence for the periodic dream; None keeps only the overflow trigger |
model |
host agent's model | model for the curation/recall side-queries |
The side-queries (digest, dream, expansion, summarization) dogfood
Runner.run with a tool-less, plugin-less sub-agent at temperature 0 — so
they reuse your provider chain and cannot recurse (the sub-agent has no
Memory plugin). Because lovia's transcript is durable and
compaction is view-only, the digest runs once over the
complete transcript: it is curation, not rescue.
Sharp edges¶
- Custom backends are shared by every run — possibly concurrently. They must be safe for concurrent use, and the plugin never closes them; their lifecycle belongs to whoever created them. (Notes read-modify-write is serialized by an internal lock; SQLite stores serialize internally.)
- Curation costs a model call per run (two when a dream is due). On
high-volume, low-value traffic, set
auto_curate=Falseand rely on theremembertool, or pointmodel=at a cheaper model. - Dream state lives under
root— the.dreamedcadence marker and theMEMORY.md.bakbackup are plain files there, even whennotes=is a custom store (the backup then holds an export of the store's facts). - Background curation is best-effort. A process that exits without
drain()can lose the last run's curation — acceptable by design (the transcript is still in the session), but surprising if you expected durability. recallis only as good as what was archived.index=Noneearlier means those conversations are simply not findable later; the archive starts when you turn it on.
See also¶
- Plugins — Memory is the flagship cross-run-state plugin
- Sessions & checkpoints — within-conversation persistence, and where the transcripts come from
- Web UI — the bundled Memory editor
- Example:
23_memory.py