Agent Memory

Engram: Encrypted Local Memory for AI Agents

Engram by EL AI Intelligence is a memory daemon that runs on your machine, keeps its vault encrypted with SQLCipher, and exposes one JSON API that MCP clients, chat bots, editors, and the browser all talk to. Here is how the vault, the deletion model, and the open format actually work.

Editorial illustration of an encrypted vault door built into a laptop, with memory threads flowing into a knowledge graph.
Generated editorial hero for this guide; not Engram artwork.

What Engram Is

Engram is a memory daemon for AI tools. It runs on your machine, stores memories in an encrypted local vault, and answers every client through one JSON API — MCP servers for coding agents, Slack/Discord/Telegram bots, a browser extension, and the CLI. The product’s own framing is blunt about the problem: AI assistants forget everything the moment a conversation ends, and the fix is a long-term memory that stays on your computer.

The company behind it is EL AI Intelligence, which describes itself as building deterministic infrastructure for AI products — Engram for memory, Guardrail as a policy engine, Amparo as an agent that acts under policy, and ELLM as the engine underneath. Engram is the piece this guide covers.

Current release: 0.3.0, published September 12, 2026 according to the product’s own version endpoint. The tagline the project uses is “Encrypted, Local, Retrievable Private Memory for AI & AI Agents,” and the three load-bearing claims are narrower than that sounds: an encrypted vault at rest, deletion that propagates rather than hides, and a storage format that is open enough to outlive the product.

Official Engram social card reading Engram, Memory for AI, One JSON API, encrypted vault, deletes that actually delete.
Official Engram social card, from the project’s own site. It names the three claims this guide tests: one JSON API, an encrypted vault, and deletes that actually delete.

Install and First Memory

Installation is one command per platform, followed by an onboarding step that creates the vault and writes a first memory in roughly five minutes:

# macOS / Linux
curl -fsSL https://engram.ellmstack.dev/install.sh | bash
engram onboarding

# Windows (PowerShell)
irm https://engram.ellmstack.dev/install.ps1 | iex
engram onboarding

# Windows package manager
choco install engramd

Windows deserves a note the project states plainly rather than hiding: the binaries are unsigned, and the installer fetches them directly so SmartScreen does not fire. If you download a binary in a browser instead, you will hit “More info → Run anyway.” Unsigned distribution is a real supply-chain caveat, not a cosmetic one — the practical mitigation is that the install path is a script you can read before running, and the daemon is local-only by default.

One packaging caveat straight from the project: the Chocolatey package is published but moderation-gated, so the public feed can still serve an older build while a new one clears review. The vendor's own guidance is to treat a freshly pushed version as in flight rather than live.

Uninstall is symmetric and provided as a script rather than a manual cleanup ritual, which matters for a daemon that holds your notes:

# macOS / Linux
curl -fsSL https://engram.ellmstack.dev/uninstall.sh | bash

# Windows
irm https://engram.ellmstack.dev/uninstall.ps1 | iex

Update notifications shipped in 0.3.0, handled without a background updater that silently mutates your install. The site publishes a changelog page, a machine-readable version endpoint, and an RSS feed, and the CLI exposes engram update-check plus an update banner in the vault UI. The installer is always the upgrade path; the checks only notify. That is the right shape for a local daemon; plenty of tools get it wrong by bundling an auto-updater.

One JSON API

The architectural decision that shapes everything else is that Engram is a daemon, not a library you embed per tool. Every surface — MCP clients, Python and JavaScript SDKs, chat bots, the browser extension, the CLI — speaks to the same local process over HTTP and JSON. The project’s phrasing is the honest version of a marketing line: if your stack can POST JSON, it has memory now.

That choice matters in daily use. The alternative is per-tool memory, where your editor has its own store, your chat app has another, and your CLI has a third. Nothing consolidates, and nothing can be deleted coherently because deletion has to be re-implemented in every tool. A daemon inverts that: capture once in one client, recall it in another, and keep a single deletion path.

The Vault: SQLCipher at Rest, Local-First

The vault is a SQLite database encrypted with SQLCipher — page-level AES-256-CBC plus HMAC, as bundled by rusqlite’s bundled-sqlcipher feature. The key is always derived locally; nothing that can open the vault leaves the device. The project publishes the derivation in its normative format specification rather than describing it in prose, and the two modes exist because the threat models genuinely differ.

Machine-key vaults (v1) derive the key from the platform hardware id:

key = hex( SHA-256( machine_id | ":" | "axiom-engram-vault-v1" ) )

# machine_id is the platform hardware id:
#   Linux   -> /etc/machine-id
#   macOS   -> IOPlatformUUID
#   Windows -> MachineGuid

The spec then states the limitation in the document itself, which is unusual candor: this key protects against offline disk cloning and stolen storage media. It does not protect against a local user on the same machine, because the machine id is world-readable and the salt is a public constant. For confidentiality against local attackers, use a passphrase vault.

Passphrase vaults (v2) derive the key with Argon2id over a 16-byte per-vault random salt, falling back to a constant-derived salt only for legacy vaults that predate the change. Choosing the passphrase mode is a real security upgrade, and the honest reading of the two modes is: machine-key vaults defend your laptop if it is stolen; passphrase vaults defend your notes from someone who already has the laptop.

Deletes That Actually Delete

This is the most interesting engineering claim in the product, because deletion is where memory systems usually cheat. The common pattern is a soft delete: the row leaves the UI, the embedding stays in the vector index, the sync server keeps a copy, and the content is recoverable by anyone with database access. Engram’s spec describes three properties that, read together, close those holes.

PropertyWhat the spec saysWhy it matters
AtomicRow, search index, and tombstone are written in one transaction.There is no window where a deleted memory is still retrievable.
Sync-awareThe tombstone rides the sync relay, so every device forgets together, and the relay physically erases the ciphertext.Deletion reaches other devices and the hosted copy, not just the machine you deleted from.
AuditableA local access ledger records every read and delete.You can see who touched what, including after the memory itself is gone.

Release 0.3.0 turned those three properties from a description into shipped behaviour, and the details are worth knowing because they are what makes the claim checkable. Deletes write a database tombstone in the same transaction as the row removal (schema v8), and deletes cascade through consolidation outputs, so links and embeddings derived from a deleted memory go with it rather than lingering as orphaned edges. The access ledger is served over HTTP at /audit/events: it reports which memories a client touched since a given timestamp, and which were deleted since then, exposing ids and hashes but never content. Clients can label their own traffic with an X-Engram-Client header, which is how the ledger names the tool that did the touching. On the sync side, deleted blobs are garbage-collected from the relay rather than left in cold storage indefinitely, with a 30-day retention window before the ciphertext is gone.

On the wire, a deletion is a tombstone: a blob with deleted = true and a higher vector clock, with empty content. Last-write-wins on the vector clock is what makes the deletion stick instead of racing a stale device that still holds the old row. The push response reports accepted counts, rejected ids for stale clocks, and any revoked devices the relay refuses.

One limitation is stated rather than glossed: the relay can block a device’s pushes, but only a re-key removes the passphrase from that device. In other words, revoking a device stops it from writing back — it does not retroactively scrub what that device already decrypted. That is a property of end-to-end encrypted sync in general, and it is the kind of sentence that tells you the spec author is thinking about attackers rather than reviewers.

MCP: Claude Code, Cursor, Windsurf

For coding agents, the integration is a single command that merges an engram entry into every supported editor’s MCP config, alongside whatever servers you already have:

engram mcp install

# Per client
claude mcp add engram -- engramd-mcp          # Claude Code
# Claude Desktop -> merged into claude_desktop_config.json
# Cursor         -> merged into ~/.cursor/mcp.json
# Windsurf       -> merged into ~/.codeium/windsurf/mcp_config.json

The design decision here is that MCP capture is explicit — the comparison table on the product site lists “MCP capture (explicit) — on call, never automatic” as a differentiator. That distinction is easy to miss, and it separates a memory system you chose to write to and a background scraper that records everything your agent touches. Anyone who has tried to audit what a self-storing agent kept will recognize why this matters.

If you are new to wiring MCP servers at all, two of our existing guides cover the plumbing: how to use MCP servers in Antigravity and MCP servers for Claude Code by job. Engram slots into the same config surface as any other MCP server, which is the point of using MCP rather than a per-editor plugin.

Chat Bots and the Browser

Team memory is where a lot of decisions actually happen, so Engram ships chat capture as first-class surfaces rather than a webhook you write yourself. These graduated from planned to shipped in 0.3.0, which introduced a dedicated engram-chat binary alongside the browser extension. The triggers are explicit, which keeps the “never automatic” property intact:

SurfaceCapture triggersTransport detail
Slack/remember, @-mentions, 📌 reactions, opt-in channel watchSocket Mode (pull)
Discord/remember, @-mentions, 📌 reactions, opt-in channel watchGateway with heartbeats (pull)
Telegram/remember, mention-entity detectiongetUpdates long polling (pull)
BrowserCurrent page: selection, title, or linkManifest V3 extension, Chrome/Edge/Brave/Firefox 121+

Two changelog details matter because they are where chat integrations usually break. First, the transports are all pull-based — Socket Mode, the Discord gateway, Telegram long polling — so no inbound port is opened on your machine. Second, bot tokens are read from the environment only, keys never leave your machine, duplicate captures are reported back to the user as an honest skip (“duplicate of …”) rather than silently dropped, and a daemon that is down is logged rather than queued for a retry storm.

The most recent change in this area moved chat captures from a generic chat source to per-platform sources — slack, discord, telegram — as first-class daemon sources in schema v7. That is a small change with a real payoff: per-platform filters work without having to inspect the context string of every memory.

How Engram Models Memory

Engram does not present a flat list of notes. Memories are typed, and the types come from the cognitive-science vocabulary the product is named after. The onboarding tour states the three layers directly, and the format spec confirms them as the layer column:

  • Episodic — things that happened. The tour’s examples: “Deployed the website on Tuesday,” “Had coffee with Sarah.”
  • Semantic — things that were learned. Facts, preferences, and rules: “You prefer short answers,” “Never deploy on Fridays.”
  • Imagined — ideas the AI came up with on its own. These stay quarantined, clearly marked and never treated as fact, until you explicitly approve (“ground”) them.

Every memory also carries two continuous properties the schema stores as real numbers: strength, which is retrieval-weighted and decays when a memory goes unused, and valence, the emotional tone on a scale from challenging to joyful, so an agent can distinguish an experience that went well from one that did not.

Engram onboarding tour showing the three memory kinds, strength and valence, a remember form, and a demo memory loader.
Engram’s onboarding tour: the three memory layers, the strength and valence properties, a manual capture form, and a demo-memory loader. Account details in the original screenshot were redacted.

The imagined layer is the most opinionated design decision in the model. Most memory systems treat everything the model writes as a fact to be stored. Quarantining self-generated content until a human grounds it is how you avoid an agent that slowly builds a private mythology out of its own suggestions. The schema backs the claim with two separate integer flags — imagined and grounded — rather than a single status field.

Forgetting, Consolidation, and QEM

The retrieval side is built on three mechanisms, and the product’s own slogan states the intent: forgetting is green, because remembering everything forever is computationally wasteful.

Ebbinghaus decay. Memories you revisit get stronger; memories you ignore fade. The site visualizes this as a decay curve with and without spaced retrieval, and the spec stores the state as strength (default 1.0) plus a retrievals counter and last_retrieved timestamp, so the curve is computed from real access history rather than an age heuristic.

Holographic associative codes (QEM). 32-bit XOR holographic codes give O(1) associative lookup — subject XOR relation maps to object — implemented as an L1 cache with write-through to the store and a prediction-error novelty filter. The spec is careful about the claim: QEM codes are an optional acceleration layer, derivable from row content, and never the only index.

Nightly consolidation. Episodes merge into semantic knowledge overnight, tracked in a consolidation_runs table so the process is auditable rather than a black box that rewrites your memories.

Underneath, retrieval uses two independent indexes over the memory table: FTS5 for exact and stemmed text search, and a 384-dimension vector index built from all-MiniLM-L6-v2, computed locally through ONNX. Results are ranked with a deterministic total order — score, then creation time, then id — so identical searches return identical orderings. Determinism sounds like a detail until you are debugging why an agent got different context for the same question twice.

There is also a deduplication rule that explains a lot about how the system behaves in daily use. Captures are hashed after normalization: content is lowercased, whitespace collapsed, and shell hook prefixes stripped, then SHA-256’d. Equal hashes strengthen the existing row instead of inserting a duplicate. Near-duplicates above 0.95 embedding cosine are skipped and reported — the human decides on merges rather than the daemon guessing. Tag normalization is similarly opinionated: lowercased, deduplicated, capped, and a denylist of zero-value tags like note and misc dropped silently.

Token Economics, Measured

The cost argument for retrieval over injection is where memory products usually start hand-waving, so the measured numbers and the assumed ones need separating. The benchmark is an A/B test on the same corpus: Engram’s retrieval versus wholesale context injection.

MeasurementRetrieval (Engram)Wholesale injection
Input tokens per question, corpus growing 4× (25 → 100 memories)+4.5%+11%, and rising linearly by construction
Where it breaks downBounded by top-k retrievalFills the entire context window around ~4,000 memories
20-person team, 10 memory-touching questions/day, 1,000 memories≈10M input tokens/day≈14M input tokens/day
Marginal cost per 1,000 memories addedBarely moves+8.4M tokens/day

Two caveats the project states itself, and they are the reason this table is more believable than most: below roughly 100–150 memories, retrieval actually costs a few percent more tokens than injection, because you are paying retrieval overhead to avoid a problem you do not have yet. And retrieval runs locally, so the only datacenter load per query is the tokens your model actually reads — hosted RAG pays for retrieval compute on top.

The carbon numbers on the site come from an interactive calculator with visible inputs — hours per day using agents, average tokens per interaction, and expected context reduction — and the site names its constant: 0.4 g CO₂e per 1,000 tokens, described as a public estimate. The vault dashboard shows the same estimate for your own activity, counted at event time and labelled plainly — “Memory tokens counted” and “CO₂ equivalent” — rather than presented as audited measurement. Treat the headline figure as an order-of-magnitude argument, not a measurement, because the constant is a third-party estimate and your actual reduction depends on how much of the corpus your queries really touch.

Sync Without Trusting the Relay

Multi-device sync is where privacy claims usually collapse, so the wire format carries the weight. Sync is version 5.1, and every blob carries a self-describing envelope:

{
  "vault_id":     "engram-local",
  "memory_id":    "3b…-uuid",
  "device_id":    "4f…-uuid",
  "vector_clock": 17,
  "ciphertext":   "<base64>",
  "hmac":         "<base64>",
  "created_at":   "2026-08-30T12:00:00Z",
  "deleted":      false
}

The cipher construction is specified down to the byte order: the ciphertext is a 12-byte random nonce prepended to AES-256-GCM output, so each stored blob is self-contained, and the HMAC is SHA-256 over vault id, memory id, vector clock, and ciphertext — which binds a blob to its identity and version, so a replayed or moved blob fails verification. Encryption and MAC keys come from separate Argon2id derivations (axiom-sync-enc-v2 and axiom-sync-hmac-v2, m=65536 KiB, t=3, p=4).

Vault identity is derived from the passphrase, so devices converge on the same vault id without a server telling them which vault is theirs — and a device can still pin an explicit id for manually named vaults. Team sharing uses a zero-knowledge handoff: the requester mints an ephemeral P-256 keypair and posts the public key to the relay’s mailbox, an owner encrypts the shared key to the ECDH shared secret, and the requester claims the seal exactly once and unseals it locally. The relay stores only public keys and ciphertext, never a private key, and the seal is erased on delivery.

The protocol itself is deliberately boring: push batches encrypted blobs, pull downloads incrementally from an RFC 3339 cursor with 1,000 blobs per page, and health reports relay counters. Conflicts resolve last-write-wins on the vector clock. For teams, that means the relay is a dumb, replaceable pipe — which is the correct threat posture, and also the reason self-hosting stays free: there is nothing clever to license.

The Open Format

The product is closed source and the format is open, and the project is explicit about where that line sits. The daemon, relay service, browser vault, and MCP servers live in a private repository. What sits in github.com/El-AI-Intelligence/engram-format under Apache-2.0 is the part that has to stay open: the crate axiom-engram (vault storage, retrieval indexes, sync format types) and FORMAT.md, the normative specification covering schema, key derivation, and the wire format.

The repo states the bet in two sentences: a memory is only yours if you can read it, and privacy that is verifiable is not the same as privacy that is promised. Concretely, that means the vault is a SQLite database with a documented, versioned schema — currently schema version 9, with the migration history published — so any tool can open it without reverse engineering. The Rust API is small enough to read:

# Cargo.toml
axiom-engram = { version = "0.1.6" }
# local semantic search (all-MiniLM-L6-v2, 384-dim, zero-config ONNX):
axiom-engram = { version = "0.1.6", features = ["onnx-embed"] }

// Open (or create) a machine-key vault
let store = EngramStore::open("/path/to/vault").await?;

// Or with real confidentiality: a passphrase-protected vault
let store = EngramStore::open_with_passphrase("/path/to/vault", "your passphrase").await?;

The spec is not a summary document; it is normative enough to implement against. A memory row is a full record with a layer, a source drawn from eighteen enumerated values (interaction, sensor, consolidation, imagined, chat, slack, discord, telegram, window, mic, agent, research, system, user, observation, ai-session, ai-tool), a privacy level (strict_local, hybrid, cloud_first, enterprise), the content and its capture context, strength, valence, retrieval count, the imagined and grounded flags, timestamps, project, tags, a dedupe hash, scope, and content type.

Around the memory table sit link, coherence, goal, consolidation, metrics, embedding, evidence, annotation, and saved-search tables — including typed links (associative, causal, analogical, temporal) with weights and foreign-key cascades, so deleting a memory takes its links with it rather than leaving dangling edges in the graph.

Pricing

The flat-rate structure is deliberate — no per-seat charges on any tier — and sync is the only thing you pay for, since the daemon and self-hosting are free:

TierPriceDevicesStorageNotes
Free$011 GiBFull local daemon; self-hosting stays free
Personal$10 / month, flat510 GiBEnd-to-end encrypted sync
Team$29 / month, flat25100 GiBShared team memory, up to 10 members
Organization$99 / month, flat50500 GiBCompany-wide memory, unlimited members

Device management is built into the account surface rather than being an enterprise upsell: a Devices panel with roster, rename, and revoke, matching engram devices list|rename|revoke in the CLI, with device labels that survive daemon restarts. Signup issues a recovery phrase you can download or print and rotate later from a modal.

How It Compares

The comparison below is the project’s own table from its site, reproduced so you can see exactly which claims it is making against Mem0, ChromaDB, and Letta. Read it with the usual caution: self-reported comparison tables pick the axes where the author wins, and the interesting line here is the last one, where Engram is the only one of the four that is not fully open source.

CapabilityEngramMem0ChromaDBLetta
Local-first✓ SQLCipher✓ self-host✓ embedded✗ server required
Ebbinghaus decay✓ built-in✗✗✗
Holographic cache✓ QEM L1✗✗✗
Nightly consolidation✓ episodic → semantic✗✗✓ agent memory
Explicit MCP capture✓ on call, never automatic✗✗✗
End-to-end encrypted sync✓ shipped✗✗✗
Deletion audit trail✓ access ledger✗✗✗
LicenseOpen-core, format Apache-2.0Apache-2.0Apache-2.0Apache-2.0

Two adjacent approaches make useful comparisons, since their trade-offs differ: Claude-Mem sits inside a single agent’s workflow, while Claude Context targets semantic code search rather than conversational memory. Engram’s bet is that the memory layer should be one local service shared by every tool, and the format should outlive whoever maintains it.

Where It Doesn’t Hold

Reading a product’s own documentation is the fastest way to find its real edges, and Engram is unusually candid in several places. The honest limitations, in its own terms:

  • Unsigned Windows binaries. Install via the script or Chocolatey; direct browser downloads require bypassing SmartScreen. If your policy forbids unsigned executables, this is a blocker regardless of how good the encryption is.
  • Machine-key vaults do not protect against local users. The spec says so directly: the machine id is world-readable and the salt is a public constant. Use a passphrase vault if that is your threat model.
  • Revoking a device does not erase what it already decrypted. Push blocking is not retroactive scrubbing; only re-keying removes the passphrase from a device.
  • Retrieval costs more below ~100–150 memories. The efficiency story starts once the corpus is big enough that injection would hurt.
  • Carbon figures rest on a third-party constant (0.4 g CO₂e per 1K tokens) plus your own usage inputs, so treat them as directional.
  • The product is closed source. The format, schema, key derivation, and wire protocol are open; the daemon and relay are not. You can always read your own vault, but you cannot audit the code that writes it.
  • Early release. Version 0.3.0 shipped September 12, 2026, and it is the release that moved chat bots and the browser extension from planned surfaces to shipped ones. Expect churn in the surfaces rather than the format.

None of those negate the design; they bound it. The pattern across all seven is that the team documents the failure mode next to the feature, which is a better signal than a longer feature list.

FAQ

What exactly is Engram?

A local memory daemon from EL AI Intelligence. It stores AI memories in an encrypted SQLCipher vault on your machine and serves every client through one JSON API. Current release is 0.3.0 (September 12, 2026).

Is my memory uploaded anywhere?

No, unless you enable sync. The vault is local and encrypted at rest. With sync turned on, encryption happens on your machine and the relay stores only ciphertext. The vault carries no trackers: the site’s own phrasing is zero telemetry, and the vendor states there are no analytics beyond a user count.

What does “deletes that actually delete” mean in practice?

The row, its search index entry, and a tombstone are written in one transaction, so there is no window where the memory is still retrievable. The tombstone syncs to other devices and the relay erases the ciphertext. A local access ledger keeps a record of reads and deletes even after the memory is gone, and 0.3.0 serves it at /audit/events (ids and hashes only, never content).

Which tools can use it?

Anything that can POST JSON. Officially wired: Claude Desktop, Claude Code, Cursor, Windsurf via engram mcp install; Slack, Discord, and Telegram bots; a Manifest V3 browser extension for Chrome, Edge, Brave, and Firefox 121+; plus Python and JavaScript SDKs and the CLI.

Is the format really open?

Yes — the axiom-engram crate and the normative FORMAT.md specification are Apache-2.0, covering the schema, key derivation, and sync wire format. The product itself (daemon, relay, browser vault, MCP servers) is closed source in a private repository.

Can I self-host?

The project states self-hosting stays free. Sync infrastructure is the paid part, in three flat tiers: $10/month personal, $29/month team, $99/month organization, with no per-seat charges.

Does it work on Windows?

Yes, with a caveat: choco install engramd or the PowerShell installer. The binaries are unsigned, so the script path avoids SmartScreen while a direct browser download will require “Run anyway.”

How is memory kept from decaying into noise?

Three mechanisms: Ebbinghaus-style decay with retrieval-based strengthening, nightly consolidation that merges episodes into semantic knowledge, and an explicit deduplication pass — identical content strengthens the existing row, near-duplicates above 0.95 cosine are skipped and reported for a human decision.

Sources

Everything above comes from first-party material: the product site, its changelog and version endpoint, the open format repository, and the normative specification. Screenshots are official product captures, with account details redacted. This post has no third-party reporting, analysis, or benchmark claims attached to it.

Official Engram logo lockup: the Engram wordmark with the EL Intelligence emblem and product tagline.
Official Engram lockup: the wordmark, the emblem, and the EL Intelligence product mark.

The through-line is simple. Memory is only useful to an agent if it can be recalled cheaply, and it is only safe to give an agent memory if you can delete it completely and prove what happened. Engram’s bet is that both properties come from putting the vault on your machine, making the format readable, and making the deletion path part of the format itself rather than a feature of the UI.