Skip to content

Supermemory for agent memory: how retrieval actually works

Preetham Reddy7 min read

An agent that forgets is a stateless text generator with good manners. Supermemory is a hosted memory API that fixes that by keeping a per-user memory graph of profiles and fact hierarchies, then returning the relevant parts for the next call. It bills by what it ingests, searches with a quoted 187ms median server time, and offers a free tier with 5 USD of credits a month, according to Supermemory's pricing page.

Most teams build memory twice. The first version is a vector store plus a cron job, the second is the one that works, and the gap between them is usually a year of learning which parts of memory aren't retrieval problems at all. Supermemory's bet is that the second version should be a service.

Four kinds of memory, and only one of them is RAG

Before picking a vendor, it helps to be precise about what you are storing. The field has settled on borrowed cognitive-science terms, and they map cleanly onto storage decisions.

Semantic memory is general knowledge with the clock stripped off, like the user prefers dark mode or this service talks to Postgres. Mem0's comparison describes its characteristic failure as the missing clock, a fact that was true once and is quietly wrong now.

Episodic memory is a dated event: on Tuesday, mocking the network call failed to fix the test. It is retrieved by similarity and recency together, and its usual failure is keeping the action and dropping the outcome, so the agent cheerfully retries the fix that already failed.

Procedural memory is a routine, a five-step deploy or a checklist for adding a REST endpoint. The graph-based agent memory taxonomy on arXiv treats procedural memory as skills, routines and immutable rules, distinct from the stable ontology of semantic facts.

Working memory is the context window itself. It isn't a store, it's a budget.

Supermemory doesn't expose these as four different endpoints. It exposes one ingest path, and the distinction shows up in how you scope writes and which retrieval mode you ask for. That is a reasonable design, as long as you know which kind of memory you think you are writing.

One ingest path, two pipelines

Everything goes in through a document add, scoped to a namespace that sits in the URL path and never in the body. The ingestion docs show the whole surface:

// illustration of the documented call shape
await supermemory.add("user_123", {
 content: "user: What's the weather?\nassistant: It's sunny today.",
 id: "conv_123", // stable id, enables diff billing
 taskType: "memory", // or "superrag"
 dreaming: "instant", // or "dynamic" (default)
});

The taskType parameter is the one that changes architecture.

So reference material gets superrag. Product manuals, policy PDFs and the API docs your support agent quotes from should be searchable, but they shouldn't shape what the system believes about a user. Conversations and anything the user says about themselves get memory. Teams that send everything through the memory pipeline end up with a profile polluted by the contents of a PDF, which is both expensive and wrong.

The dreaming parameter is the other one worth understanding early, because it will confuse you in development. The default "dynamic" groups related documents so memories form from coherent units, and the docs warn plainly that a fresh namespace can show zero memories and an empty profile for several minutes. Set dreaming: "instant" in tests and quickstarts, where you add a document and immediately read it back.

Retrieval: profile, query, or both

Retrieval is where the product earns its keep. The Vercel AI SDK integration wraps a model so memory injection happens without touching your call sites:

// illustration from the documented SDK surface
const modelWithMemory = withSupermemory(openai("gpt-5"), {
 namespace: "user-123",
 id: "conversation-456",
 mode: "full",
});

query searches memories against the current message, which is what you want for a support or research agent where the question determines what is relevant. full does both and costs both.

Start with query. Profile mode injects context on every turn whether or not the turn needs it, and an agent that gets told the user's dietary preferences before answering a billing question is paying tokens for noise. Move to full when you can show that profile context changes answers.

Two defaults in that middleware deserve a review before production. Memory saving is on by default with addMemory: "always", so every conversation persists unless you set "never". And skipMemoryOnError defaults to true, meaning if Supermemory is unreachable or retrieval hits its internal time limit, the call still runs with the original prompt and no memories. That is the correct default for a chat product and the wrong one for an agent whose behaviour depends on knowing what it already did. Set skipMemoryOnError: false where a silent amnesia event is worse than an error.

There is one more default to think about. Tool calls are dropped from saved conversations, since tool payloads are large and low-signal. For a multi-step agent, that is exactly the episodic record you wanted. includeToolCalls: true persists the full round trip.

Conversations and documents

Namespace per user or tenant

Memory pipeline extracts facts

SuperRAG chunks and indexes only

User profile

Semantic search and graph

Context injected into the model call

One namespace, two ingest pipelines, three retrieval modes.

Pay for new tokens, not for history

The billing model is the most interesting engineering decision in the product, because it dictates how you structure writes.

Supermemory meters in SM tokens, and re-ingesting content under the same id bills only the net-new delta. The billing docs give the arithmetic directly: billableTokens = max(0, fullTokenCount - previousTokenCount). Unchanged content isn't billed again on the token meter.

That is why the documented chat pattern keeps a whole session as one document with a stable id. Re-send the conversation, pay for the new turns. The requirements are narrow and worth memorising. The id must be stable and at most 255 characters, the org must be the same, and it must be an update instead of a full replace, because isFullReplace treats previous tokens as zero and you pay for the history again.

Plans run Free at 5 USD of included credits, Pro at 19 USD a month with 20 included, Max at 100 with 130, and Scale at 399 with 600. Subscription credits reset monthly and don't roll over; purchased top-ups persist until used.

Read the operations meter carefully if you are tempted by dreaming: "instant" everywhere. At 0.0001 USD each it's cheap per call and not cheap at a million documents a month, and it buys you latency you probably only need in tests.

For anyone putting patient or financial conversation history into a third-party memory service, the relevant line is that the pricing page lists SOC 2, a HIPAA BAA and a self-hosted option on the Scale plan.

Where this fits, and where it strains

Supermemory is a good fit when the thing you need is continuity per user across sessions, and when the content is conversational.

It is a worse fit when your retrieval target is governed enterprise data. A memory service that holds extracted facts about a user isn't the place for the fact table behind your revenue reporting; that belongs where the lineage, the row filters and the column masks already live. A memory layer and a governed lakehouse solve different halves of the problem, and the agent calls both.

One scheduling note. The v3 and v4 APIs are deprecated and will be shut down on 31 December 2026, per the API reference. v5 makes namespace scope explicit in the URL. Anything you build now should be v5.

Run one namespace through both retrieval modes before you commit

Compare the injected context, not just the answers. That tells you within an afternoon whether your problem is recall, personalization or neither, and the free tier's 5 USD of monthly credits covers the experiment.

Frequently asked questions

What is Supermemory?

Supermemory is a hosted memory API that gives an agent long-term memory of every user, storing a per-user memory graph of profiles and extracted facts that is searchable through one call.

How is Supermemory different from a vector database?

Supermemory runs the whole retrieval pipeline so there are no embeddings or vectors to manage, and it adds a memory layer on top, with fact extraction, user profiles and graph traversal alongside semantic search. A vector database stores and searches embeddings you generate and chunk yourself.

How much does Supermemory cost?

Plans are Free at 0 USD with 5 USD of included credits a month, Pro at 19 USD a month with 20 USD included, Max at 100 USD with 130 USD included, and Scale at 399 USD with 600 USD included, with Enterprise priced on committed spend. Usage draws from that credit balance at the same rates on every plan, with memory ingest at 5 USD per million SM tokens for plain text and search at 5 USD per million queries.

What is the difference between taskType memory and superrag?

The memory task type chunks and embeds content for search, extracts facts, updates the user profile and links the content into the memory graph, while superrag only chunks, embeds and indexes for search at 5x cheaper per token. Use superrag for reference material that should be searchable without shaping what the system knows about a user.