MCP server for persistent, searchable memory for AI agents.
Persistent memory for AI agents. Claude Code and Codex share one memory per project, and your own agents get the same pipeline: store, retrieve, summarize, and prune.
Website and install guides: memstack.stalewell.com
Tell one agent how your project works. The other already knows. Claude Code learns a convention; a new Codex session recalls it and gets the task right:
npm install -g @memstack/cli @memstack/mcp better-sqlite3@^11.10.0 # MemStack never installs storage drivers
memstack init # LLM provider (optional) + store
memstack connect claude-code
memstack connect codex
Every session then starts with the project's key memories, and both agents save what you ask them to remember. How harness memory works.
Building your own agent?
# Use MemStack in your application
npm install @memstack/core
# Give your coding agents the MemStack skill (-g: every project; omit it for this project only)
npx skills add isiomaC/memstack -g
@memstack/core is the runtime SDK; the Agent Skill teaches compatible coding agents how to integrate and operate MemStack correctly. Without -g, the skill installs into the current project (.agents/skills/, plus a .claude/skills/ link for Claude Code). Update it later with npx skills update.
The problem: AI agents forget. Every interaction starts from zero. You either stuff everything into the context window (expensive, slow, degrades output quality) or the agent has no memory of past conversations.
What MemStack does: A persistent memory pipeline that lives between your agent and the LLM. It stores every interaction, retrieves only what's relevant, summarizes old memories to save tokens, and prunes stale ones automatically. One method call, no infrastructure required.
Think of it as the open-source alternative to Mem0 — pluggable storage, bring your own LLM, zero vendor lock-in.
LLMs have context windows, not memory. The difference matters.
| Approach | Problem |
|---|---|
| Stuff everything in context | Cost is O(n²). 100 conversations = thousands of tokens = dollars per call. Quality degrades from "lost in the middle" effect. |
| Use a vector DB directly | You get similarity search. You don't get summarization, pruning, recency weighting, deduplication, or token budget management. You're building the pipeline yourself. |
| Use Mem0 | Proprietary, cloud-only with their hosted API. You don't control where your data lives. |
| Use MemStack | Full pipeline. Pluggable everything. Your data, your infrastructure. Open source. |
What MemStack handles that raw vector DBs don't:
compileContext() tells you how many tokens you're spending before the LLM callPersistent memory across agent harnesses: what you tell Claude Code, Codex recalls in the same project, and the reverse.
npm install -g @memstack/cli @memstack/mcp better-sqlite3@^11.10.0 # MemStack never installs storage drivers for you
memstack init # choose an LLM provider and a store
memstack connect claude-code
memstack connect codex
Then, in Claude Code: "Remember that this project uses Hono." In Codex, in the same repository: "What framework does this project use?" Codex answers Hono. See Harness Memory for how it works.
npm install @memstack/core
import { MemStack, OpenAILLMAdapter, OpenAIEmbeddingAdapter, InMemoryStorageAdapter } from "@memstack/core";
const llm = new OpenAILLMAdapter({ apiKey: process.env.OPENAI_API_KEY! });
const memstack = new MemStack({
llm,
embedding: new OpenAIEmbeddingAdapter({ apiKey: process.env.OPENAI_API_KEY! }),
storage: new InMemoryStorageAdapter(),
});
DeepSeek provides chat completions but has no embedding API. Use the OpenAI-compatible LLM adapter with baseURL and omit the embedding adapter — retrieval falls back to keyword + recency + importance ranking. You still get the full pipeline: store, summarize, prune, and compileContext.
import { MemStack, OpenAILLMAdapter, InMemoryStorageAdapter } from "@memstack/core";
const llm = new OpenAILLMAdapter({
apiKey: process.env.DEEPSEEK_API_KEY!,
baseURL: "https://api.deepseek.com/v1",
defaultModel: "deepseek-flash",
});
const memstack = new MemStack({
llm,
storage: new InMemoryStorageAdapter(),
// No embedding adapter — retrieval uses keyword matching
});
Same pattern — change baseURL and defaultModel:
// OpenRouter
const llm = new OpenAILLMAdapter({
apiKey: process.env.OPENROUTER_API_KEY!,
baseURL: "https://openrouter.ai/api/v1",
defaultModel: "openai/gpt-4o-mini",
});
// Together AI
const llm = new OpenAILLMAdapter({
apiKey: process.env.TOGETHER_API_KEY!,
baseURL: "https://api.together.xyz/v1",
defaultModel: "meta-llama/Llama-3.3-70B-Instruct-Turbo",
});
// Gemini (OpenAI-compatible endpoint)
const llm = new OpenAILLMAdapter({
apiKey: process.env.GEMINI_API_KEY!,
baseURL: "https://generativelanguage.googleapis.com/v1beta/openai",
defaultModel: "gemini-2.0-flash",
});
// 1. Store what happened
await memstack.memory.store({
actorId: "support-bot-42",
content: "User reports login failing with error 503 on Chrome 125.",
tags: ["login", "bug", "chrome"],
importance: 0.8,
});
// 2. Later, retrieve relevant context
const memories = await memstack.memory.retrieve({
actorId: "support-bot-42",
query: "login error",
strategy: "hybrid",
});
// 3. Assemble an LLM-ready context
const ctx = await memstack.memory.compileContext({
actorId: "support-bot-42",
maxTokens: 2000,
});
const response = await llm.complete({
system: `You are a support bot. Here is what you remember:\n${ctx.systemPrompt}`,
user: "The user is back and still can't log in. What do you do?",
});
console.log(response.text);
// "Based on our history, the user has been experiencing 503 errors on Chrome 125..."
// 4. Every 100 interactions, summarization triggers automatically.
// Old interactions are compressed into a paragraph. Token costs stay flat.
Coding agents forget everything between sessions, and they don't share what they learn with each other. MemStack gives Claude Code and Codex one memory per project: a decision made with one agent is known to the other, and to every later session.
memory_store,
memory_retrieve, memory_get, memory_delete, memory_stats) through
the MCP harness profile. When you say "remember that…", or the agent learns
a durable fact, decision, preference, or rule, it calls memory_store.
MemStack asks your LLM for a few topic tags, so "uses Hono" can later be
found by "which framework?". If tagging fails, the memory is still saved.memory_retrieve with plain questions.
Recall runs locally with keyword ranking (BM25 with stemming) and never
calls the LLM, so it is fast and works offline.scope: "global") and are recalled in
every project. One project can never read or delete another's memories.~/.memstack/config.json (readable only by you), never in agent configs.
Agents are told never to store secrets, and MemStack checks too: a memory
that contains a credential (a private key, a cloud or source-host token, a
JWT, a password in a connection string, a password = "..." assignment) is
refused before the LLM or the store sees it, and the error names the kind of
secret without repeating it. The check is pattern-based, so a secret in a
format it doesn't know will pass; MEMSTACK_SECRET_POLICY=redact stores the
memory with [REDACTED:<kind>] in place of the secret, and off disables
the check.| Command | What it does |
|---|---|
memstack init | Choose an LLM provider (optional) and a store. Verifies the key with a real request and writes ~/.memstack/config.json. Without a key, memories are saved without topic tags and recall works the same. Non-interactive: --provider, --base-url, --model, --api-key-env <VAR>, --no-llm, --store, --path, --url, --yes. |
memstack connect <claude-code|codex> | Registers MemStack with the agent. Checks the server works first, and undoes everything if a step fails. --dry-run shows the changes; --no-hooks and --no-agents-md skip those parts. |
memstack disconnect <claude-code|codex> | Removes everything connect added. Your memories are kept. |
memstack status | Config, storage, the current project, and what each agent has connected. |
memstack doctor | Diagnoses setup problems and prints the fix for each. --live also tests the LLM key. |
memstack memories [query] | Lists or searches this project's memories. --global for global ones, --delete <id> to remove one. |
memstack project | Shows this repository's project ID. project pin <id> fixes it in .memstack.json; project merge <old-id> moves memories from an old ID. |
connect changesconnect changes only these entries, backs up each file first, and
disconnect restores each file exactly:
| Agent | File | Change |
|---|---|---|
| Claude Code | ~/.claude.json | A user-scope memstack MCP server, added with claude mcp add-json. |
| Claude Code | ~/.claude/settings.json | A SessionStart hook running memstack-mcp hook session-start. |
| Codex | ~/.codex/config.toml | A memstack MCP server, added with codex mcp add. |
| Codex | ~/.codex/hooks.json | A SessionStart hook. Codex runs it after you approve it once with /hooks. |
| Codex | ~/.codex/AGENTS.md | A short block between memstack:begin/memstack:end markers telling Codex to save memories with memory_store. |
The registered command is the memstack-mcp you installed, run by absolute
path, so it works even when the agent starts without your shell's PATH.
Restart Claude Code, or start a new Codex session, to load it.
The project ID comes from the repository itself, with nothing stored, so it stays the same across clones, worktrees, renamed remotes, moved folders, new machines, and storage switches:
.memstack.json at the repository root
(memstack project pin <id>; commit it to share it with your team);origin remote;A fork shares its first commit with the original repository, so on one store
the two share memories until one of them runs memstack project pin.
Any supported store works; pick it in memstack init. MemStack never installs
storage drivers: install the one you need alongside @memstack/mcp.
| Store | Driver to install | Notes |
|---|---|---|
sqlite | better-sqlite3@^11.10.0 | Default; one local file, safe for both agents at once. |
postgres | postgres@^3.4.9 | Shared across machines. |
redis | ioredis@^5.11.1 | |
disk, markdown | none | Single process only; doctor warns when both agents use them. |
Switching stores keeps project IDs; move memories with memstack export and
memstack import.
Run memstack doctor first; it checks the config file and its permissions,
the LLM key, the storage driver, each agent's registration, hook, and
guidance, and starts the server to confirm it answers.
npm install -g … command it prints./hooks.memstack status shows the
AGENTS.md guidance as present; reconnect if not.memstack project shows which project you are
in. Use memstack project merge <old-id> to bring memories from an old ID.Details on the MCP server itself: docs/MCP_SETUP.md.
MemStack's core is a five-stage pipeline. Each stage can be used independently.
Every agent interaction becomes a Memory with metadata that controls how it's retrieved, summarized, and pruned later.
interface Memory {
id: string;
actorId: string; // Who this memory belongs to (user ID, agent ID, session ID)
memoryType: MemoryType; // "interaction" | "summary" | "observation" | "fact" | "reflection" | "preference" | "decision" | "instruction"
content: string; // The actual text
importance: number; // 0-1 — higher = survives pruning, ranks higher in retrieval
emotionalValence: number; // -1 to 1 — for tone-aware retrieval
tags: string[]; // Filter by tag: "bug", "billing", "urgent", etc.
embedding?: number[]; // Computed automatically if embedding adapter is configured
metadata?: Record<string, unknown>; // Your custom fields
expiresAt?: Date; // Auto-pruned after this date
sourceId?: string; // Link back to the originating event
createdAt: Date;
}
// Simple store
await ms.memory.store({
actorId: "agent-7",
content: "Customer asked about refund policy for Q2 purchases.",
tags: ["billing", "refund"],
});
// Batch store — embeddings are batched into one API call for efficiency
await ms.memory.storeBatch([
{ actorId: "agent-7", content: "First interaction" },
{ actorId: "agent-7", content: "Second interaction" },
{ actorId: "agent-7", content: "Third interaction" },
]);
Pull back what's relevant — by keyword, by meaning (semantic), by recency, or by importance.
const memories = await ms.memory.retrieve({
actorId: "agent-7", // Scope to one actor
query: "refund policy", // What to search for
strategy: "hybrid", // How to rank: "recent" | "important" | "semantic" | "hybrid"
limit: 10, // Max results
memoryTypes: ["interaction"], // Only certain types
tags: ["billing"], // Only certain tags
});
Strategy behavior:
| Strategy | Sorts by | Requires embeddings | Best for |
|---|---|---|---|
recent | Newest first | No | Knowing what just happened |
important | Highest importance first | No | Filtering noise, keeping signal |
semantic | Cosine similarity to query | Yes | "Find memories about X" |
hybrid | Semantic + importance blend | Yes | Best of both worlds |
No embedding adapter? semantic and hybrid fall back to keyword matching + importance sort. No API costs, just less precise.
compileContext() takes retrieval results and assembles an LLM-ready system prompt — deduplicated, sorted by recency and importance, with a token estimate so you know the cost before calling the LLM.
const ctx = await ms.memory.compileContext({
actorId: "agent-7",
maxTokens: 2000, // Budget — assembler stops when it hits this
memoryTypes: ["interaction", "summary"],
});
// ctx.systemPrompt:
// ## Important Memories
// - The customer has been attempting login for 3 days. (importance: 0.85)
// - Refund was processed for order #4521 on Jan 12. (importance: 0.72)
//
// ## Recent Interactions
// - Customer asked about refund policy for Q2 purchases.
// - Customer reported login error 503 on Chrome 125.
console.log(ctx.tokenEstimate); // ~280
// Inject into your LLM call
const currentMessage = "The user is asking about their refund status.";
const response = await llm.complete({
system: ctx.systemPrompt,
user: currentMessage,
});
console.log(response.text);
compileContext() handles deduplication, token budgeting, and splits context into important-vs-recent sections. Without it, you'd be concatenating raw retrieval results and risking context-window overflow.
When an actor has hundreds of interactions, retrieval gets expensive and context gets bloated. Summarization compresses old interactions into a single paragraph using the configured LLM.
const { summary, deletedCount } = await ms.memory.summarize({
actorId: "agent-7",
olderThan: new Date(Date.now() - 7 * 86400000), // Older than 7 days
skipMostRecent: 10, // Never touch the 10 most recent
targetCount: 50, // Summarize at most 50 memories
memoryTypes: ["interaction"],
keepOriginals: false, // Delete originals after summary
});
// summary.content:
// "Over the past week, the customer reported recurring login failures (error 503)
// on Chrome 125. Multiple troubleshooting attempts including cache clearing and
// password reset were unsuccessful. A refund was processed for order #4521."
console.log(deletedCount); // 47 — 47 interactions compressed into 1 summary memory
Auto-summarization: Set summarizationThreshold in config (default: 100). Every 100th interaction for an actor triggers summarization automatically.
Warning: keepOriginals: false deletes the summarized memories. Set keepOriginals: true to preserve them alongside the summary.
Custom summarization prompt:
const ms = new MemStack({
llm,
defaults: {
summarizationPrompt:
"You are an enterprise support memory compressor. Highlight: customer name,
product, severity, resolution status, and any open issues.",
},
});
Not all memories deserve to live forever. Pruning removes low-value memories to keep storage and retrieval fast.
// Remove memories older than 30 days
await ms.memory.prune({ type: "byAge", maxAge: 30 * 86400000 });
// Keep only memories above importance 0.3
await ms.memory.prune({ type: "byImportance", minImportance: 0.3 });
// Keep at most 500 memories per actor
await ms.memory.prune({ type: "byCount", maxPerActor: 500 });
// Remove specific types
await ms.memory.prune({ type: "byType", memoryTypes: ["observation"] });
// Custom logic
await ms.memory.prune({
type: "custom",
shouldRemove: (memory) => memory.content.includes("[RESOLVED]"),
});
// Dry run first — see what would be removed
const { wouldPrune, count } = await ms.memory.dryRunPrune({
type: "byAge",
maxAge: 86400000,
});
console.log(`Would remove ${count} memories:`, wouldPrune);
Auto-prune on every process() call by setting pruneStrategy in config:
const ms = new MemStack({
llm,
defaults: {
pruneStrategy: { type: "byImportance", minImportance: 0.05 },
},
});
// detectUrgency and classifyIntent are your own business logic.
// They could be simple keyword matchers, regex, or an LLM call.
function detectUrgency(msg: string): number {
if (msg.match(/urgent|asap|immediately/i)) return 0.9;
if (msg.match(/error|fail|broken/i)) return 0.7;
return 0.5;
}
function classifyIntent(msg: string): string[] {
const tags: string[] = [];
if (msg.match(/bill|refund|charge|payment/i)) tags.push("billing");
if (msg.match(/error|bug|fail|crash/i)) tags.push("bug");
if (msg.match(/login|password|account/i)) tags.push("account");
return tags;
}
// Every customer message becomes a memory
async function handleMessage(customerId: string, message: string) {
await ms.memory.store({
actorId: `customer:${customerId}`,
content: message,
importance: detectUrgency(message),
tags: classifyIntent(message),
});
// Retrieve everything relevant to this customer's history
const ctx = await ms.memory.compileContext({
actorId: `customer:${customerId}`,
maxTokens: 1500,
});
const response = await llm.complete({
system: `You are a support agent. Customer history:\n${ctx.systemPrompt}`,
user: message,
});
return response.text;
}
// Every 100th interaction, old history auto-compresses.
// A customer with 10,000 messages still fits in a $0.02 LLM call.
// Suppose you have documents from your knowledge base
const documents = [
{ text: "Authentication uses JWT tokens with 15-minute expiry.", url: "/docs/auth", section: "security" },
{ text: "Refunds are processed within 5-10 business days.", url: "/docs/billing", section: "billing" },
];
// Index documents as observation memories
for (const doc of documents) {
await ms.memory.store({
actorId: "knowledge-base",
content: doc.text,
memoryType: "observation",
metadata: { source: doc.url, section: doc.section },
});
}
// Query with semantic search
const relevantDocs = await ms.memory.retrieve({
actorId: "knowledge-base",
query: "How does authentication work?",
strategy: "semantic",
limit: 5,
});
const ctx = await ms.memory.compileContext({
actorId: "knowledge-base",
memoryTypes: ["observation"],
});
// Prompt the LLM with retrieved context
const answer = await llm.complete({
system: `Answer using only these documents:\n${ctx.systemPrompt}`,
user: "How does authentication work?",
});
// Each user gets their own memory space
async function chat(userId: string, message: string) {
await ms.memory.store({
actorId: userId,
content: message,
});
const ctx = await ms.memory.compileContext({
actorId: userId,
maxTokens: 1000,
});
return llm.complete({
system: `You are a friendly assistant. Conversation history with this user:\n${ctx.systemPrompt}`,
user: message,
});
}
// Get stats
const total = await ms.memory.count();
const userCount = await ms.memory.count({ actorId: "user-42" });
| Type | Purpose | Example |
|---|---|---|
interaction | Default. Direct exchanges between agent and user/other agent. | "User asked about billing." |
summary | Compressed collection of old interactions. Created by summarize(). | "Over 3 weeks, user reported 5 login failures..." |
observation | Passive knowledge — facts, documents, things the agent knows but didn't interact with. | "Company refund policy is 30 days from purchase." |
fact | Verified knowledge — discrete truths the agent has confirmed. | "The user's subscription tier is Enterprise." |
reflection | Self-generated insight — the agent thinking about its own experiences. | "I tend to over-explain billing policies — should be more concise." |
preference | How the user likes things done. | "Prefer small pull requests with one concern each." |
decision | A choice made, ideally with its reason. | "Chose Hono over Express for edge runtime support." |
instruction | A standing rule to follow. | "Never commit directly to main." |
Types control retrieval behavior — compileContext() treats interaction and summary differently from observation. Use types to separate "what happened" from "what I know."
Four strategies, each with a purpose:
// "What just happened?" — most recent first
await ms.memory.retrieve({ actorId: "x", strategy: "recent", limit: 3 });
// "What matters most?" — highest importance, ignoring age
await ms.memory.retrieve({ actorId: "x", strategy: "important" });
// "What relates to this query?" — cosine similarity search (needs embeddings)
await ms.memory.retrieve({ actorId: "x", query: "login bug", strategy: "semantic" });
// "Balance relevance and importance" — semantic + importance blend
await ms.memory.retrieve({ actorId: "x", query: "login bug", strategy: "hybrid" });
Choosing a strategy:
recent for chatbots, ongoing conversations, anything time-sensitiveimportant for long-running agents where signal-to-noise matterssemantic for RAG, document search, knowledge base querieshybrid for most agent memory — it balances meaning with significanceLexicalRetriever answers natural questions without embeddings or an LLM
call, and behaves the same on every storage adapter. It loads the memories
in the given actors through retrieve(), then ranks them in MemStack with
BM25 over content and tags, with stemming and prefix matching. When nothing
matches, it returns the most important memories instead.
import { LexicalRetriever } from "@memstack/core";
const retriever = new LexicalRetriever(storage);
const { hits, fallback } = await retriever.recall({
actorIds: ["project:abc", "global"], // Searched together
query: "What framework does this project use?",
limit: 10, // Max results
maxChars: 8000, // Max total content; the top hit is always returned
});
Up to 2,000 memories per actor are ranked (candidateLimit). Only returned
memories are marked as accessed.
Embeddings power semantic search. They're optional — without them, retrieval uses keyword matching.
With embeddings (embedding adapter configured):
import { MemStack, OpenAILLMAdapter, OpenAIEmbeddingAdapter, InMemoryStorageAdapter } from "@memstack/core";
const ms = new MemStack({
llm: new OpenAILLMAdapter({ apiKey: process.env.OPENAI_API_KEY! }),
embedding: new OpenAIEmbeddingAdapter({ apiKey: process.env.OPENAI_API_KEY! }),
storage: new InMemoryStorageAdapter(),
});
// store() computes a 1536-dim vector automatically
await ms.memory.store({
actorId: "agent-7",
content: "Customer asked about refund policy for Q2 purchases.",
});
// retrieve() with "semantic" or "hybrid" uses cosine similarity
// Query: "refund" finds the refund policy memory even though the word "refund"
// appears differently across stored memories.
const results = await ms.memory.retrieve({
actorId: "agent-7",
query: "how do I get my money back",
strategy: "semantic",
});
// Matches "Customer asked about refund policy" — semantic match, not keyword match.
Without embeddings (no embedding adapter):
const ms = new MemStack({
llm: new OpenAILLMAdapter({ apiKey: process.env.OPENAI_API_KEY! }),
storage: new InMemoryStorageAdapter(),
// no embedding adapter
});
// store() works identically, just no vector computed
await ms.memory.store({
actorId: "agent-7",
content: "Customer asked about refund policy for Q2 purchases.",
});
// retrieve() with "semantic" or "hybrid" falls back to keyword matching
// plus importance/recency sorting. No API costs, no setup required.
const results = await ms.memory.retrieve({
actorId: "agent-7",
query: "refund",
strategy: "hybrid", // falls back to keyword + importance
});
// Still works — finds "refund" via substring match. Less precise for
// paraphrased queries ("money back" won't match "refund").
Batch embedding: storeBatch() sends all texts in one embedding API call, reducing cost and latency.
// Disable auto-embedding if you only need keyword search
const ms = new MemStack({
llm,
embedding: new OpenAIEmbeddingAdapter({ apiKey }),
defaults: { embedOnStore: false },
});
Different embedding models produce vectors of different lengths. Cosine similarity only works between vectors of the same dimension. If you change embedding models, existing vectors become incompatible — they can't be compared to new ones.
| Adapter | Default model | Dimensions |
|---|---|---|
OpenAIEmbeddingAdapter | text-embedding-3-small | 1536 |
OpenAIEmbeddingAdapter | text-embedding-3-large | 3072 |
CohereEmbeddingAdapter | embed-english-v3.0 | 1024 |
CohereEmbeddingAdapter | embed-english-light-v3.0 | 384 |
CohereEmbeddingAdapter | embed-english-v2.0 | 4096 |
CohereEmbeddingAdapter | embed-multilingual-v3.0 | 1024 |
What happens if dimensions don't match: If you store memories with one model (e.g., 1536 dims) then switch to another model (e.g., 1024 dims), the storage adapter receives query vectors and stored vectors of different lengths. Cosine similarity between vectors of different dimensions is undefined — results depend on the storage backend's behavior. Most will either error, return empty results, or produce meaningless scores.
Recommendation: Pick one embedding model per storage instance and stick with it. If you need to switch models, create a new storage instance and re-embed from scratch.
DeepSeek users: DeepSeek has no embeddings API. If you use DeepSeek as your LLM, you must either:
"recent" or "important" retrieval strategies (no API costs, less precise)MemStack is provider-agnostic. Every boundary is an interface — bring your own LLM, embedding model, and storage backend.
Used by summarize() and compileContext(). Ships with OpenAI, Anthropic, Ollama, and Groq built-in — and via baseURL, the OpenAI adapter works with any OpenAI-compatible API (DeepSeek, Mistral, Gemini, Together AI, Perplexity, Fireworks, xAI, and dozens more).
// OpenAI
import { OpenAILLMAdapter } from "@memstack/core";
const llm = new OpenAILLMAdapter({ apiKey: "..." });
// Any OpenAI-compatible API — just change baseURL
const deepseek = new OpenAILLMAdapter({ apiKey: "...", baseURL: "https://api.deepseek.com/v1" });
const mistral = new OpenAILLMAdapter({ apiKey: "...", baseURL: "https://api.mistral.ai/v1" });
const together = new OpenAILLMAdapter({ apiKey: "...", baseURL: "https://api.together.xyz/v1" });
// Anthropic
import { AnthropicLLMAdapter } from "@memstack/core";
const llm = new AnthropicLLMAdapter({
apiKey: process.env.ANTHROPIC_API_KEY!,
defaultModel: "claude-sonnet-4-5-20250929",
});
// Ollama (built-in)
import { OllamaLLMAdapter } from "@memstack/core";
const llm = new OllamaLLMAdapter({
baseURL: "http://localhost:11434",
defaultModel: "llama3.2",
});
Used by semantic retrieval. Ships with OpenAI and Cohere built-in — and via baseURL, the OpenAI adapter works with any OpenAI-compatible embedding API (Together AI, Voyage AI, Jina, Nomic, and more).
import { OpenAIEmbeddingAdapter, CohereEmbeddingAdapter } from "@memstack/core";
// OpenAI
new OpenAIEmbeddingAdapter({ apiKey: "...", model: "text-embedding-3-small" }); // 1536 dims
// Cohere
new CohereEmbeddingAdapter({ apiKey: "..." }); // embed-english-v3.0, 1024 dims
// Any OpenAI-compatible embedding API
new OpenAIEmbeddingAdapter({ apiKey: "...", baseURL: "https://api.voyageai.com/v1", model: "voyage-3" });
MemStack contains 18 storage-adapter implementations. Twelve are exported from @memstack/core; six remain experimental source implementations. Core has no runtime dependencies, and database clients are injected by callers.
Support levels:
pnpm test:e2e.Built-in (zero external deps):
| Adapter | Backend | Use case |
|---|---|---|
InMemoryStorageAdapter | In-memory Map | Testing, prototyping |
DiskStorageAdapter | Local JSON files | Simple local persistence |
MarkdownStorageAdapter | Append-only .md files | Human-readable, git-diffable, debug-friendly |
HybridStorageAdapter | Compose any two StorageProviders | Cache + durable, edge + durable |
Relational / SQL:
| Adapter | Backend | Vector search |
|---|---|---|
PostgresStorageAdapter | PostgreSQL + pgvector | HNSW native |
SQLiteStorageAdapter | SQLite (better-sqlite3) | Cosine in-memory |
Vector databases:
| Adapter | Backend |
|---|---|
QdrantStorageAdapter | Qdrant |
WeaviateStorageAdapter | Weaviate |
LanceDBStorageAdapter | LanceDB |
MongoDBStorageAdapter | MongoDB Atlas Vector Search |
Cache / KV:
| Adapter | Backend |
|---|---|
RedisStorageAdapter | Redis (ioredis) |
Graph:
| Adapter | Backend |
|---|---|
Neo4jStorageAdapter | Neo4j |
These implementations are available to source contributors but are not part of the published package API.
| Adapter | Backend | Blocker |
|---|---|---|
TursoStorageAdapter | Turso (libsql) | Cloud-only (needs Turso account) |
ChromaStorageAdapter | ChromaDB | Embedding function dependency |
PineconeStorageAdapter | Pinecone | Cloud-only (needs API key) |
UpstashStorageAdapter | Upstash Redis + Vector | Cloud-only (needs API key) |
Mem0StorageAdapter | Mem0 OSS or Cloud | Cloud-only (needs API key) |
ZepStorageAdapter | Zep Cloud or CE | Cloud-only (needs API key) |
Live cloud compatibility remains unverified for Pinecone, Upstash, Mem0, Zep, and Turso. Chroma's real-client E2E suite is skipped when its optional default embedding function is unavailable. LLM and embedding-provider tests use mocks; live-provider testing is opt-in and is not part of CI.
Quick-start per backend:
// Postgres
import { PostgresStorageAdapter } from "@memstack/core";
const storage = new PostgresStorageAdapter({ connectionString: "postgres://..." });
// Redis
import Redis from "ioredis";
import { RedisStorageAdapter } from "@memstack/core";
const storage = new RedisStorageAdapter({ redis: new Redis() });
// Markdown (append-only, human-readable)
import { MarkdownStorageAdapter } from "@memstack/core";
const storage = new MarkdownStorageAdapter({ dir: "./memories" });
// Hybrid (Redis cache + Postgres durable)
import { HybridStorageAdapter } from "@memstack/core";
const storage = new HybridStorageAdapter({
cache: new RedisStorageAdapter({ redis: new Redis() }),
durable: new PostgresStorageAdapter({ connectionString: "postgres://..." }),
});
Custom storage:
import type { StorageProvider, MemoryStoreInput } from "@memstack/core";
class MyStorage implements StorageProvider {
async store(input: MemoryStoreInput): Promise<Memory> { /* ... */ }
async get(id: string): Promise<Memory | null> { /* ... */ }
async retrieve(query: MemoryRetrieveQuery, embedding?: number[]): Promise<Memory[]> { /* ... */ }
async count(filter?: MemoryCountFilter): Promise<number> { /* ... */ }
async delete(id: string): Promise<void> { /* ... */ }
async deleteMany(ids: string[]): Promise<number> { /* ... */ }
async storeBatch(inputs: MemoryStoreInput[]): Promise<Memory[]> { /* ... */ }
async initialize(): Promise<void> { /* ... */ }
async close(): Promise<void> { /* ... */ }
}
| Backend | Vector search | Touch | Status |
|---|---|---|---|
| InMemory | Cosine in-memory | Yes | ✅ Production |
| Disk (JSON) | Keyword + importance | Yes | ✅ Production |
| Markdown | Keyword + importance | No | ✅ Production |
| Postgres | pgvector HNSW | Yes | ✅ Production |
| Redis | RediSearch KNN (auto-detect) | Yes | ✅ Production |
| Qdrant | ANN native | No | ✅ Production |
| Weaviate | BM25 + vector hybrid | No | ✅ Production |
| LanceDB | DiskANN native | No | ✅ Production |
| MongoDB | Atlas Vector Search | No | ✅ Production |
| Neo4j | Neo4j vector index | No | ✅ Production |
| Hybrid | Delegates to cache/durable | If durable supports | ✅ Production |
| SQLite | Cosine in-memory | Yes | ✅ Production |
import { MemStack } from "@memstack/core";
const ms = new MemStack({
llm: LLMProvider, // Required — for summarization
embedding?: EmbeddingProvider, // Optional — for semantic search
storage?: StorageProvider, // Optional — defaults to InMemoryStorageAdapter
defaults?: {
summarizationThreshold?: number, // Auto-summarize every N process() calls. Default: 100
embedOnStore?: boolean, // Auto-embed on store(). Default: true
pruneStrategy?: PruneStrategy, // Auto-prune during process() (throttled). Default: disabled
pruneInterval?: number, // Run auto-prune every N process() calls. Default: 100
autoImportance?: boolean, // LLM-score importance in process() when not provided. Default: false
autoTags?: boolean, // LLM-extract tags in process() when not provided. Default: false
summarizationPrompt?: string, // Custom prompt for the summarizer
},
hooks?: {
onMemoryStored?: (memory: Memory) => void;
onMemoryPruned?: (ids: string[]) => void;
onSummaryCreated?: (summary: Memory, deletedCount: number) => void;
onError?: (error: Error, context: string) => void;
},
});
Auto-behaviors run inside
process(), notstore().process()tracks a per-actor call count: summarization fires everysummarizationThresholdcalls, and pruning fires everypruneIntervalcalls (whenpruneStrategyis set).store()is the low-level write and never triggers these.
All methods accessible via ms.memory.*:
// Store
ms.memory.store(input: MemoryStoreInput): Promise<Memory>
ms.memory.storeBatch(inputs: MemoryStoreInput[]): Promise<Memory[]>
// Retrieve
ms.memory.retrieve(query: MemoryRetrieveQuery): Promise<Memory[]>
ms.memory.get(id: string): Promise<Memory | null>
// Context assembly
ms.memory.compileContext(options: ContextOptions): Promise<CompiledContext>
// Lifecycle
ms.memory.summarize(options: SummarizeOptions): Promise<{ summary: Memory; deletedCount: number }>
ms.memory.prune(strategy: PruneStrategy): Promise<{ pruned: string[]; count: number }>
ms.memory.dryRunPrune(strategy: PruneStrategy): Promise<{ wouldPrune: string[]; count: number }>
// Management
ms.memory.count(filter?: MemoryCountFilter): Promise<number>
ms.memory.delete(id: string): Promise<void>
ms.memory.deleteMany(ids: string[]): Promise<number>
ms.memory.touch(id: string): Promise<void>
ms.memory.purgeActor(actorId: string): Promise<number>
ms.memory.merge(ids: string[]): Promise<Memory>
ms.memory.stats(actorId?: string): Promise<MemoryStats>
ms.memory.summarizeStream(options: SummarizeOptions): AsyncIterable<{ chunk: string; text: string }>
Snapshot and restore full state for persistence, backups, or migration:
import * as fs from "node:fs";
// Save
const snapshot = await ms.export();
fs.writeFileSync("state.json", JSON.stringify(snapshot, null, 2));
// Restore
const data = JSON.parse(fs.readFileSync("state.json", "utf-8"));
await ms2.import(data);
Each memory's original createdAt is preserved on import, so export → import is a lossless round-trip — safe for backups and cross-backend migration (e.g. disk → Postgres). All storage adapters honor a createdAt supplied on store()/storeBatch(); when omitted, they default to the current time.
const status = await ms.health();
// { storage: true, llm: true, embedding: true }
await ms.close(); // graceful shutdown
The harness features are built on HarnessMemory, which works with any
storage adapter. Use it to build your own agent integration:
import { HarnessMemory, defaultRecallNamespaces, projectNamespace } from "@memstack/core";
const memory = new HarnessMemory({ storage, llm });
await memory.remember({
namespace: projectNamespace("acme-api"), // or "global"
content: "This project uses Hono",
kind: "decision", // fact | preference | decision | instruction | observation | ...
source: { harness: "my-agent", project: "acme-api" },
});
const { hits, fallback } = await memory.recall({
namespaces: defaultRecallNamespaces("acme-api"), // project, then global
query: "Which framework do we use?",
limit: 10,
maxChars: 8000,
});
await memory.get(id, namespaces); // null outside namespaces
await memory.forget(id, namespaces); // refuses ids outside namespaces
await memory.stats(namespaces); // { "project:acme-api": 3, global: 1 }
await memory.moveNamespace(from, to); // e.g. when a project's ID changes
remember asks the LLM for topic tags (disable with autoTags: false); every
other method makes no LLM call. recall uses LexicalRetriever.
const ms = new MemStack({
llm: new OpenAILLMAdapter({ apiKey: "..." }),
// Defaults control auto-behavior (all applied during process())
defaults: {
summarizationThreshold: 50, // Summarize every 50 process() calls (default: 100)
embedOnStore: false, // Don't auto-embed — saves API costs
pruneStrategy: { // Auto-clean during process(), throttled by pruneInterval
type: "byAge",
maxAge: 90 * 86400000, // 90 days
},
pruneInterval: 100, // Run the prune check every 100 process() calls (default: 100)
autoImportance: true, // Let the LLM score importance when you don't pass one
autoTags: true, // Let the LLM extract tags when you don't pass any
},
// Hooks for observability
hooks: {
onMemoryStored: (m) => logger.debug("memory:stored", { id: m.id, actor: m.actorId }),
onMemoryPruned: (ids) => logger.info("memory:pruned", { count: ids.length }),
onSummaryCreated: (summary, n) => logger.info("memory:summarized", { count: n }),
onError: (err, context) => logger.error("memory:error", { context, message: err.message }),
},
});
The CLI and the MCP harness profile read ~/.memstack/config.json (or
$MEMSTACK_HOME/config.json), which memstack init writes with permissions
readable only by you:
{
"version": 1,
"llm": { "provider": "openai-compatible", "apiKey": "…", "baseURL": "https://api.deepseek.com", "model": "deepseek-flash" },
"storage": { "type": "sqlite", "path": "/Users/you/.memstack/memstack.db" }
}
Environment variables override the file one section at a time: any LLM
variable (OPENAI_API_KEY, ANTHROPIC_API_KEY, MEMSTACK_OPENAI_BASE_URL,
MEMSTACK_LLM_MODEL) replaces the whole llm section, and MEMSTACK_STORAGE
replaces the whole storage section, so a key is never sent to another
provider's URL.
Implement StorageProvider for any database. The interface is 9 methods. See the reference section above for the full contract.
Optional members, none of them required:
capabilities: { multiProcess?, textSearch? } declares whether several
processes can share the store safely and whether search() is native.search(query) provides native full-text search. LexicalRetriever uses
it when textSearch is declared and ranks memories itself otherwise.retrieve() should honor touch: false by returning memories without
marking them as accessed.SQLiteStorageAdapter enables WAL and a 5-second busy timeout so several
processes can share one database file. Set walMode: false or
busyTimeoutMs to change this.
Implement LLMProvider or EmbeddingProvider for any service:
import type { LLMProvider } from "@memstack/core";
class TogetherAIAdapter implements LLMProvider {
async complete(req: { system: string; user: string; model?: string }) {
const res = await fetch("https://api.together.xyz/v1/chat/completions", {
headers: { Authorization: `Bearer ${this.apiKey}`, "Content-Type": "application/json" },
body: JSON.stringify({ model: req.model, messages: [{ role: "system", content: req.system }, { role: "user", content: req.user }] }),
});
const data = await res.json() as any;
return { text: data.choices[0].message.content, tokens: { prompt: data.usage.prompt_tokens, completion: data.usage.completion_tokens, total: data.usage.total_tokens } };
}
}
Monitor memory operations without modifying code:
const ms = new MemStack({
llm,
hooks: {
onMemoryStored: (m) => metrics.increment("memory.stored"),
onSummaryCreated: (_, n) => metrics.gauge("memory.summarized_count", n),
onMemoryPruned: (ids) => metrics.increment("memory.pruned", ids.length),
},
});
MemStack includes a Val benchmark pack for retrieval-only evidence-session recall on the public, cleaned LongMemEval-S split. The runner stores each question's sessions through MemStack's harness memory API, uses the local lexical retriever and in-memory storage, and does not call an answer model or judge. Secret scanning is disabled only for this benchmark's public corpus in the temporary in-memory store; do not use the runner with private data. Its recall metrics are not end-to-end QA accuracy and are not directly comparable to the official LongMemEval QA leaderboard. The published dataset is public, not a hidden evaluation set.
The dataset is not included in this repository. Validate the pack, then fetch only its explicitly named split with [Val](https://github.com/
Source-derived launch command. Check the maintainer’s required arguments and credentials before running:
npx -y @memstack/mcpMerge this template into ~/Library/Application Support/Claude/claude_desktop_config.json. Keep existing servers. Add any arguments, credentials, and permissions required by the maintainer; this template has not been install-tested.
{
"mcpServers": {
"io-github-isiomac-memstack": {
"command": "npx",
"args": [
"-y",
"@memstack/mcp"
]
}
}
}Restart Claude Desktop completely for changes to take effect. Confirm the server appears connected in the client’s tool list, then try a read-only example from its documentation.
Claude Desktop setup reference@memstack/mcpnpmio.github.isiomaC/memstack works with any MCP-compatible client. Copy the config snippet from the Configuration section above and add it to the file shown for your client, then restart the application.
~/Library/Application Support/Claude/claude_desktop_config.jsonRestart Claude Desktop completely for changes to take effect.~/.cursor/mcp.jsonRestart Cursor for changes to take effect..vscode/mcp.jsonReload VS Code window for changes to take effect.~/.codeium/windsurf/mcp_config.jsonRestart Windsurf for changes to take effect..mcp.jsonSave at the project root, then start Claude Code in that project and review the MCP server approval prompt. Keep real credentials out of shared files.