A TypeScript chatbot that starts with no database, provisions a MongoDB Atlas Ephemeral Cluster (in Public Preview) at runtime, and then remembers useful facts across restarts with MongoDB Vector Search and Atlas Automated Embedding.
Repository URL: https://github.com/mongodb-developer/ephemeral-agent-memory
Companion repo for Give Your Chatbot or Agent Persistent Memory in One API Call.
- A terminal chatbot and browser UI that share the same memory module.
- One unauthenticated API call creates a temporary Atlas cluster when memory is needed.
- The model can call
provision_memoryitself, or you can trigger provisioning manually. - Conversation history and extracted facts are stored in MongoDB.
- Semantic recall uses
$vectorSearchwithquery: { text }, so the app never calls an embedding API. threadIdanduserIdfilters keep one conversation's memories out of another.
flowchart LR
User[User] --> CLI[Terminal chatbot]
User --> Web[Browser UI]
CLI --> Turn[src/converse.ts]
Web --> Turn
Turn --> LLM[src/llm.ts]
Turn --> Memory[src/memory]
Memory --> MongoDB[(Atlas Ephemeral Cluster)]
MongoDB --> Messages[messages collection]
MongoDB --> Facts[memories collection]
Facts --> Vector[Atlas Vector Search autoEmbed on fact]
Turn --> Provision[src/provision.ts]
Provision --> Atlas[Ephemeral Cluster API]
The terminal and browser paths both call converse() in src/converse.ts, so recall, model interaction, fact extraction, and optional cluster provisioning stay in one place. MongoDB connection setup and index creation live in src/memory/store.ts, and semantic recall lives in src/memory/index.ts.
git clone https://github.com/mongodb-developer/ephemeral-agent-memory
cd ephemeral-agent-memory
npm install
cp .env.example .env
npm run chatIn .env, set your chat model provider:
LLM_PROVIDER=anthropic
LLM_API_KEY=sk-ant-...Or use an OpenAI-compatible endpoint:
LLM_PROVIDER=openai
LLM_API_KEY=sk-...For local Ollama:
LLM_PROVIDER=openai
LLM_BASE_URL=http://localhost:11434/v1
LLM_CHAT_MODEL=llama3.1
LLM_EXTRACT_MODEL=llama3.1Leave MONGODB_URI empty on the first run. The bot can create a temporary cluster later.
To seed synthetic memory data after setting MONGODB_URI, run:
npm run seed
npm run inspectStart memoryless:
npm run chatSay something durable, exit, restart, and ask what it knows. It should know nothing because no database was attached.
Then ask the bot to remember you:
you > Can you remember this next time we talk?
The model receives provision_memory only while no memory is attached. If it calls the tool, the app creates an Ephemeral Cluster, waits for the Vector Search index to become queryable, backfills the current in-memory conversation, and continues the turn.
You can also provision manually:
you > /provision
The command prints a clusterId, claimUrl, expiry time, and MONGODB_URI. Put that URI in .env if you want this local app to reconnect to the same memory after restart. Open the claim URL if you want to keep the temporary cluster.
npm run webOpen http://localhost:3000.
The browser uses the same converse() function and src/memory module as the terminal. It adds a status pill, a facts panel, a thread switcher, and an expiry banner. See web/README.md for UI notes.
| Command | What it does |
|---|---|
/provision |
Create an Ephemeral Cluster and attach it now |
/attach URI |
Attach a cluster you already have |
/memories |
Show stored facts for the current thread |
/forget |
Wipe memories and transcript for the current thread |
/status |
Show whether memory is attached |
/exit |
Quit |
| Script | What it does |
|---|---|
npm run chat |
Terminal chatbot |
npm run web |
Browser demo at http://localhost:3000 |
npm run provision |
Create an Ephemeral Cluster directly |
npm run seed |
Insert synthetic transcript and memory documents into MONGODB_URI |
npm run inspect |
Print stored messages and memories |
npm run typecheck |
Run TypeScript checks |
| Path | Purpose |
|---|---|
src/converse.ts |
One turn: recall, call the model, remember, and optionally provision memory |
src/llm.ts |
The only file that imports model SDKs |
src/provision.ts |
The one API call that creates an Ephemeral Cluster |
src/memory/config.ts |
Database, index, model, and endpoint settings |
src/memory/store.ts |
MongoDB connection and index setup |
src/memory/index.ts |
Public memory API: remember, recall, history, forget |
src/memory/extract.ts |
Decides what is worth promoting into durable facts |
web/ |
HTTP server and static browser UI |
examples/ |
MCP and LangGraph adapters |
PROMPT.md |
Pasteable prompts for coding agents using this demo |
EDD.md |
Entity Document Diagram and MongoDB schema source of truth |
AGENTS.md |
Repo-specific instructions for coding agents |
| Kind | Collection | Shape | Index |
|---|---|---|---|
| Working memory | In RAM | Short per-thread buffer | None |
| Episodic memory | messages |
One document per message | { threadId, createdAt } |
| Semantic memory | memories |
One document per subject | Unique { threadId, subject } |
| Semantic recall | memories |
$vectorSearch query over fact |
autoEmbed on fact, filters on threadId and userId |
Facts are upserted by { threadId, subject }. If the user changes their city, diet, or goal, the newer fact replaces the older one instead of creating contradictions.
This repo does not store an embedding field and does not require VOYAGE_API_KEY.
The Vector Search index declares fact as an autoEmbed field:
{
type: "autoEmbed",
path: "fact",
modality: "text",
model: "voyage-4",
}Atlas embeds stored facts and query text using the model declared in the index. recall() therefore passes text directly:
{
$vectorSearch: {
index: VECTOR_INDEX,
path: "fact",
query: { text: question },
filter: {
threadId: { $eq: scope.threadId },
userId: { $eq: scope.userId },
},
}
}The filter block is required. Removing it can leak memories between threads.
- Atlas Ephemeral Clusters for runtime provisioning without asking the user to create an account first:
src/provision.ts. - MongoDB Node.js driver connection with
appName=devrel.content.ephemeral-agent-memory:src/memory/store.ts. - Standard indexes for transcript and fact upserts:
ensureRegularIndexes(). - Atlas Vector Search with Automated Embedding on the
factfield:ensureVectorIndex(). - Thread-scoped semantic recall with
$vectorSearch.filter:recall(). - Schema contract and relationships:
EDD.md.
This demo needs a database that can be attached after the chatbot is already running, store raw transcript lines and evolving fact documents, and search memories semantically without a separate vector database. MongoDB keeps the transcript and memory data model simple, uses standard indexes for history and upserts, and uses Atlas Vector Search filters to keep recall scoped to one threadId and userId. The same Atlas cluster handles operational data and semantic search, which keeps the demo small enough for a tutorial while still matching a production agent-memory pattern.
Atlas Automated Embedding uses Voyage AI from inside the Vector Search index. This repo selects voyage-4 in src/memory/config.ts because it gives high-quality general text embeddings for short user facts without adding embedding API calls, API keys, vector dimensions, or batch jobs to the application. Embeddings are generated and stored by Atlas for the fact field, then queried with plain text from recall(). That trades direct embedding control for lower app complexity and fewer moving parts in a runnable example.
- Create an Atlas Ephemeral Cluster
- MongoDB Atlas Vector Search
- Atlas Vector Search automatic embedding
- MongoDB Node.js Driver
- MongoDB Agent Skills
Recommended GitHub description: Give a TypeScript chatbot persistent, thread-scoped memory at runtime with Atlas Ephemeral Clusters, MongoDB Vector Search, and Automated Embedding.
Recommended topics: mongodb, mongodb-search, vector-search, atlas-vector-search, automated-embedding, voyage-ai, agent-memory, chatbot, ai-agents, typescript, nodejs, rag.
| Symptom | Likely cause | Fix |
|---|---|---|
| Search returns zero results after setup | Vector Search index is still building | Wait for queryable; store.ts already polls this |
| A newly stored fact is not recalled | Atlas Automated Embedding is asynchronous | Wait briefly and query again |
Embedding provider rate limit exceeded |
Auto-embedding is rate limited | Retry with backoff |
406 INVALID_VERSION_DATE during provisioning |
Wrong Accept header |
Use application/vnd.atlas.preview+json |
| Another thread's facts appear | Missing Vector Search filter | Restore the threadId and userId filter |
Ephemeral clusters have open network access during the temporary window. Unclaimed clusters pause 2 days after creation and are deleted 7 days after creation. Treat connectionString, claimUrl, and clusterId as secrets: anyone with the connection string can read and write data, and anyone with the claim URL can claim the cluster. Do not store secrets or real personal data in an unclaimed ephemeral cluster. Claim and secure the cluster before using it for anything private.