Skip to content

About

A TypeScript chatbot that starts with no database, provisions a MongoDB Atlas Ephemeral Cluster at runtime, and then remembers useful facts across restarts with MongoDB Vector Search and Atlas Automated Embedding. Companion repo for Give Your Chatbot or Agent Persistent Memory in One API Call tutorial.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Repository files navigation

Ephemeral Agent Memory

A TypeScript chatbot that starts with no database, provisions a MongoDB Atlas Ephemeral Cluster (in Public Preview) at runtime, and then remembers useful facts across restarts with MongoDB Vector Search and Atlas Automated Embedding.

Repository URL: https://github.com/mongodb-developer/ephemeral-agent-memory

Companion repo for Give Your Chatbot or Agent Persistent Memory in One API Call.

Capabilities

  • A terminal chatbot and browser UI that share the same memory module.
  • One unauthenticated API call creates a temporary Atlas cluster when memory is needed.
  • The model can call provision_memory itself, or you can trigger provisioning manually.
  • Conversation history and extracted facts are stored in MongoDB.
  • Semantic recall uses $vectorSearch with query: { text }, so the app never calls an embedding API.
  • threadId and userId filters keep one conversation's memories out of another.

Architecture Overview

flowchart LR
  User[User] --> CLI[Terminal chatbot]
  User --> Web[Browser UI]
  CLI --> Turn[src/converse.ts]
  Web --> Turn
  Turn --> LLM[src/llm.ts]
  Turn --> Memory[src/memory]
  Memory --> MongoDB[(Atlas Ephemeral Cluster)]
  MongoDB --> Messages[messages collection]
  MongoDB --> Facts[memories collection]
  Facts --> Vector[Atlas Vector Search autoEmbed on fact]
  Turn --> Provision[src/provision.ts]
  Provision --> Atlas[Ephemeral Cluster API]
Loading

The terminal and browser paths both call converse() in src/converse.ts, so recall, model interaction, fact extraction, and optional cluster provisioning stay in one place. MongoDB connection setup and index creation live in src/memory/store.ts, and semantic recall lives in src/memory/index.ts.

Quick Start

git clone https://github.com/mongodb-developer/ephemeral-agent-memory
cd ephemeral-agent-memory
npm install
cp .env.example .env
npm run chat

In .env, set your chat model provider:

LLM_PROVIDER=anthropic
LLM_API_KEY=sk-ant-...

Or use an OpenAI-compatible endpoint:

LLM_PROVIDER=openai
LLM_API_KEY=sk-...

For local Ollama:

LLM_PROVIDER=openai
LLM_BASE_URL=http://localhost:11434/v1
LLM_CHAT_MODEL=llama3.1
LLM_EXTRACT_MODEL=llama3.1

Leave MONGODB_URI empty on the first run. The bot can create a temporary cluster later.

To seed synthetic memory data after setting MONGODB_URI, run:

npm run seed
npm run inspect

Try The Memory Flow

Start memoryless:

npm run chat

Say something durable, exit, restart, and ask what it knows. It should know nothing because no database was attached.

Then ask the bot to remember you:

you > Can you remember this next time we talk?

The model receives provision_memory only while no memory is attached. If it calls the tool, the app creates an Ephemeral Cluster, waits for the Vector Search index to become queryable, backfills the current in-memory conversation, and continues the turn.

You can also provision manually:

you > /provision

The command prints a clusterId, claimUrl, expiry time, and MONGODB_URI. Put that URI in .env if you want this local app to reconnect to the same memory after restart. Open the claim URL if you want to keep the temporary cluster.

Browser Demo

npm run web

Open http://localhost:3000.

The browser uses the same converse() function and src/memory module as the terminal. It adds a status pill, a facts panel, a thread switcher, and an expiry banner. See web/README.md for UI notes.

Commands

Command What it does
/provision Create an Ephemeral Cluster and attach it now
/attach URI Attach a cluster you already have
/memories Show stored facts for the current thread
/forget Wipe memories and transcript for the current thread
/status Show whether memory is attached
/exit Quit

Scripts

Script What it does
npm run chat Terminal chatbot
npm run web Browser demo at http://localhost:3000
npm run provision Create an Ephemeral Cluster directly
npm run seed Insert synthetic transcript and memory documents into MONGODB_URI
npm run inspect Print stored messages and memories
npm run typecheck Run TypeScript checks

Project Layout

Path Purpose
src/converse.ts One turn: recall, call the model, remember, and optionally provision memory
src/llm.ts The only file that imports model SDKs
src/provision.ts The one API call that creates an Ephemeral Cluster
src/memory/config.ts Database, index, model, and endpoint settings
src/memory/store.ts MongoDB connection and index setup
src/memory/index.ts Public memory API: remember, recall, history, forget
src/memory/extract.ts Decides what is worth promoting into durable facts
web/ HTTP server and static browser UI
examples/ MCP and LangGraph adapters
PROMPT.md Pasteable prompts for coding agents using this demo
EDD.md Entity Document Diagram and MongoDB schema source of truth
AGENTS.md Repo-specific instructions for coding agents

How Memory Is Stored

Kind Collection Shape Index
Working memory In RAM Short per-thread buffer None
Episodic memory messages One document per message { threadId, createdAt }
Semantic memory memories One document per subject Unique { threadId, subject }
Semantic recall memories $vectorSearch query over fact autoEmbed on fact, filters on threadId and userId

Facts are upserted by { threadId, subject }. If the user changes their city, diet, or goal, the newer fact replaces the older one instead of creating contradictions.

Automated Embedding

This repo does not store an embedding field and does not require VOYAGE_API_KEY.

The Vector Search index declares fact as an autoEmbed field:

{
  type: "autoEmbed",
  path: "fact",
  modality: "text",
  model: "voyage-4",
}

Atlas embeds stored facts and query text using the model declared in the index. recall() therefore passes text directly:

{
  $vectorSearch: {
    index: VECTOR_INDEX,
    path: "fact",
    query: { text: question },
    filter: {
      threadId: { $eq: scope.threadId },
      userId: { $eq: scope.userId },
    },
  }
}

The filter block is required. Removing it can leak memories between threads.

MongoDB Features Demonstrated

  • Atlas Ephemeral Clusters for runtime provisioning without asking the user to create an account first: src/provision.ts.
  • MongoDB Node.js driver connection with appName=devrel.content.ephemeral-agent-memory: src/memory/store.ts.
  • Standard indexes for transcript and fact upserts: ensureRegularIndexes().
  • Atlas Vector Search with Automated Embedding on the fact field: ensureVectorIndex().
  • Thread-scoped semantic recall with $vectorSearch.filter: recall().
  • Schema contract and relationships: EDD.md.

Why MongoDB?

This demo needs a database that can be attached after the chatbot is already running, store raw transcript lines and evolving fact documents, and search memories semantically without a separate vector database. MongoDB keeps the transcript and memory data model simple, uses standard indexes for history and upserts, and uses Atlas Vector Search filters to keep recall scoped to one threadId and userId. The same Atlas cluster handles operational data and semantic search, which keeps the demo small enough for a tutorial while still matching a production agent-memory pattern.

Why Voyage AI?

Atlas Automated Embedding uses Voyage AI from inside the Vector Search index. This repo selects voyage-4 in src/memory/config.ts because it gives high-quality general text embeddings for short user facts without adding embedding API calls, API keys, vector dimensions, or batch jobs to the application. Embeddings are generated and stored by Atlas for the fact field, then queried with plain text from recall(). That trades direct embedding control for lower app complexity and fewer moving parts in a runnable example.

Additional Resources

Repository Metadata

Recommended GitHub description: Give a TypeScript chatbot persistent, thread-scoped memory at runtime with Atlas Ephemeral Clusters, MongoDB Vector Search, and Automated Embedding.

Recommended topics: mongodb, mongodb-search, vector-search, atlas-vector-search, automated-embedding, voyage-ai, agent-memory, chatbot, ai-agents, typescript, nodejs, rag.

Troubleshooting

Symptom Likely cause Fix
Search returns zero results after setup Vector Search index is still building Wait for queryable; store.ts already polls this
A newly stored fact is not recalled Atlas Automated Embedding is asynchronous Wait briefly and query again
Embedding provider rate limit exceeded Auto-embedding is rate limited Retry with backoff
406 INVALID_VERSION_DATE during provisioning Wrong Accept header Use application/vnd.atlas.preview+json
Another thread's facts appear Missing Vector Search filter Restore the threadId and userId filter

Safety

Ephemeral clusters have open network access during the temporary window. Unclaimed clusters pause 2 days after creation and are deleted 7 days after creation. Treat connectionString, claimUrl, and clusterId as secrets: anyone with the connection string can read and write data, and anyone with the claim URL can claim the cluster. Do not store secrets or real personal data in an unclaimed ephemeral cluster. Claim and secure the cluster before using it for anything private.

About

A TypeScript chatbot that starts with no database, provisions a MongoDB Atlas Ephemeral Cluster at runtime, and then remembers useful facts across restarts with MongoDB Vector Search and Atlas Automated Embedding. Companion repo for Give Your Chatbot or Agent Persistent Memory in One API Call tutorial.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Used by

Contributors

Languages