Memory Server

What is Memory Server about?

Memory Server offers a simple local RAG (Retrieval‑Augmented Generation) service that lets you memorize arbitrary text – single sentences, multiple entries, or full PDF documents – and later retrieve the most relevant fragments through semantic similarity.

How to use Memory Server?

Clone the repository and install the fast Python package manager uv.
Start the supporting services with Docker Compose (docker‑compose up). This launches ChromaDB and Ollama.
Pull the embedding model inside the Ollama container (docker exec -it ollama ollama pull all-minilm:l6-v2).
Run the MCP server using the configuration provided (see serverConfig).
Interact via the MCP tool:
- memorize_text for a single passage.
- memorize_multiple_texts for a list.
- memorize_pdf_file to ingest PDFs in 20‑page chunks.
- Ask natural‑language queries to retrieve relevant stored texts.

Key features of Memory Server

Semantic memorization: Stores text as embeddings, not just keywords.
Chunked PDF ingestion: Reads PDFs in 20‑page batches, auto‑chunks, and memorizes each segment.
Conversational chunking: LLM can split large texts into meaningful pieces interactively.
Local deployment: All components run locally via Docker; no external API keys required.
ChromaDB admin GUI: Web UI at http://localhost:8322 for browsing and managing the vector store.

Use cases of Memory Server

Building a personal knowledge base that can be queried with natural language.
Augmenting chatbots with domain‑specific documents (e.g., policies, manuals).
Research assistants that retain and retrieve excerpts from large PDFs or corpora.
Rapid prototyping of RAG pipelines without relying on cloud services.

FAQ

Q: Do I need an internet connection? A: Only for the initial ollama pull of the embedding model. After that the service runs entirely offline.

Q: Which embedding model is used? A: all-minilm:l6-v2 from Ollama, a lightweight sentence‑embedding model.

Q: Can I change the storage port? A: Yes, adjust CHROMADB_PORT and OLLAMA_PORT in the server configuration.

Q: How large a PDF can be processed? A: Any size; the tool processes it in 20‑page increments, looping until the end.

Q: Is there a way to delete memorized texts? A: Use the ChromaDB admin GUI or invoke ChromaDB’s collection deletion APIs.

Memory Server (mcp-rag-local)

This MCP server provides a simple API for storing and retrieving text passages based on their semantic meaning, not just keywords. It uses Ollama for generating text embeddings and ChromaDB for vector storage and similarity search. You can "memorize" any text and later retrieve the most relevant stored texts for a given query.

Example Usage

Memorize a Text

You can simply ask the LLM to memorize a text for you in natural language:

User: Memorize this text: "Singapore is an island country in Southeast Asia."

LLM: Text memorized successfully.

Memorize Multiple Texts

You can also ask the LLM to memorize several texts at once:

User: Memorize these texts:

Singapore is an island country in Southeast Asia.
It is about one degree of latitude north of the equator.
It is a major financial and shipping hub.

LLM: All texts memorized successfully.

This will store all provided texts for later semantic retrieval.

Memorize a PDF File

You can also ask the LLM to memorize the contents of a PDF file via memorize_pdf_file. The MCP tool will read up to 20 pages at a time from the PDF, return the extracted text, and have the LLM chunk it into meaningful segments. The LLM then uses the memorize_multiple_texts tool to store these chunks.

This process is repeated: the MCP tool continues to read the next 20 pages, the LLM chunks and memorizes them, and so on, until the entire PDF is processed and memorized.

User: Memorize this PDF file: C:\path\to\document.pdf

LLM: Reads the first 20 pages, chunks the text, stores the chunks, and continues with the next 20 pages until the whole document is memorized.

You can also specify a starting page if you want to begin from a specific page:

MCP to LLM: Memorize this PDF file starting from page 40: C:\path\to\document.pdf

LLM: Reads pages 40–59, chunks and stores the text, then continues with the next set of pages until the end of the document.

Example: Conversational Chunking and Memorizing Large Text

If you have a long text, you can ask the LLM to help you split it into short, meaningful chunks and store them. For example:

User: Please chunk the following long text and memorize all the chunks.

{large body of text}

LLM: Splits the text into short, relevant segments and calls memorize_multiple_texts to store them. If the text is too long to store in one go, the LLM will continue chunking and storing until the entire text is memorized.

User: Are all the text chunks stored?

LLM: Checks and, if not all are stored, continues until the process is complete.

This conversational approach ensures that even very large texts are fully chunked and memorized, with the LLM handling the process interactively.

Retrieve Similar Texts

To recall information, just ask the LLM a question:

User: What is Singapore?

LLM: Returns the most relevant stored texts along with a human-readable description of their relevance.

Setup Instructions

0. Clone this repository

First, clone this git repository and change into the cloned directory:

git clone <repository-url>
cd mcp-rag-local

1. Install uv

Install uv (a fast Python package manager):

curl -LsSf https://astral.sh/uv/install.sh | sh

1a. Windows Installation

If you are on Windows, install uv using PowerShell:

powershell -ExecutionPolicy ByPass -c "irm https://astral.sh/uv/install.ps1 | iex"

2. Start the services

Run the following command to start ChromaDB and Ollama using Docker Compose:

docker-compose up

3. Pull the embedding model

After the containers are running, pull the embedding model for Ollama:

docker exec -it ollama ollama pull all-minilm:l6-v2

4. MCP Server Config

Add the following to your MCP server configuration:

"mcp-rag-local": {
  "command": "uv",
  "args": [
    "--directory",
    "path\\to\\mcp-rag-local",
    "run",
    "main.py"
  ],
  "env": {
    "CHROMADB_PORT": "8321",
    "OLLAMA_PORT": "11434"
  }
}

5. Viewing and Managing Memory (ChromaDB Admin GUI)

A web-based GUI for ChromaDB(Memory Server's db) is included for easy inspection and management of stored memory.

The admin UI is available at: http://localhost:8322
You can use this interface to browse, search, and manage the vector database contents.

Memory Server

Memory Server Overview

Memory Server

What is Memory Server about?

How to use Memory Server?

Key features of Memory Server

Use cases of Memory Server

FAQ

Memory Server's README

Memory Server (mcp-rag-local)

Example Usage

Memorize a Text

Memorize Multiple Texts

Memorize a PDF File

Example: Conversational Chunking and Memorizing Large Text

Retrieve Similar Texts

Setup Instructions

0. Clone this repository

1. Install uv

1a. Windows Installation

2. Start the services

3. Pull the embedding model

4. MCP Server Config

5. Viewing and Managing Memory (ChromaDB Admin GUI)

Memory Server Reviews

Login Required

Similar MCP Servers like Memory Server

Memory

Cognee

Basic Memory

Obsidian Model Context Protocol

Mcp Server Chatsum

Minima

Mcp Server Qdrant

MCP Memory Service

Context Portal

Actions

Memory Server's Information

Configuration

Claude Code (Terminal)

Configure Clients