by kagura-ai
Provides adaptive memory for AI agents and teams, combining hybrid semantic‑keyword search with a self‑learning neural memory graph that improves automatically as searches are performed.
Kagura Memory Cloud delivers an adaptive, team‑scale knowledge store that goes beyond traditional RAG. It stores memories as structured records, indexes them with BM25, vector embeddings (Qdrant), and a Hebbian‑learned graph, then continuously strengthens connections between related items each time they are retrieved.
git clone https://github.com/kagura-ai/memory-cloud.git && cd memory-cloud && ./setup.sh.python -m src.cli.setup_env inside the backend folder to generate secrets and add your OpenAI or self‑hosted embedding provider keys.docker compose up -d.python -m src.cli.create_admin.http://localhost:8080, Web UI at http://localhost:3000, API docs at http://localhost:8080/redoc..mcp.json file pointing to http://localhost:8080/mcp/w/{workspace_id}.explore() surfaces hidden relationships.Q: Do I need an OpenAI API key?
A: Only if you use OpenAI embeddings. You can alternatively run a self‑hosted inference server (Ollama, vLLM) and set EMBEDDING_PROVIDER=self_hosted.
Q: Can I run the whole stack without Docker?
A: Yes. Start PostgreSQL, Qdrant, and Redis manually, then uvicorn src.main:app --host 0.0.0.0 --port 8080 after installing the Python dependencies.
Q: How does the graph learn without expensive LLM calls?
A: Hebbian learning updates edge weights whenever recall() returns multiple memories together, requiring no extra model inference.
Q: Is the system safe for sensitive data? A: Secrets are stored encrypted (zero‑knowledge); the vector store can be run on‑premise, and all traffic can be terminated behind TLS.
Q: What is the difference between recall() and explore()?
A: recall() performs precision‑oriented hybrid search (vector + BM25) possibly with AI reranking. explore() traverses the neural graph to surface related memories that keyword search might miss.
Your AI forgets everything after each conversation. Kagura fixes that — and gets smarter every time you search.
Most AI memory tools are just vector databases with a chat wrapper. Kagura is different — it implements the full LLM Knowledge Base pattern (Karpathy's LLM Wiki) at team scale:
| Approach | Storage | Compounding | Scale |
|---|---|---|---|
| Vector DB / RAG | Embedded chunks | None — retrieve-only | Any |
| Karpathy's LLM Wiki | Markdown files | LLM rewrites pages | Personal (~100 pages) |
| Kagura Memory Cloud | PostgreSQL + Qdrant + Neural graph | Hebbian + Sleep Maintenance | Team / org |
| Feature | Description |
|---|---|
| Adaptive Memory | Every search automatically strengthens connections between related memories. The more you use it, the better explore() discovers hidden relationships. |
| Hybrid Search | Semantic (OpenAI / self-hosted) + BM25 keyword — 96% top-1 accuracy |
| AI Reranking | Self-hosted (Ollama/vLLM — local, free), Voyage AI, or Cohere — cross-encoder reranking for precision |
| Neural Memory Graph | Hebbian learning builds a knowledge graph in the background. explore() traverses it for serendipitous discovery. |
| Agent Memory Substrate | Beyond a knowledge store: delivery modes (pinned / time-triggered), a server-stamped trust boundary, an agent state lane, and a retrieval-feedback signal — the primitives an autonomous agent loop needs. |
| Agent Control Plane (preview) | Workspace-scoped Agent Registry, subtractive context bindings, agent-bound member keys, lifecycle kill switches, and one-call session bootstrap. Introduced in v0.49.0. |
| 64 MCP Tools | Memory, Agent Substrate, Agent Control Plane, Neural edges, Contexts, Tags, Files (R2), Analyses (Memory Analysis), Resources, Secrets, Sleep Maintenance, Usage, API-Key Bindings |
| Multi-Provider | OpenAI or self-hosted (Ollama, vLLM — local, private, zero cost) for embeddings |
| Team Ready | Workspaces, RBAC, context isolation, shared memory |
| Web UI | Next.js dashboard — contexts, search settings, member management |
| 5-Minute Setup | ./setup.sh and you're done |
Workspace (team/org)
├── Context A ("my-project") ← like a folder
│ ├── Memory 1 ← 3-layer: summary / context / content
│ ├── Memory 2
│ └── Neural edges (Hebbian) ← automatic connections
├── Context B ("learning-notes")
│ └── ...
└── Members (Owner/Admin/Member/Viewer)
Karpathy's LLM Wiki pattern describes a 5-layer "living knowledge base" — beyond traditional RAG. Kagura implements all 5 layers at team scale:
| Layer | Kagura Implementation | Difference from Karpathy's pattern |
|---|---|---|
| Ingest | REST /api/v1/memory, MCP remember, R2 file storage, resource tokens |
+ binary blobs, + multi-tenant |
| Compile | MCP-as-compile-API — chat agent compiles via structured tool calls (remember(summary, content, type, tags)) + Sleep Maintenance for batch consolidation |
Continuous micro-compile (not batch wiki rewrite) — schema-enforced output |
| Index | Triple index: BM25 (keyword) + Qdrant (semantic) + Hebbian graph (relational) — all auto-maintained | No manual index.md upkeep |
| Query | Hybrid Search + AI Reranker + explore graph traversal |
Beyond markdown grep — supports semantic + relational queries |
| Enhance | Hebbian learning — every recall() strengthens edges between co-retrieved memories. Sleep Maintenance consolidates periodically. |
Background graph evolution (zero LLM cost) vs LLM-driven page rewrites |
Compounding loop: Currently explicit (user/agent calls remember() after synthesizing answers). Auto-write-back of synthesized answers is intentionally opt-in to keep noise low.
Kagura separates precision search and discovery into two independent paths, each optimized for its purpose:
recall() ──→ Hybrid Search (semantic + BM25) ──→ [Reranker] ──→ Precise results
│
└──→ Hebbian Learning (background) ──→ Graph edges grow
│
explore() ──→ Graph Traversal (Neural Memory) ←─────────────────┘ Related discoveries
recall() — Precision search. Hybrid (semantic 60% + BM25 40%) with optional AI reranking. Returns the most relevant memories.explore() — Discovery. Traverses the Neural Memory graph to find related memories that keyword search would miss.recall() silently strengthens edges between co-retrieved memories. No explicit training needed — the graph grows organically as you use the system.This separation is intentional: mixing graph signals into recall degrades precision (validated via benchmarks). Instead, each path does what it's best at.
Data isolation: All data is filtered by workspace_id → context_id → user_id. Memories never leak across boundaries. Single Qdrant collection with payload filtering.
Tech stack: FastAPI (async) · PostgreSQL · Qdrant · Redis · Next.js 16 · OAuth2 · MCP over Streamable HTTP
Vector backend: Qdrant by default. A single-process self-hosted / CLI / edge deployment can instead run the embedded LanceDB backend — "Kagura Lite" (preview) with no separate Qdrant server (KAGURA_VECTOR_BACKEND=lance, cd backend && uv sync --locked --extra lite). Not for multi-worker / SaaS (LanceDB is single-writer). See Deployment → Embedded Vector Backend.
| Minimum | Recommended | |
|---|---|---|
| CPU | 2 cores | 4+ cores |
| RAM | 4 GB | 8+ GB |
| Disk | 10 GB free | 20+ GB free |
One-line setup:
git clone https://github.com/kagura-ai/memory-cloud.git
cd memory-cloud
./setup.sh
With Claude Code:
git clone https://github.com/kagura-ai/memory-cloud.git
cd memory-cloud
claude # then run /setup
Step-by-step setup:
# 1. Clone
git clone https://github.com/kagura-ai/memory-cloud.git
cd memory-cloud
# 2. Configure environment (generates secrets, prompts for API keys)
(cd backend && python3 -m src.cli.setup_env)
# 3. Start all services
docker compose up -d
# 4. Run migrations
(cd backend && alembic upgrade head)
# 5. Create admin account (interactive — sets password, MFA, API key, embedding provider)
(cd backend && python3 -m src.cli.create_admin)
# Backend API: http://localhost:8080
# Frontend UI: http://localhost:3000
# API docs: http://localhost:8080/redoc
.env.local settings (auto-configured by setup_env):
| Setting | Required | Description |
|---|---|---|
API_KEY_SECRET |
Yes | Secret for API key encryption (auto-generated) |
JWT_SECRET |
Yes | Secret for JWT tokens (auto-generated) |
OPENAI_API_KEY |
Yes* | OpenAI API key for embeddings |
SELF_HOSTED_BASE_URL |
No | Self-hosted backend URL (default: http://localhost:11434) |
EMBEDDING_PROVIDER |
No | openai (default) or self_hosted |
GOOGLE_CLIENT_ID/SECRET |
No | Google OAuth2 login (optional — password login available) |
GITHUB_CLIENT_ID/SECRET |
No | GitHub OAuth2 login (optional) |
* Either OPENAI_API_KEY or a running self-hosted inference server (e.g. Ollama) is required for memory features.
| Command | Purpose |
|---|---|
python3 -m src.cli.setup_env |
Generate secrets + configure .env.local (run before Docker) |
python3 -m src.cli.create_admin |
Create admin + workspace + API key + .mcp.json + embedding setup |
python3 -m src.cli.reset_password |
Reset password and/or MFA |
python3 -m src.cli.delete_admin |
Delete admin (for re-creation) |
python3 -m src.cli.transfer_context_creator |
Move contexts created by one user to another (CLI admin → OAuth account, see docs/deployment.md) |
Run from
backend/directory. Docker API container must be running.
brew install python@3.11 nodesudo apt install docker.io docker-compose-v2 python3.11 nodejs npm.env.local (DATABASE_URL, QDRANT_URL, ENVIRONMENT=production, CORS_ORIGINS)frontend/.env.example to frontend/.env.local and set:
NEXT_PUBLIC_API_URL — backend URL (default: http://localhost:8080)NEXT_PUBLIC_APP_URL — frontend URL for metadataNEXT_PUBLIC_PLAN_FREE_DISPLAY_NAME / BASIC / PRO / PROMAX — plan display name customization (default: S/M/L/XL)Works with Claude Code, Claude Desktop, Claude Chat, ChatGPT, Gemini CLI, and any Streamable-HTTP MCP client.
Claude Code (3 steps):
http://localhost:3000/workspace/integrations/api-keys to create an API key.mcp.json.example to .mcp.json and fill in your workspace ID and API key:cp .mcp.json.example .mcp.json
# Edit .mcp.json — set workspace_id (from URL bar) and API key
.mcp.json.example ships with the all-tools URL. Set "url" to one of:
http://localhost:8080/mcp/w/{workspace_id}http://localhost:8080/mcp/w/{workspace_id}?profile=corePick core when your client loads every tool schema at session start (it is about 65% smaller). It lists the 12 memory and context tools and leaves out Sleep, analyses, files, edges, secrets, resources and the agent control plane — those stay callable, they are just not listed; switch back to the default URL to see them. See Tool Profiles.
You: "Remember: our API uses JWT with 1h expiry and refresh token rotation"
→ AI calls remember() — stored permanently
You: "What do we know about auth?"
→ AI calls recall() — finds it instantly, even months later
.mcp.jsonis in.gitignore— never commit it (contains API keys).
Full setup guide — every client, the memory-sync hook, the ready-to-use .claude/ templates, the kagura-memory Claude Code plugin (skills + tool-guardrail hooks), and the WSL2 networking note: MCP Client Setup
64 tools across 13 categories: Memory (remember / recall / explore …), Agent Substrate (pinned + time-triggered delivery, state, measurements, feedback), Agent Control Plane (preview), Neural Edges, Contexts, Tags, Files (R2), Analyses (Memory Analysis), Resources, Secrets (zero-knowledge), Sleep Maintenance, Usage, and API-Key Bindings — each with per-role access control.
Tool-by-tool reference with required roles: MCP Tools Reference
A client does not have to list all 64: the core URL above (?profile=core) lists 12, and ?tools=remember,recall lists exactly the tools you name — see Tool Profiles.
In addition to MCP tools, a full REST API is available:
/api/v1/memory/*)/api/v1/contexts/*)/api/v1/agents/*)/api/v1/files/*, up to 100 MiB); legacy /api/v1/attachments/* routes return 410 Gone/api/v1/analyses/*)/api/v1/resources/*)/api/v1/workspaces/*)/api/v1/admin/*)/api/v1/config/secrets/*)Full API documentation: http://localhost:8080/redoc
Two OAuth2 providers are supported:
GOOGLE_CLIENT_ID and GOOGLE_CLIENT_SECRETGITHUB_CLIENT_ID and GITHUB_CLIENT_SECRETUsers with the same email address across providers share a single account. Password + MFA login is available without any OAuth provider (see Quick Start). An existing account can also add a password from its profile and then sign in with its verified email address and that password; see Deployment → Email + password sign-in.
Plans control per-workspace resource limits (contexts / memories / MCP calls per day). Four tiers ship by default: S (free), M (basic), L (pro) and XL (promax). For self-hosted single-user setups, assign the XL (promax) plan to your workspace — it is the only tier that may create resources, connectors and public contexts (numeric limits are env-overridable; a tier's feature set is not). Defaults, environment-variable overrides, and optional Stripe self-service billing: Deployment → Plan Tiers
This project is designed to be developed with Claude Code and Kagura Memory Cloud itself — pre-configured slash commands, safety hooks, sub-agents, and rules load automatically from .claude/. Setup and the full tooling reference: Contributing → Development with Claude Code
API reference — two complementary entry points:
http://localhost:8080/redoc — auto-generated from FastAPI, always in sync with the running backendConcepts & guides:
KaguraClient (MCP) plus REST clients for resources, files, secrets, workspaces, and agent bootstrap, and a document FileIngestorSee CONTRIBUTING.md for development setup, code style, and PR workflow.
Please log in to share your review and rating for this MCP.
Explore related MCPs that share similar capabilities and solve comparable challenges
by modelcontextprotocol
A basic implementation of persistent memory using a local knowledge graph. This lets Claude remember information about the user across chats.
by topoteretes
Provides dynamic memory for AI agents through modular ECL (Extract, Cognify, Load) pipelines, enabling seamless integration with graph and vector stores using minimal code.
by basicmachines-co
Enables persistent, local‑first knowledge management by allowing LLMs to read and write Markdown files during natural conversations, building a traversable knowledge graph that stays under the user’s control.
by agentset-ai
Provides an open‑source platform to build, evaluate, and ship production‑ready retrieval‑augmented generation (RAG) and agentic applications, offering end‑to‑end tooling from ingestion to hosting.
by smithery-ai
Provides read and search capabilities for Markdown notes in an Obsidian vault for Claude Desktop and other MCP clients.
by chatmcp
Summarize chat messages by querying a local chat database and returning concise overviews.
by dmayboroda
Provides on‑premises conversational retrieval‑augmented generation (RAG) with configurable Docker containers, supporting fully local execution, ChatGPT‑based custom GPTs, and Anthropic Claude integration.
by qdrant
Provides a Model Context Protocol server that stores and retrieves semantic memories using Qdrant vector search, acting as a semantic memory layer.
by doobidoo
Provides a universal memory service with semantic search, intelligent memory triggers, OAuth‑enabled team collaboration, and multi‑client support for Claude Desktop, Claude Code, VS Code, Cursor and over a dozen AI applications.