
MARU is an open-source RAG (Retrieval-Augmented Generation) chatbot engine built for enterprise environments. The core principle behind MARU is effective integration with existing corporate data β the key to any successful enterprise RAG system.
To provide seamless user experiences and easy compatibility with enterprise infrastructure, MARU is designed to align with corporate document management and access systems β including first-class support for Korean document formats (HWP/HWPX). We open-sourced MARU to help developers who face similar real-world challenges in enterprise AI integration.
- Broad format support β PDF, DOCX, PPTX, XLSX, CSV, HTML, JSON, Markdown, plain text/code, and Korean documents (HWP / HWPX / HWPML) parsed via the KorDoc MCP server
- Pluggable parser routing β LangChain loaders by default; route any share of dual-support formats (pdf/docx/xls/xlsx) to KorDoc with
kordoc_mcp_ratioto A/B-compare parsing quality (the parser used is recorded in document metadata) - Change-aware re-upload β documents are identified by (team, path); re-uploading a modified file updates the document in place (embeddings replaced), and
/ingest/checkskips unchanged files by fingerprint - Failure recovery β failed documents keep their error message and are re-processed via the retry API (per document, or per folder in queue mode)
- Safe deletion β single-document and folder-subtree deletion with a cooperative-cancel state machine: in-flight ingests are marked
deletingand finalized by the worker, so deletes never race embedding writes (chunks, DB rows, and storage files are all cleaned up) - Scales out with a task queue (optional) β offload embedding to ARQ workers over Redis, with a shared Chroma server so the API and workers see one consistent vector store
- Search/no-search routing β a classifier node decides whether a question needs document retrieval (biased toward search for fact/regulation questions)
- Memory-aware conversations β user facts/preferences and rolling session summaries are loaded up front; follow-up questions ("그건μ?") are rewritten into self-contained search queries using prior context
- RAG with self-correction β intent rewrite β keyword extraction β retrieval β sufficiency evaluation with retry β reranking (optional cross-encoder with
reranker_min_scorefiltering) - Per-message team scoping β document search is restricted to the requester's teams (and optionally a subset per message)
- Feedback collection β interrupt/resume flow for answer scoring and reasons, persisted per conversation
- Observability β every turn carries
user_id/team_ids/session_id/graph_idinto LangSmith traces and checkpoint metadata
- Auth β access (2h) / refresh (30d) token flow, corporate email verification (SMTP OTP), domain allowlist, role hierarchy (anonymous / editor / admin)
- Teams β N:N membership, team-admin-gated destructive actions, invitation flow for not-yet-registered users
- Sessions β server-owned session ids; the last session is resumed only within a 7-day idle window (older conversations stay in history)
- Audit trail β upload / re-upload / delete / ingest success / ingest error per document
graph LR
Client[Web UI / CLI]
Client -->|REST| API["FastAPI<br/>auth Β· teams Β· sessions Β· memory Β· ingest Β· config"]
Client -->|WebSocket| WS["/chat/connect"]
WS --> Router{"graph_router<br/>(team-scoped selection)"}
Router --> Chat["chat graph"]
Router -.-> Other["other graphs<br/>(registry, opt-in per team)"]
API --> RDB[("Relational DB<br/>User Β· Team Β· Session<br/>Conversation Β· UserMemory")]
Chat --> RDB
Chat --> CP[("Checkpointer<br/>(turn state)")]
Chat --> VDB[("Vector DB<br/>Chroma (embedded or server)")]
API -->|queue on: enqueue| Q[("Redis / ARQ")]
Q --> Worker["Ingest worker(s)"]
Worker --> VDB
API -->|queue off: in-process| Ingest["Ingest graph<br/>sync β parse β process"]
Ingest --> VDB
- Two routing levels:
graph_router(L1) picks which graph to run within the team's accessible set; each graph then routes internally (L2). - Ingest graph: one compiled graph serves every entry point β CLI sync, API upload (in-process or via the ARQ worker), and retries.
- Memory loop: the chat graph reads memory up front (
context_builder) and writes it back at the end (summarize,memory_extractor) β see below.
Regenerate with
python scripts/draw_graph.pyβ traced from the actualcreate_rag_graph()topology.
graph TD;
__start__([<p>__start__</p>]):::first
context_builder(context_builder)
route(route)
generate(generate)
search_entry(search_entry)
intent(intent)
keywords(keywords)
retrieve(retrieve)
evaluate(evaluate)
rerank(rerank)
format(format)
score(score)
reason(reason)
summarize(summarize)
memory_extractor(memory_extractor)
__end__([<p>__end__</p>]):::last
__start__ --> context_builder;
context_builder --> route;
evaluate -. retry .-> keywords;
evaluate -.-> rerank;
format --> generate;
generate -.-> score;
generate -.-> summarize;
intent --> keywords;
keywords --> retrieve;
reason --> summarize;
rerank --> format;
retrieve --> evaluate;
route -.-> generate;
route -.-> search_entry;
score -.-> reason;
score -.-> summarize;
search_entry --> intent;
summarize --> memory_extractor;
memory_extractor --> __end__;
classDef default fill:#f2f0ff,line-height:1.2
classDef first fill-opacity:0
classDef last fill:#bfb6fc
- context_builder β loads user memory (facts/preferences) + session summary/recent turns.
- route β classifies search vs. direct answer.
- search_entry β β¦ β format β RAG retrieval with an evaluate/retry loop.
- generate β produces the answer (with retrieved + memory context).
- score/reason β optional feedback collection (interrupt/resume).
- summarize β memory_extractor β write-back: turn/session summaries + durable user memory.
- Python 3.10+
- Node.js / npx β for Korean document parsing via KorDoc (set
kordoc_mcp_enabled: falseto run without it) - Docker (optional) β Redis + Chroma server for the task-queue deployment mode
git clone https://github.com/kc-ml2/MARU-Lang.git
cd MARU-Lang
python -m venv venv && source venv/bin/activate
pip install -e .
maru install # scaffolds maru_app/ with maru_config.yamlOne file configures everything. The keys you'll touch first:
# --- LLM (OpenAI-compatible servers like vLLM work via the openai provider) ---
llms:
- name: my-llm
provider: openai # openai | anthropic | google | ollama | vllm
model_name: gpt-4o
api_key: ${ENV:OPENAI_API_KEY}
# base_url: http://my-vllm:8000/v1
# --- Embedding & retrieval ---
embedding_model: BAAI/bge-m3
retriever_top_k: 5
reranker_enabled: false # cross-encoder reranking (+ reranker_min_score filter)
# --- Vector store ---
vector_db_url: chroma://data/chroma/maru # embedded (single process)
# vector_db_url: chroma+http://localhost:8001/maru # Chroma server (required for queue mode)
# --- Ingest task queue (optional) ---
task_queue_enabled: false # true β embedding runs on ARQ workers (needs redis_url + Chroma server)
redis_url: redis://localhost:6379
# --- Korean document parser ---
kordoc_mcp_enabled: true # hwp/hwpx/hwpml via KorDoc MCP (needs Node/npx)
# kordoc_mcp_ratio: 0.0 # share of pdf/docx/xls/xlsx routed to KorDoc (A/B comparison)${ENV:VAR} / ${ENV:VAR:default} interpolation is supported throughout.
Simplest β single process (embedded Chroma, in-process embedding):
maru run # API server + interactive chat REPL in one commandProduction-style β task queue with shared stores:
# one-time: shared infrastructure
docker run -d --name maru-redis --restart unless-stopped -p 6379:6379 redis:7
docker run -d --name maru-chroma --restart unless-stopped -p 8001:8000 \
-v $HOME/maru-chroma-data:/data chromadb/chroma
# maru_config.yaml: task_queue_enabled: true, vector_db_url: chroma+http://localhost:8001/maru
maru serve --worker 1 # API + co-launched ARQ ingest worker (use under systemd)
maru run --attach # attach a chat REPL to the running server (same machine)Chat REPL commands: /team switch teams Β· /ingest <path> upload & embed Β· /status document states Β· /retry [force] re-process failed (or all) docs Β· /llms Β· /function feedback Β· /help
Safe deletion from the CLI:
# Target one document by ID or its exact stored path
maru remove document <document-id-or-path> --team-id <team-id>
# Delete a folder and its complete subtree
maru remove group <group-id> --team-id <team-id>
# Add --force / -f to skip the confirmation promptThese commands remove relational records, vector embeddings, and stored upload files together. In-flight ingests are marked for cooperative deletion and are finalized by the ingest worker.
- Interactive docs:
http://localhost:8000/docs(Swagger) - Guides for frontend integration:
docs/ingest-api.mdβ upload / status / check / retry / delete (incl. folder-level operations and the queue processing model)docs/teams-api.mdβ teams, members, invitations
GET /configreturns client bootstrap data β includingsupported_extensions, computed from the current parser configuration (Korean formats appear only when KorDoc is enabled)
MARU was presented at π Open Source Summit Korea 2025, Making RAG Chatbots Enterprise-Ready: Group Permissions and Pluggable Backends.
Distributed under the MIT License
We welcome contributions from the community! Feel free to:
- Open an issue for bugs or suggestions
- Submit a pull request for improvements