Portable, modular, local AI orchestration platform that runs from a USB drive.
Monad (codename: Ultron) is not a new language model. It is an orchestration operating system that coordinates multiple open-source LLMs — routing, reasoning, coding, and creative work across specialized models — all running locally on your machine, from a USB drive.
Plug the USB into any compatible Windows PC, double-click monad.bat, and you have a full local AI workstation.
- 🔌 Fully portable — Python, dependencies, models, and code all live on the USB
- 🧩 Modular architecture — Every subsystem is swappable via dependency injection
- 🤖 Multi-model brain — LongCat 2 (reasoning) + GLM-5 (code) + Llama 2 (creative)
- 🛠️ Tool framework — Filesystem, Python sandbox, Git, terminal, browser, PDFs
- 🧠 Local memory — SQLite + ChromaDB vector store, no cloud
- 🔒 Approval-gated — All impactful actions require your explicit approval
- 🔧 Plugin-based — Extend without touching core (JCode, ZeroLang, and more)
- 📴 Offline-first — Once installed, needs zero internet
- Laptop: ASUS G615JHR-S5005WS (or any Windows 11 PC with an NVIDIA GPU)
- GPU: RTX 5070 Laptop (or any CUDA 12.x GPU with ≥ 8 GB VRAM)
- USB: 128 GB minimum, USB 3.2 recommended for speed
- Plug your 128 GB USB drive into your desktop.
- Download this repo and extract it anywhere on your desktop.
- Double-click
installer/install_to_usb.bat. - The installer will:
- Detect your USB drive (asks you to confirm)
- Copy the Monad codebase to the USB
- Download portable Python 3.12 (~30 MB) to the USB
- Create a virtual environment on the USB
- Install all Python dependencies on the USB
- Download the 3 GGUF models (~15 GB) to
USB:/models/ - Create the launcher
monad.batat the USB root
- Eject and plug the USB into your laptop.
- Double-click
monad.baton the USB. Done. 🎉
- Copy the entire
Monad-Ultron/folder to your USB drive. - Follow
docs/INSTALL.mdfor manual Python + model setup.
USER
│
▼
┌──────────────────┐
│ CLI / Dashboard │
└────────┬─────────┘
▼
Application Manager
│
┌───────┼───────┐
▼ ▼ ▼
Config DI Plugins
└───────┼───────┘
▼
Environment Manager
▼
Resource Manager
▼
Prompt Management
▼
Router / Intent Engine
│
┌───────┼───────┐
▼ ▼ ▼
LongCat GLM-5 Llama-2
(reason) (code) (creative)
└───────┼───────┘
▼
Response Synthesizer
▼
Policy Gate
▼
Memory & Retrieval
▼
Tool Framework
┌────┬────┬────┬────┐
▼ ▼ ▼ ▼ ▼
FS Python Browser Term Git
│
┌───┴───┐
▼ ▼
JCode ZeroLang
│
▼
Final Response
Full spec: see docs/ARCHITECTURE.md.
Monad-Ultron/
├── monad/ # Core Python package (all real code, no stubs)
│ ├── core/ # App manager, DI container, logger, env, resources
│ ├── config/ # YAML configuration system
│ ├── models/ # Model manager, loader, runtime, registry
│ ├── inference/ # LLM providers (llama.cpp + speculative decoding)
│ ├── prompts/ # Prompt templates & context builder
│ ├── router/ # Intent classifier
│ ├── chat/ # Single-model chat engine
│ ├── orchestration/ # ✅ Multi-model + fusion + cache + adaptive + streaming
│ ├── cognition/ # ✅ 82 organs + memory + reasoning + executive + self-model
│ ├── evolution/ # ✅ Self-improvement (propose/test/approve/rollback)
│ ├── memory/ # ✅ Real SQLite + ChromaDB + RRF hybrid retrieval
│ ├── tools/ # ✅ Filesystem, Python sandbox, terminal, HTTP
│ ├── policy/ # ✅ Real approval gate (5 modes + SQLite audit)
│ ├── scheduler/ # ✅ Thread-based periodic + one-shot jobs
│ ├── api/ # ✅ FastAPI + HTML dashboard + streaming endpoints
│ ├── plugins/ # Plugin manager + example plugins
│ ├── ui/ # CLI (Typer + Rich)
│ └── utils/ # Shared utilities
├── webapp/ # ✅ Next.js 15 landing + chat UI
├── installer/ # Windows installer scripts
├── launcher/ # USB launcher (.bat files)
├── docs/ # ARCHITECTURE, INSTALL, USAGE, BUILDS
├── tests/ # Unit tests
├── scripts/ # Dev/maintenance scripts
├── config.yaml # Main runtime config
├── models.yaml # Model download manifest
├── requirements.txt
├── pyproject.toml
├── LICENSE # MIT
├── run.py # Main entry point
└── README.md # (this file)
Monad is being built in ~120 small, testable milestones.
| Phase | Milestones | Status |
|---|---|---|
| Foundation (project setup, config, logging, CLI) | #001–#010 | ✅ Complete |
| Model framework & single-model chat | #011–#013 | ✅ Complete |
| Routing, inference, prompts | #014–#016 | ✅ Complete |
| Self-improvement framework (self-update, self-extend, self-debug) | #017a | ✅ Complete |
| Multi-model orchestration (5 strategies + confidence scoring) | #017 | ✅ Complete |
| llama.cpp perf upgrades (speculative decoding, KV quant, flash attn) | #017b | ✅ Complete |
| Cognitive architecture (9 layers, 82 canonical organs, Cognee, MCP) | Phases 1-6 | ✅ Complete |
| Real memory layer (SQLite + ChromaDB + RRF hybrid retrieval) | #026 | ✅ Complete |
| Tool framework (Filesystem, Python sandbox, Terminal, HTTP) | #036–#039 | ✅ Complete |
| Real policy gate (allow/deny/prompt + SQLite audit) | #056 | ✅ Complete |
Cognition→Orchestrator wiring (monad ask --cognition) |
#017f | ✅ Complete |
FastAPI + HTML dashboard (monad serve) |
#059 | ✅ Complete |
| Background scheduler (periodic + one-shot jobs) | #070 | ✅ Complete |
| LLM Fusion (all models → ONE unified answer via Chain / EnsembleTokens / Logits) | #080 | ✅ Complete |
| Web app (Next.js 15 landing + chat UI + natural-language commands) | #090 | ✅ Complete |
| Streaming (Server-Sent Events + typewriter effect) | #018 | ✅ Complete |
| Adaptive routing (Thompson sampling over strategies, learns from usage) | #020 | ✅ Complete |
| Response cache (LRU + SQLite persistent, 2-tier) | #024 | ✅ Complete |
| One-click USB installer (wizard + profiles + progress) | #100 | ✅ Complete |
| Remaining polish / tutorials / release prep | #101–#120 | 🟡 Partial |
Detailed tracker: docs/BUILDS.md.
All the previously-stubbed subsystems are now real code: memory · tools · policy · scheduler · api · streaming · caching · adaptive routing. The only remaining items are documentation polish, extra tests, and release prep — none of which are stubs, just extras.
# Clone
git clone https://github.com/YOUR_USERNAME/Monad-Ultron.git
cd Monad-Ultron
# Create venv
python -m venv .venv
.venv\Scripts\activate # Windows
# or: source .venv/bin/activate # Linux/Mac
# Install
pip install -e .
# Run
python run.py
# or
monad --helpMonad is built one small milestone at a time. See docs/BUILDS.md for the build queue. Each PR should implement exactly one milestone.
MIT — see LICENSE.
- Read the docs in
docs/ - Run
python run.py doctorfor diagnostics - Open an issue on GitHub
⚠️ Note on GitHub username: All URLs in this repo currently useYOUR_USERNAMEas a placeholder. Before pushing to GitHub, run:# Windows PowerShell Get-ChildItem -Recurse -File | ForEach-Object { (Get-Content $_.FullName -Raw) -replace 'YOUR_USERNAME', 'your-actual-username' | Set-Content $_.FullName }Then push with
git remote add origin https://github.com/your-actual-username/Monad-Ultron.git && git push -u origin main.