Ask questions about your own documents, answered by Software Tailor AI Server with citations — entirely on your own machines. Nothing is uploaded anywhere.
This is retrieval-augmented generation (RAG) in its honest, minimal form: three files, no dependencies, no vector database, no framework. You can read the whole thing in a sitting and see exactly what RAG actually does.
$ python index.py ./my-notes
indexed my-notes/handbook.md (12 chunks)
indexed my-notes/pricing.md (5 chunks)
Done. 2 file(s) indexed, 0 unchanged. 17 chunks in docqa.sqlite3.
$ python ask.py "What does Pro unlock compared to the free tier?"
Sources:
[1] my-notes/pricing.md (part 1, similarity 0.627)
[2] my-notes/handbook.md (part 4, similarity 0.579)
Answer:
According to [1], Pro unlocks network serving (LAN / all interfaces)…index.pywalks a folder, splits each text file into overlapping chunks, and asks AI Server for an embedding of each one (POST /v1/embeddings) — a vector capturing its meaning. Vectors go into a local SQLite file.ask.pyembeds your question with the same model, finds the nearest chunks by cosine similarity, and passes them to the model as context (POST /v1/chat/completions, streamed), instructing it to answer only from that context and cite its sources.
Two different model families are involved, which is the thing most RAG tutorials gloss over: an embedding model turns text into vectors, and a chat model writes the answer. A chat model cannot do the first job.
- AI Server running, with an API key from AI Server → API keys.
- An embedding model installed.
nomic-embed-textis the usual choice — add it from AI Server's Models page. - A chat model (any — it's discovered automatically).
- Python 3.8+. No
pip install.
export AISERVER_BASE_URL="http://192.168.1.42:11436/v1" # from AI Server's Server page
export AISERVER_API_KEY="ai-suite_..." # from AI Server -> API keys
python index.py ./my-notes
python ask.py "your question"PowerShell:
$env:AISERVER_BASE_URL = "http://192.168.1.42:11436/v1"
$env:AISERVER_API_KEY = "ai-suite_..."
python index.py .\my-notes
python ask.py "your question"| Variable | Default | Meaning |
|---|---|---|
AISERVER_BASE_URL |
http://localhost:11436/v1 |
Base URL including /v1 |
AISERVER_API_KEY |
— | Key from AI Server → API keys (required) |
AISERVER_EMBED_MODEL |
enginea/nomic-embed-text/latest |
Embedding model id |
AISERVER_MODEL |
auto | Chat model; discovered if unset |
DOCQA_DB |
docqa.sqlite3 |
Where the index lives |
DOCQA_TOP_K |
5 |
How many chunks to feed the model |
Re-running index.py is incremental — files whose contents haven't changed are skipped, so you can
point it at a folder you keep editing.
Indexes plain-text formats (.md, .txt, .rst, source code, .json, .csv, …). PDFs and Office
files would need a third-party parser, which this sample deliberately avoids.
- Batch your embeddings.
index.pysends 32 chunks per request. One request per chunk is the single biggest performance mistake in a naive indexer. - Send
inputas an array. AI Server requires it; OpenAI accepts both. An array works everywhere. - Embedding models aren't listed by
/v1/models. Chat models are discoverable, embedding models aren't — so name yours viaAISERVER_EMBED_MODELrather than trying to auto-detect it. - Chunk with overlap (here 1200 chars, 200 overlapping), breaking at paragraph boundaries so a chunk rarely ends mid-sentence and a split idea still matches.
- Query and index must share an embedding model. Vectors from different models aren't comparable;
ask.pydetects a size mismatch and tells you to re-index rather than returning nonsense. - A linear scan is fine. Cosine over a few thousand chunks is milliseconds. Reach for a vector index when you can measure that this hurts — not before.
"The server could not produce embeddings … (500)". AISERVER_EMBED_MODEL isn't an embedding model,
or none is installed. Add one in AI Server → Models.
"No documents indexed yet." Run index.py <folder> first.
"Indexed vectors have a different size to the query vector." The embedding model changed since you
indexed. Re-index, or delete docqa.sqlite3 and start over.
"Could not reach …". Check AISERVER_BASE_URL ends in /v1. Serving from another machine needs
that server bound to your network and allowed through its firewall — Network diagnostics → Test
connectivity on that machine checks all of it.
- AI Server API & clients
- ai-server-quickstarts — the API in five languages
- ai-server-chat-web — a streaming chat UI
MIT — copy any of this into your own project.