A minimalist, RAG engine powered by LangChain, ChromaDB, FastEmbed, and Ollama.
Wisdom RAG is a fully local, privacy-first question-answering pipeline designed for structured knowledge bases. It ingests markdown documents with YAML frontmatter, creates hierarchical semantic chunks, embeds them locally with ultra-fast ONNX multilingual transformers, and streams responses through a terminal UI with typewriter animations and source citations.
ββββββββββββββββββββββββββββββββββββββββββββ
β ~/knowledge_base β― RAG_ENGINE_v1.0 β
ββββββββββββββββββββββββββββββββββββββββββββ
flowchart TD
subgraph Ingestion ["1. Ingestion Pipeline (`src/ingest.py`)"]
A["Raw Markdown Files\n(`data/raw/*.md`)"] --> B["YAML Frontmatter Parser\n& Comment Cleaner"]
B --> C["Markdown Header Splitter\n(`#`, `##`, `###`)"]
C --> D["Recursive Character Splitter\n(`chunk_size=450, overlap=50`)"]
D --> E["FastEmbed Model (ONNX)\n(`paraphrase-multilingual-MiniLM-L12-v2`)"]
E --> F[("ChromaDB Vector Store\n(`chroma_db/`)")]
end
subgraph Retrieval ["2. Retrieval & Generation (`src/query.py`)"]
G["User Query (CLI Prompt)"] --> H["Vector Similarity Search\n(`top_k=3`)"]
F -.-> H
H --> I["Context Synthesizer\n& Prompt Template"]
I --> J["Ollama Local LLM Daemon\n(`mistral:latest` via `localhost:11434`)"]
J --> K["Streaming Response + Sources Referenced"]
end
- 100% Offline & Private: Zero external API calls. Everything runs locally on your machine.
- Fast Multilingual Embeddings: Uses
FastEmbedwith optimized ONNX runtime weights (out-of-the-box support for multilingual semantics). - Metadata Preservation: Extracts YAML frontmatter attributes (e.g.
uploader,title,url,tags) and tags every chunk so the AI can attribute sources precisely. - Resilient Health Checking: Verifies Ollama connectivity and model availability before prompting.
- Minimalist Dark-Mode Terminal Aesthetic: Cyan/Green ANSI color scheme with smooth token streaming and interactive session management.
- Operating System: Linux, macOS, or Windows (WSL / PowerShell)
- Python:
3.10or higher - Ollama: Installed and running on
http://localhost:11434 - RAM: Minimum 8GB recommended for running 7B models (
mistral)
Clone the repository and run the automated installer for your operating system:
git clone https://github.com/uckix/RAG-engine.git
cd rag-engine
chmod +x install.sh
./install.shgit clone https://github.com/uckix/RAG-engine.git
cd rag-engine
Set-ExecutionPolicy -ExecutionPolicy RemoteSigned -Scope CurrentUser
.\install.ps1- Validates Python 3.10+ installation.
- Creates and activates an isolated virtual environment (
.venv). - Installs all required Python dependencies from
requirements.txt. - Checks for Ollama and installs/prompts if missing.
- Pulls the
mistralLLM model and pre-caches the FastEmbed model. - Generates
.envfrom.env.exampleand configures directory scaffolds (data/raw,data/processed).
Copy .env.example to .env to configure paths and parameters:
# ==============================================================================
# RAG PIPELINE CONFIGURATION
# ==============================================================================
# Data Paths
RAW_DATA_PATH=./data/raw
DB_PATH=./chroma_db
# Ollama LLM Service Configuration
OLLAMA_BASE_URL=http://localhost:11434
OLLAMA_MODEL=mistral
# Embedding Model Configuration (FastEmbed ONNX)
EMBEDDING_MODEL_NAME=sentence-transformers/paraphrase-multilingual-MiniLM-L12-v2
# Ingestion & Splitting Hyperparameters
CHUNK_SIZE=450
CHUNK_OVERLAP=50
# Retrieval Hyperparameters
TOP_K_RESULTS=3Place .md files into the data/raw/ directory (or the path defined in RAW_DATA_PATH).
Documents can include optional YAML frontmatter at the top. Any metadata keys (like title, uploader, url) are indexed and injected into citations:
---
title: "Agency Agents GitHub Repository"
uploader: "Dmitry"
url: "https://github.com/example/repo"
tags: "ai agents, prompt engineering, roles"
---
# Agency Agents GitHub Repository for Role-Based Prompts
## Overview
A repository collecting role-based system prompts for software teams.
## Key Features
- Frontend engineer role prompts
- Architect design prompts
- Data science pipelinesExecute the ingestion script to process, split, embed, and store your knowledge base:
# Using active venv:
python src/ingest.py
# Or pass custom paths:
python src/ingest.py --source ./data/raw --dest ./chroma_db --chunk-size 450The CLI displays real-time progress bars via tqdm for parsing, splitting, and vector indexing.
Launch the terminal chat loop:
python src/query.py- Type your question and hit
Enterto stream answers. clear: Wipes the terminal screen and redraws the header.help: Displays available commands.exitorquit: Gracefully exits the session.
rag-engine/
βββ .env.example # Template configuration for environment variables
βββ .gitignore # Production ignore rules (caches, DBs, raw data)
βββ requirements.txt # Pinned & versioned Python dependencies
βββ install.sh # 1-click installer for Linux/macOS
βββ install.ps1 # 1-click installer for Windows PowerShell
βββ README.md # Complete system documentation
βββ data/
β βββ raw/ # Source Markdown files (.gitkeep)
β βββ processed/ # Processed artifacts cache (.gitkeep)
βββ chroma_db/ # Local persistent ChromaDB vector storage
βββ src/
βββ __init__.py # Package marker
βββ config.py # Typed configuration loader (.env parser)
βββ db.py # Vector database & embedding connection module
βββ ingest.py # Ingestion, YAML parsing & chunking pipeline
βββ query.py # Dark-mode terminal UI & streaming RAG loop
| Issue | Cause | Solution |
|---|---|---|
Cannot connect to Ollama daemon |
Ollama service is stopped | Run ollama serve or systemctl start ollama |
Model 'mistral' not detected |
Model not downloaded | Run ollama pull mistral |
Vector database is unpopulated |
No embeddings indexed | Place .md files in data/raw/ and run python src/ingest.py |
Python 3.10+ required |
Outdated Python runtime | Upgrade Python using your package manager or pyenv |
This project is licensed under the MIT License β see the LICENSE file for details.