Skip to content

Latest commit

Β 

History

2 Commits

Folders and files

NameName
Last commit message
Last commit date
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 

Repository files navigation

🧠 Wisdom RAG Engine

A minimalist, RAG engine powered by LangChain, ChromaDB, FastEmbed, and Ollama.

Python 3.10+ License: MIT FastEmbed Ollama


⚑ Overview

Wisdom RAG is a fully local, privacy-first question-answering pipeline designed for structured knowledge bases. It ingests markdown documents with YAML frontmatter, creates hierarchical semantic chunks, embeds them locally with ultra-fast ONNX multilingual transformers, and streams responses through a terminal UI with typewriter animations and source citations.

 β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
 β”‚  ~/knowledge_base ❯ RAG_ENGINE_v1.0      β”‚
 β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜

πŸ—οΈ Architecture Overview

flowchart TD
    subgraph Ingestion ["1. Ingestion Pipeline (`src/ingest.py`)"]
        A["Raw Markdown Files\n(`data/raw/*.md`)"] --> B["YAML Frontmatter Parser\n& Comment Cleaner"]
        B --> C["Markdown Header Splitter\n(`#`, `##`, `###`)"]
        C --> D["Recursive Character Splitter\n(`chunk_size=450, overlap=50`)"]
        D --> E["FastEmbed Model (ONNX)\n(`paraphrase-multilingual-MiniLM-L12-v2`)"]
        E --> F[("ChromaDB Vector Store\n(`chroma_db/`)")]
    end

    subgraph Retrieval ["2. Retrieval & Generation (`src/query.py`)"]
        G["User Query (CLI Prompt)"] --> H["Vector Similarity Search\n(`top_k=3`)"]
        F -.-> H
        H --> I["Context Synthesizer\n& Prompt Template"]
        I --> J["Ollama Local LLM Daemon\n(`mistral:latest` via `localhost:11434`)"]
        J --> K["Streaming Response + Sources Referenced"]
    end
Loading

Key Architectural Highlights

  • 100% Offline & Private: Zero external API calls. Everything runs locally on your machine.
  • Fast Multilingual Embeddings: Uses FastEmbed with optimized ONNX runtime weights (out-of-the-box support for multilingual semantics).
  • Metadata Preservation: Extracts YAML frontmatter attributes (e.g. uploader, title, url, tags) and tags every chunk so the AI can attribute sources precisely.
  • Resilient Health Checking: Verifies Ollama connectivity and model availability before prompting.
  • Minimalist Dark-Mode Terminal Aesthetic: Cyan/Green ANSI color scheme with smooth token streaming and interactive session management.

πŸ“‹ Prerequisites

  • Operating System: Linux, macOS, or Windows (WSL / PowerShell)
  • Python: 3.10 or higher
  • Ollama: Installed and running on http://localhost:11434
  • RAM: Minimum 8GB recommended for running 7B models (mistral)

πŸš€ 1-Click Installation

Clone the repository and run the automated installer for your operating system:

Linux / macOS

git clone https://github.com/uckix/RAG-engine.git
cd rag-engine
chmod +x install.sh
./install.sh

Windows (PowerShell)

git clone https://github.com/uckix/RAG-engine.git
cd rag-engine
Set-ExecutionPolicy -ExecutionPolicy RemoteSigned -Scope CurrentUser
.\install.ps1

What the installer automates:

  1. Validates Python 3.10+ installation.
  2. Creates and activates an isolated virtual environment (.venv).
  3. Installs all required Python dependencies from requirements.txt.
  4. Checks for Ollama and installs/prompts if missing.
  5. Pulls the mistral LLM model and pre-caches the FastEmbed model.
  6. Generates .env from .env.example and configures directory scaffolds (data/raw, data/processed).

βš™οΈ Configuration (.env)

Copy .env.example to .env to configure paths and parameters:

# ==============================================================================
# RAG PIPELINE CONFIGURATION
# ==============================================================================

# Data Paths
RAW_DATA_PATH=./data/raw
DB_PATH=./chroma_db

# Ollama LLM Service Configuration
OLLAMA_BASE_URL=http://localhost:11434
OLLAMA_MODEL=mistral

# Embedding Model Configuration (FastEmbed ONNX)
EMBEDDING_MODEL_NAME=sentence-transformers/paraphrase-multilingual-MiniLM-L12-v2

# Ingestion & Splitting Hyperparameters
CHUNK_SIZE=450
CHUNK_OVERLAP=50

# Retrieval Hyperparameters
TOP_K_RESULTS=3

πŸ“– Usage Guide

1. Format Your Markdown Documents

Place .md files into the data/raw/ directory (or the path defined in RAW_DATA_PATH).

Documents can include optional YAML frontmatter at the top. Any metadata keys (like title, uploader, url) are indexed and injected into citations:

---
title: "Agency Agents GitHub Repository"
uploader: "Dmitry"
url: "https://github.com/example/repo"
tags: "ai agents, prompt engineering, roles"
---

# Agency Agents GitHub Repository for Role-Based Prompts

## Overview
A repository collecting role-based system prompts for software teams.

## Key Features
- Frontend engineer role prompts
- Architect design prompts
- Data science pipelines

2. Run the Ingestion Pipeline

Execute the ingestion script to process, split, embed, and store your knowledge base:

# Using active venv:
python src/ingest.py

# Or pass custom paths:
python src/ingest.py --source ./data/raw --dest ./chroma_db --chunk-size 450

The CLI displays real-time progress bars via tqdm for parsing, splitting, and vector indexing.

3. Launch the Interactive Chat CLI

Launch the terminal chat loop:

python src/query.py

Terminal Commands:

  • Type your question and hit Enter to stream answers.
  • clear: Wipes the terminal screen and redraws the header.
  • help: Displays available commands.
  • exit or quit: Gracefully exits the session.

πŸ“‚ Repository Structure

rag-engine/
β”œβ”€β”€ .env.example              # Template configuration for environment variables
β”œβ”€β”€ .gitignore                # Production ignore rules (caches, DBs, raw data)
β”œβ”€β”€ requirements.txt          # Pinned & versioned Python dependencies
β”œβ”€β”€ install.sh                # 1-click installer for Linux/macOS
β”œβ”€β”€ install.ps1               # 1-click installer for Windows PowerShell
β”œβ”€β”€ README.md                 # Complete system documentation
β”œβ”€β”€ data/
β”‚   β”œβ”€β”€ raw/                  # Source Markdown files (.gitkeep)
β”‚   └── processed/            # Processed artifacts cache (.gitkeep)
β”œβ”€β”€ chroma_db/                # Local persistent ChromaDB vector storage
└── src/
    β”œβ”€β”€ __init__.py           # Package marker
    β”œβ”€β”€ config.py             # Typed configuration loader (.env parser)
    β”œβ”€β”€ db.py                 # Vector database & embedding connection module
    β”œβ”€β”€ ingest.py             # Ingestion, YAML parsing & chunking pipeline
    └── query.py              # Dark-mode terminal UI & streaming RAG loop

πŸ›‘οΈ Error Handling & Troubleshooting

Issue Cause Solution
Cannot connect to Ollama daemon Ollama service is stopped Run ollama serve or systemctl start ollama
Model 'mistral' not detected Model not downloaded Run ollama pull mistral
Vector database is unpopulated No embeddings indexed Place .md files in data/raw/ and run python src/ingest.py
Python 3.10+ required Outdated Python runtime Upgrade Python using your package manager or pyenv

πŸ“„ License

This project is licensed under the MIT License β€” see the LICENSE file for details.

About

Wisdom RAG is a fully local, privacy-first question-answering pipeline designed for structured knowledge bases.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages