A multimodal AI/ML teaching assistant that retrieves and explains technical diagrams using a Plan-and-Execute agent architecture. Features SigLIP 2 for semantic image retrieval and Gemini 3 Flash for vision-augmented explanations.
- π€ LangGraph Agent: Plan-and-Execute architecture with CoT planning and ReAct execution
- πΌοΈ Multimodal Vision: Agent sees retrieved diagrams and provides contextual descriptions
- π SigLIP 2 Retrieval: State-of-the-art bi-encoder (76% top-1 accuracy on ML diagrams)
- π MCP Tool Server: Diagram retrieval runs as a separate MCP subprocess (stdio); the agent calls it like any LangChain tool
- π LangSmith Eval: Traced retrieval (SigLIP vs CLIP) and end-to-end agent eval against a versioned dataset
- π¬ Conversation Memory: Redis-backed checkpointing with 24h TTL
- π¨ 92 Diagrams: Curated from Jay Alammar's Illustrated ML posts
- β‘ SSE Streaming: Real-time token streaming with thinking/planning visibility
- π₯οΈ React Frontend: Chat UI with inline diagrams and collapsible teaching plans
# Clone repo with big_vision dependency
git clone https://github.com/jessecui/mentor-ml.git
cd mentor-ml
git clone --quiet --branch=main --depth=1 https://github.com/google-research/big_vision big_vision_repo
# Create virtual environment
python -m venv venv
source venv/bin/activate
pip install -r requirements.txt
# Download Gemma tokenizer (used by SigLIP 2)
curl -L -o model/gemma_tokenizer.model https://storage.googleapis.com/big_vision/paligemma_tokenizer.model
# Set up environment variables
cp .env.example .env # Then edit with your keys
# Start Redis (required for conversation memory)
brew services start redis # or: docker run -p 6379:6379 redis
# Install frontend dependencies
cd frontend && npm install && cd ..# Terminal 1: Backend
uvicorn main:app --reload --port 8080
# Terminal 2: Frontend (with hot reload)
cd frontend && npm run dev# Build frontend
cd frontend && npm run build && cd ..
# Run server (serves API + frontend)
uvicorn main:app --port 8080# Stream chat responses with SSE
curl -N -X POST http://localhost:8080/chat/stream \
-H "Content-Type: application/json" \
-d '{"message": "How do transformers work?", "thread_id": "user-123"}'SSE Events:
thinking- Planning tokens (JSON teaching plan)diagram- Retrieved diagram metadatatoken- Response text tokensplan- Parsed teaching plan objectdone- Final state with referenced diagrams
# Chat with the agent (blocking)
curl -X POST http://localhost:8080/chat \
-H "Content-Type: application/json" \
-d '{"message": "How do transformers work?", "thread_id": "user-123"}'{
"response": "Transformers use self-attention to process sequences...",
"diagrams": [
{
"id": "diagram_042",
"score": 0.00045,
"query": "transformer self-attention mechanism",
"description": "Diagram showing Q, K, V matrices...",
"vision_description": "This diagram illustrates the scaled dot-product attention...",
"vision_latency_s": 2.5,
"post_url": "https://jalammar.github.io/illustrated-transformer/"
}
],
"plan": {
"topic": "Transformers",
"steps": ["Explain self-attention", "Describe Q, K, V matrices", "..."],
"diagrams_needed": ["attention mechanism", "encoder-decoder"]
}
}The scorer uses SigLIP 2 (So400m/14 @ 384px), Google's state-of-the-art contrastive vision-language model optimized for retrieval.
from model.scorer import SigLIPScorer
scorer = SigLIPScorer()
# Score single image-query pair
score = scorer.score("diagram.png", "transformer attention mechanism")
# Batch scoring (efficient - encodes query once)
scores = scorer.score_batch(["img1.png", "img2.png"], "self-attention layer")| Component | Specification |
|---|---|
| Model | SigLIP 2 So400m/14 |
| Image Size | 384Γ384 |
| Parameters | ~400M |
| Checkpoint | ~1.5GB (auto-downloaded) |
| Tokenizer | Gemma (256k vocab) |
| Framework | JAX/Flax (big_vision) |
The checkpoint downloads automatically from Google Cloud Storage on first run:
https://storage.googleapis.com/big_vision/siglip2/siglip2_so400m14_384.npz
Evaluate SigLIP vs CLIP on technical ML diagram retrieval using 92 diagrams from Jay Alammar's Illustrated series.
| Model | Top-1 Accuracy |
|---|---|
| SigLIP 2 | 76.1% (70/92) |
| CLIP ViT-L/14 | 46.7% (43/92) |
+29.3 percentage points improvement (+62.8% relative)
# 1. Scrape diagrams from Jay Alammar's blog
python benchmark/scripts/scrape_ai_ml_diagrams.py
# 2. Generate queries (requires GEMINI_API_KEY in .env)
python benchmark/scripts/generate_queries.py
# 3. Evaluate SigLIP vs CLIP (offline, self-contained)
python benchmark/scripts/evaluate.pyThe retrieval benchmark is also runnable via LangSmith for traced runs, a versioned dataset, and a side-by-side comparison UI. End-to-end agent quality (tool-call correctness + LLM-judge on explanation quality) ships as a separate script.
# Requires LANGSMITH_API_KEY in .env
# One-time: upload benchmark queries as a LangSmith dataset
python benchmark/scripts/upload_to_langsmith.py
# Retrieval eval (SigLIP + CLIP as evaluate() targets)
python benchmark/scripts/langsmith_evaluate_retrieval.py
# Full agent eval (LangGraph end-to-end; ~30 min, ~$1-3 in Gemini calls)
python benchmark/scripts/langsmith_evaluate_agent.py --limit 5 # smoke test
python benchmark/scripts/langsmith_evaluate_agent.py # full runThe benchmark uses diagrams from the top 5 Illustrated posts:
- The Illustrated Transformer
- The Illustrated BERT
- The Illustrated GPT-2
- The Illustrated Word2vec
- The Illustrated Stable Diffusion
benchmark/
βββ corpus/
β βββ images/diagrams/ # Downloaded ML diagrams
β βββ metadata/
β βββ corpus.json # Scraped metadata
β βββ corpus_with_queries.json
βββ queries/
β βββ benchmark_queries.json # Query-image ground truth
βββ results/
β βββ siglip_evaluation_results.json
β βββ siglip_evaluation_summary.txt
βββ scripts/
βββ scrape_ai_ml_diagrams.py
βββ generate_queries.py
βββ evaluate.py # Offline SigLIP-vs-CLIP baseline
βββ upload_to_langsmith.py # Upload dataset to LangSmith
βββ langsmith_evaluate_retrieval.py # Retrieval eval via LangSmith
βββ langsmith_evaluate_agent.py # End-to-end agent eval via LangSmith
model/
βββ scorer.py # SigLIP scorer
βββ gemma_tokenizer.model # Tokenizer (~4MB)
βββ siglip2_so400m14_384.npz # Checkpoint (~1.5GB, gitignored)
agent/
βββ graph.py # LangGraph agent (Plan-and-Execute)
βββ tools.py # MCP client: persistent stdio session, wrapped tools
mcp_server/
βββ diagram_server.py # MCP server: SigLIP scorer + retrieve_diagram tool
Create a .env file in the project root:
# Required
GOOGLE_API_KEY=your_gemini_api_key
# Optional
REDIS_URL=redis://localhost:6379 # Default
ENABLE_VISION=true # Enable vision review (default: true)
# Optional - LangSmith tracing & eval
LANGSMITH_API_KEY=your_langsmith_key
LANGSMITH_TRACING=true # Auto-trace live agent runs
LANGSMITH_PROJECT=mentorml-prod # Default project: "default"User Query β Plan Node (CoT) β Execute Node (ReAct) β Tools β Response
β
retrieve_diagram βββ MCP stdio βββ diagram_server.py
(SigLIP scorer +
Gemini vision review)
frontend/src/
βββ components/
β βββ Chat.tsx # Main container
β βββ ChatInput.tsx # Input with send/stop/clear
β βββ DiagramCard.tsx # Diagram display with source link
β βββ Message.tsx # Message bubble + ThinkingSection
β βββ MessageList.tsx # Message list + empty state
βββ hooks/
β βββ useStreamChat.ts # SSE streaming hook
βββ types.ts # TypeScript interfaces
- Python 3.10+
- Node.js 18+ (for frontend)
- Redis (for conversation memory)
- ~4GB disk space (SigLIP checkpoint + diagrams)
- Gemini API key