This directory is the active Agentic RAG project.
Use this project directory as the working directory for development, testing, indexing, and runtime commands.
Requires Python 3.11+. Install with:
python -m pip install -e .The main entrypoint is graph_rag.py.
- graph_rag.py: CLI entrypoint
- graph_rag_app: application code
- agent: local retrieval index files
- runtime: generated logs and checkpoints
- skills: default harness skills loaded through
load_skill - reports: generated academic research reports
Run these commands from this project directory:
cd D:\Code\agent_learning\agent_ragCheck your Python version:
python --versionStart, inspect, and stop the service with the project scripts:
.\scripts\start.ps1
.\scripts\status.ps1
.\scripts\stop.ps1Build an index:
python graph_rag.py index build --kb-path "<knowledge-base-path>" --output-dir .\agentInspect an existing index:
python graph_rag.py index inspect --index-dir .\agentRun direct retrieval without the agent:
python graph_rag.py query run --index-dir .\agent --question "openai agent"Run the agent with local retrieval and web fallback:
python graph_rag.py ask --index-dir .\agent --question "What changed recently about OpenAI agents?"Generate an academic Markdown research report from a completed agent session:
python graph_rag.py report --session <session-id> --index-dir .\agent --output-dir reportsThe report includes Abstract, Introduction, Methods, Findings, Discussion,
Conclusion, and References sections. Sources from local retrieval, web fetch,
and scholar search are automatically extracted and cited as [S1], [S2], etc.
A citation audit at the end of the report flags missing references and
potentially unsupported claims. Use --model to override the chat model for
report generation.
Configure the chat model in .env with OpenAI-compatible provider settings:
# Any OpenAI-compatible endpoint — e.g. GPT-5.5, DeepSeek, or DashScope
RAG_MODEL_API_KEY=<provider-api-key>
RAG_MODEL_API_BASE=http://127.0.0.1:58410/v1
RAG_MODEL_NAME=gpt-5.5DASHSCOPE_API_KEY is still supported as a backwards-compatible fallback, but
new model/provider switches should use the RAG_MODEL_* variables.
Run the browser question-answering UI:
python graph_rag.py ui --index-dir .\agentThe UI is served by the FastAPI backend and listens at http://127.0.0.1:8765
by default. Use --host and --port to change the bind address.
Run the FastAPI backend service explicitly:
python graph_rag.py serve --index-dir .\agent --host 127.0.0.1 --port 8765Configure web search providers in .env:
RAG_WEB_SEARCH_PROVIDER=searxng
RAG_SEARXNG_URL=http://127.0.0.1:8080
RAG_SEARXNG_ENGINES=google,bing,duckduckgo
RAG_SEARXNG_CATEGORIES=general
RAG_SEARXNG_LANGUAGE=zh-CNSearXNG must enable JSON output for /search?q=...&format=json to work. In a
self-hosted SearXNG instance, make sure settings.yml includes json in
search.formats. A local Docker setup is included:
docker compose -f docker-compose.searxng.yml up -dThen set RAG_WEB_SEARCH_PROVIDER=searxng in .env and restart the API. If
SearXNG is unavailable or fails, the backend falls back to DuckDuckGo HTML search
and records provider failures in search debug metadata.
Useful HTTP endpoints:
GET /healthz
GET /api/config
GET /api/status
POST /api/ask
POST /api/retrieve
GET /api/index/inspect?index_dir=agent
POST /api/web/search
POST /api/web/fetch
POST /api/scholar/search
Every HTTP response includes an X-Request-ID header for log correlation. Pass
your own X-Request-ID header to reuse a caller-provided id.
Runtime status does not expose secret values:
Invoke-RestMethod -Uri "http://127.0.0.1:8765/api/status"Errors use a stable envelope:
{
"error": {
"code": "validation_error",
"message": "Request validation failed.",
"details": {},
"request_id": "..."
}
}Ask through the API:
Invoke-RestMethod `
-Method Post `
-Uri "http://127.0.0.1:8765/api/ask" `
-ContentType "application/json" `
-Body '{"question":"What does the knowledge base say about agents?","index_dir":"agent","session_id":"demo"}'The ask response includes answer traceability. Local knowledge-base hits are
returned as sources entries with source_type: "local", source_path,
optional section metadata, score, and snippet text. External web evidence is
returned as source_type: "web" entries with title, url, optional search
provider metadata, and snippet text. retrieval_debug.source_count reports the
number of extracted sources.
Run retrieval without invoking the LLM:
Invoke-RestMethod `
-Method Post `
-Uri "http://127.0.0.1:8765/api/retrieve" `
-ContentType "application/json" `
-Body '{"query":"agent tools","index_dir":"agent","top_k":3,"strategy":"hybrid"}'Run Google Scholar search and export the results to Markdown:
python graph_rag.py scholar search --topic "graph rag" --count 5 --save-mdThe agent keeps the existing LangGraph runtime and adds three optional harness layers around it:
- Skill loading: default skills live under
skills/<name>/SKILL.md. The agent sees an inventory of available skills and can callload_skillwhen a skill is relevant to the current research task. - Research plan: the agent can call
research_plan_updateto record the current question, planned steps, gathered evidence, open gaps, and next action. The latest plan is re-injected into context on later model calls. - Structured trace: runtime events are written as JSONL under
runtime/traces. Trace events include context construction, LLM decisions, tool calls, cache hits, tool results, checkpoints, and final answer previews.
Configure these with .env:
RAG_SKILLS_DIR=skills
RAG_STRUCTURED_TRACE_ENABLED=true
RAG_STRUCTURED_TRACE_DIR=runtime/tracesBuilt-in skills:
retrieval-debug: local retrieval, chunking, index, query rewrite, and metadata rerank diagnosis.web-scholar-research: web search, fetch-before-answer, official-source use, and Google Scholar research.agent-evaluation: layered evaluation for retrieval, trajectory, sources, and final answers.grounded-answering: final answer synthesis from local, web, and scholar evidence.
Custom skill format:
skills/retrieval-debug/SKILL.md
---
name: retrieval-debug
description: Diagnose retrieval failures and evidence gaps.
---
Inspect index metadata, compare retrieved source titles, and identify whether
the failure is chunking, query framing, or answer synthesis.The service applies a small tool permission policy before using caller-provided
paths or URLs. By default, index directories must stay under this project
directory, while local/private web fetches remain enabled for local development.
Tighten those boundaries in .env when exposing the API beyond localhost:
RAG_ALLOWED_INDEX_ROOTS=agent;runtime
RAG_ALLOW_LOCAL_WEB_FETCH=false
RAG_ALLOW_PRIVATE_WEB_FETCH=false
RAG_JOB_RUNTIME_DIR=runtime/jobs
RAG_JOB_MAX_LOG_CHARS=12000Submit a long index build job through the API:
Invoke-RestMethod `
-Method Post `
-Uri "http://127.0.0.1:8765/api/jobs/index-build" `
-ContentType "application/json" `
-Body '{"kb_path":"docs","output_dir":"agent/job-index"}'Submit an evaluation job:
Invoke-RestMethod `
-Method Post `
-Uri "http://127.0.0.1:8765/api/jobs/eval-run" `
-ContentType "application/json" `
-Body '{"dataset":"evals/datasets/agent-smoke.jsonl","index_dir":"agent","output_dir":"runtime/evals"}'Check status and read logs:
Invoke-RestMethod -Uri "http://127.0.0.1:8765/api/jobs/<job_id>"
Invoke-RestMethod -Uri "http://127.0.0.1:8765/api/jobs/<job_id>/log"The CLI can inspect persisted job records and logs from the same runtime directory:
python graph_rag.py job status --job-id <job_id>
python graph_rag.py job log --job-id <job_id>Install development-only quality tools:
python -m pip install -r requirements-dev.txtRun the same checks locally that CI runs:
python -m ruff format --check .
python -m ruff check .
python -m pyrightGitHub Actions runs these commands from this project directory.
The existing local evaluation flow remains the source of truth for pass/fail
reports under runtime/evals/*. LangSmith tracing is optional and adds run
observation for ask and eval run without changing the local report format.
Enable tracing by setting:
$env:RAG_LANGSMITH_ENABLED="true"
$env:LANGCHAIN_TRACING_V2="true"
$env:LANGCHAIN_API_KEY="<langsmith-api-key>"
$env:LANGCHAIN_PROJECT="agent-rag"Then run the usual commands:
python graph_rag.py ask --index-dir .\agent --question "What changed recently about OpenAI agents?"
python graph_rag.py eval run --dataset evals\datasets\agent-smoke.jsonl --index-dir .\agentEvaluation runs attach dataset metadata such as dataset_name, case_id,
group, and tags to the LangSmith run context.
Current LangSmith workflow:
- Set
RAG_LANGSMITH_ENABLED=true,LANGCHAIN_TRACING_V2=true, and the usual LangSmith credentials. askandresumeruns create a LangSmith chain run with metadata such assession_id,index_dir, andresume.eval runcreates a dataset-level LangSmith run namedgraph_rag.eval.dataset:<dataset_name>.- Each evaluation case creates its own LangSmith run named
graph_rag.eval.case:<case_id>. - Case runs inherit evaluation metadata and derived tags such as
mode:eval,dataset:<name>,group:<group>, andcase:<id>. - After a case completes, the local evaluation system syncs feedback back to
the same LangSmith run:
case_passedlayer_questionlayer_retrievallayer_trajectorylayer_sourceslayer_web_search_qualitylayer_source_qualitylayer_answertrajectory_summaryanswer_judge_score,answer_grounded,answer_completewhen the answer judge is enabled
- After the dataset run finishes, the dataset-level LangSmith run receives
aggregate feedback such as
dataset_pass_rateanddataset_layer_<layer_name>.
That means the local files under runtime/evals/* remain the release gate,
while LangSmith becomes the observation and analysis layer for:
- case trace debugging
- layer-by-layer evaluation feedback
- answer judge groundedness/completeness review
- dataset-level pass-rate tracking
Generated runtime files are stored under runtime:
- runtime/checkpoints.db
- runtime/logs/graph_rag.log
runtime/service-stdout.logruntime/service-stderr.log
- Use agent for the local index directory.
- Use runtime for checkpoints and logs.
- If a command depends on relative paths, run it from this project directory.