Yongjian Li1,2,*, HaoCheng Chu3,*, Yukun Yan4,†, Zhenghao Liu8,
Shi Yu4, Zheni Zeng4, Ruobing Wang9, Sen Song1,2,†,
Zhiyuan Liu4,5,6,7, Maosong Sun4,5,6,7
Affiliations
1School of Biomedical Engineering, Tsinghua Medicine, Tsinghua University
2Tsinghua Laboratory of Brain and Intelligence
3School of Informatics, Xiamen University
4Department of Computer Science and Technology, Tsinghua University, Beijing
5Beijing National Research Center for Information Science and Technology, Tsinghua University, Beijing
6Institute for Artificial Intelligence, Tsinghua University, Beijing
7Jiangsu Collaborative Innovation Center for Language Ability, Jiangsu Normal University, Xuzhou
8School of Computer Science and Engineering, Northeastern University
9Institute of Information Engineering, Chinese Academy of Sciences
*Equal contribution. †Corresponding author.
KARE-RAG trains RAG generators using expert corrections to evidence graphs. The original and corrected graphs form preference pairs for token-weighted DDPO. After training, the model answers directly from retrieved documents using standard Vanilla RAG.
Use Python 3.10+ with a CUDA 12-compatible NVIDIA driver. Install dependencies from the repository root:
pip install -r requirements.txtFlashRAG and GPU FAISS are included in requirements.txt. If another FAISS package is already installed, uninstall it first.
Set CUDA_VISIBLE_DEVICES to select GPUs for indexing, retrieval, and training.
Download MuSiQue and the Wikipedia corpus from FlashRAG datasets. Keep the original format under datasets/flashrag/<dataset>/<split>.jsonl. Data and models are not included in this repository.
export MODEL_PATH=/path/to/Llama-3.1-8B-Instruct
export RETRIEVER_MODEL=/path/to/bge-large-en-v1.5
export CORPUS_PATH=datasets/flashrag/retrieval-corpus/wiki18_100w.jsonl
export INDEX_PATH=outputs/indexes/bge_Flat.index
mkdir -p outputs/indexes logs/indexing
python -m flashrag.retriever.index_builder \
--retrieval_method bge --model_path "$RETRIEVER_MODEL" \
--corpus_path "$CORPUS_PATH" --save_dir outputs/indexes \
--max_length 180 --batch_size 256 --use_fp16 \
--pooling_method cls --faiss_type Flat --sentence_transformer --faiss_gpu \
> logs/indexing/build.log 2>&1Set the generator endpoint and expert API key. GPT-4o-mini is the default expert; another model can be set with EXPERT_MODEL and EXPERT_BASE_URL.
export GENERATOR_BASE_URL=http://127.0.0.1:8000/v1
export GENERATOR_MODEL=Llama-3.1-8B-Instruct
export GENERATOR_API_KEY=EMPTY
export EXPERT_API_KEY=your-api-key
python scripts/data/construct_data.py \
--tokenizer-path "$MODEL_PATH" \
--retrieval-model-path "$RETRIEVER_MODEL" \
--corpus-path "$CORPUS_PATH" --index-path "$INDEX_PATH"This retrieves documents for MuSiQue training questions, constructs and refines evidence graphs, and saves preference pairs to outputs/data_generation/musique/pairs/kare_dpo_data.jsonl.
python scripts/train/train_ddpo.py \
--model-path "$MODEL_PATH" \
--train-file outputs/data_generation/musique/pairs/kare_dpo_data.jsonl \
--output-dir outputs/checkpoints/kare \
--bf16 --gradient-checkpointingTraining uses LoRA with DDPO, a learning rate of 5e-5, and one epoch by default. The adapter and tokenizer are saved to the output directory.
Use an OpenAI-compatible endpoint serving the trained model with its adapter. Set --model to its served model name:
python scripts/main_experiments/eval_openai.py \
--base-url "$GENERATOR_BASE_URL" --model KARE-Llama-3.1-8B-Instruct \
--tokenizer-path "$MODEL_PATH" \
--retrieval-model-path "$RETRIEVER_MODEL" \
--corpus-path "$CORPUS_PATH" --index-path "$INDEX_PATH" \
--datasets musique --split devAlso supported: nq, hotpotqa, popqa, truthful_qa, and zsre. Predictions and metrics are saved under outputs/evaluation/kare/; logs are written to logs/. See each script's --help for additional options.
src/
├── data_construction/ # Graph generation and preference pairs
├── training/ # DDPO trainer
├── evaluation/ # FlashRAG retrieval and OpenAI inference
└── utils/ # Prompt formatting, metrics, and logging
Built on FlashRAG, TRL, and PEFT.
If you use KARE-RAG in your research, please cite our work:
@misc{li2025kareragknowledgeawarerefinementenhancement,
title={KARE-RAG: Knowledge-Aware Refinement and Enhancement for RAG},
author={Yongjian Li and HaoCheng Chu and Yukun Yan and Zhenghao Liu and Shi Yu and Zheni Zeng and Ruobing Wang and Sen Song and Zhiyuan Liu and Maosong Sun},
year={2025},
eprint={2506.02503},
archivePrefix={arXiv},
primaryClass={cs.CL},
url={https://arxiv.org/abs/2506.02503},
}Contact Yongjian Li.
