Skip to content
thunlpPublic

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Latest commit

 

History

1 Commit

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

KARE-RAG: Knowledge-Aware Refinement and Enhancement for RAG

KARE-RAG paper KARE-RAG code FlashRAG framework

Yongjian Li1,2,*, HaoCheng Chu3,*, Yukun Yan4,†, Zhenghao Liu8,
Shi Yu4, Zheni Zeng4, Ruobing Wang9, Sen Song1,2,†,
Zhiyuan Liu4,5,6,7, Maosong Sun4,5,6,7

Affiliations

1School of Biomedical Engineering, Tsinghua Medicine, Tsinghua University
2Tsinghua Laboratory of Brain and Intelligence
3School of Informatics, Xiamen University
4Department of Computer Science and Technology, Tsinghua University, Beijing
5Beijing National Research Center for Information Science and Technology, Tsinghua University, Beijing
6Institute for Artificial Intelligence, Tsinghua University, Beijing
7Jiangsu Collaborative Innovation Center for Language Ability, Jiangsu Normal University, Xuzhou
8School of Computer Science and Engineering, Northeastern University
9Institute of Information Engineering, Chinese Academy of Sciences

*Equal contribution. †Corresponding author.

Overview · Setup · Usage · Citation

📖 Introduction

KARE-RAG trains RAG generators using expert corrections to evidence graphs. The original and corrected graphs form preference pairs for token-weighted DDPO. After training, the model answers directly from retrieved documents using standard Vanilla RAG.

KARE-RAG data construction, expert refinement, DDPO training, and standard RAG inference.

⚙️ Setup

Use Python 3.10+ with a CUDA 12-compatible NVIDIA driver. Install dependencies from the repository root:

pip install -r requirements.txt

FlashRAG and GPU FAISS are included in requirements.txt. If another FAISS package is already installed, uninstall it first.

Set CUDA_VISIBLE_DEVICES to select GPUs for indexing, retrieval, and training.

🔧 Usage

1. Prepare Data

Download MuSiQue and the Wikipedia corpus from FlashRAG datasets. Keep the original format under datasets/flashrag/<dataset>/<split>.jsonl. Data and models are not included in this repository.

export MODEL_PATH=/path/to/Llama-3.1-8B-Instruct
export RETRIEVER_MODEL=/path/to/bge-large-en-v1.5
export CORPUS_PATH=datasets/flashrag/retrieval-corpus/wiki18_100w.jsonl
export INDEX_PATH=outputs/indexes/bge_Flat.index

mkdir -p outputs/indexes logs/indexing
python -m flashrag.retriever.index_builder \
  --retrieval_method bge --model_path "$RETRIEVER_MODEL" \
  --corpus_path "$CORPUS_PATH" --save_dir outputs/indexes \
  --max_length 180 --batch_size 256 --use_fp16 \
  --pooling_method cls --faiss_type Flat --sentence_transformer --faiss_gpu \
  > logs/indexing/build.log 2>&1

2. Construct Training Pairs

Set the generator endpoint and expert API key. GPT-4o-mini is the default expert; another model can be set with EXPERT_MODEL and EXPERT_BASE_URL.

export GENERATOR_BASE_URL=http://127.0.0.1:8000/v1
export GENERATOR_MODEL=Llama-3.1-8B-Instruct
export GENERATOR_API_KEY=EMPTY
export EXPERT_API_KEY=your-api-key

python scripts/data/construct_data.py \
  --tokenizer-path "$MODEL_PATH" \
  --retrieval-model-path "$RETRIEVER_MODEL" \
  --corpus-path "$CORPUS_PATH" --index-path "$INDEX_PATH"

This retrieves documents for MuSiQue training questions, constructs and refines evidence graphs, and saves preference pairs to outputs/data_generation/musique/pairs/kare_dpo_data.jsonl.

3. Train

python scripts/train/train_ddpo.py \
  --model-path "$MODEL_PATH" \
  --train-file outputs/data_generation/musique/pairs/kare_dpo_data.jsonl \
  --output-dir outputs/checkpoints/kare \
  --bf16 --gradient-checkpointing

Training uses LoRA with DDPO, a learning rate of 5e-5, and one epoch by default. The adapter and tokenizer are saved to the output directory.

4. Evaluate

Use an OpenAI-compatible endpoint serving the trained model with its adapter. Set --model to its served model name:

python scripts/main_experiments/eval_openai.py \
  --base-url "$GENERATOR_BASE_URL" --model KARE-Llama-3.1-8B-Instruct \
  --tokenizer-path "$MODEL_PATH" \
  --retrieval-model-path "$RETRIEVER_MODEL" \
  --corpus-path "$CORPUS_PATH" --index-path "$INDEX_PATH" \
  --datasets musique --split dev

Also supported: nq, hotpotqa, popqa, truthful_qa, and zsre. Predictions and metrics are saved under outputs/evaluation/kare/; logs are written to logs/. See each script's --help for additional options.

📁 Code Structure

src/
├── data_construction/   # Graph generation and preference pairs
├── training/            # DDPO trainer
├── evaluation/          # FlashRAG retrieval and OpenAI inference
└── utils/               # Prompt formatting, metrics, and logging

📄 Acknowledgements

Built on FlashRAG, TRL, and PEFT.

🥰 Citation

If you use KARE-RAG in your research, please cite our work:

@misc{li2025kareragknowledgeawarerefinementenhancement,
  title={KARE-RAG: Knowledge-Aware Refinement and Enhancement for RAG},
  author={Yongjian Li and HaoCheng Chu and Yukun Yan and Zhenghao Liu and Shi Yu and Zheni Zeng and Ruobing Wang and Sen Song and Zhiyuan Liu and Maosong Sun},
  year={2025},
  eprint={2506.02503},
  archivePrefix={arXiv},
  primaryClass={cs.CL},
  url={https://arxiv.org/abs/2506.02503},
}

📧 Contact

Contact Yongjian Li.

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages