Skip to content

About

This project implements an end-to-end Retrieval-Augmented Generation (RAG) pipeline to extract, analyze, and answer questions from research papers with minimal hallucination.

Resources

Stars

1 star

Watchers

0 watching

Forks

Latest commit

Β 

History

10 Commits

Folders and files

NameName
Last commit message
Last commit date
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 

Repository files navigation

ResearchPortal β€” AI-Powered Research Paper Analyzer

Python Flask LangChain Groq FAISS HuggingFace


πŸ“– Table of Contents


Overview

ResearchPortal is an end-to-end Retrieval-Augmented Generation (RAG) system designed to extract, analyze, and answer questions from research paper PDFs with minimal hallucination.

By combining semantic search powered by FAISS with LLM reasoning via Groq and LLaMA 3, the platform delivers accurate, context-aware responses for academic research workflows.

How It Works

Upload PDF  β†’  Select Section  β†’  View Summary  β†’  Ask Questions

Key Features

Feature Description
πŸ“€ PDF Upload Upload research papers in PDF format with a simple web interface
πŸ” Section Detection Automatically identifies and splits structured academic sections using regex-based parsing
πŸ€– LLM Refinement Cleans and normalizes noisy text extracted from PDFs using Groq-powered LLM processing
πŸ“ Section-wise Summaries Generates detailed, AI-driven summaries for each detected section of the paper
πŸ’¬ RAG Chat System Interactive Q&A interface to query the contents of the uploaded paper
🧠 Semantic Search FAISS vector store combined with HuggingFace sentence embeddings for precise retrieval
🌐 Flask Web UI Lightweight, interactive web interface built with Flask

Architecture & Workflow

The system follows a multi-stage pipeline from PDF ingestion to answer generation:

β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚   PDF Upload    β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”˜
         β–Ό
β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚  Text Extraction β”‚   ← PyPDF2
β””β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”˜
         β–Ό
β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚ Section Detectionβ”‚  ← Regex-based parsing
β””β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”˜
         β–Ό
β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚  LLM Refinement  β”‚  ← Groq (text cleanup)
β””β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”˜
         β–Ό
β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚  Text Chunking   β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”˜
         β–Ό
β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚   Embedding      β”‚  ← HuggingFace Embeddings
β””β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”˜
         β–Ό
β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚ FAISS Vector     β”‚  ← Semantic search store
β”‚ Search           β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”˜
         β–Ό
β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚  LLM Answer       β”‚  ← Groq + LLaMA 3
β”‚  Generation       β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜

Project Structure

ResearchPortal/
β”œβ”€β”€ app.py                          # Flask application entry point
β”œβ”€β”€ src/
β”‚   β”œβ”€β”€ load_and_extract_text.py    # PDF text extraction (PyPDF2)
β”‚   β”œβ”€β”€ detect_and_split_sections.py # Regex-based section detection
β”‚   β”œβ”€β”€ get_summary.py              # LLM-powered section summarization
β”‚   β”œβ”€β”€ create_vector_db.py         # FAISS vector store creation
β”‚   └── RAG_retrival_chain.py       # RAG retrieval & answer chain
β”œβ”€β”€ templates/
β”‚   └── index.html                  # Web UI template
β”œβ”€β”€ uploads/                        # Uploaded PDFs (git-ignored)
β”œβ”€β”€ .env                            # Environment variables (git-ignored)
β”œβ”€β”€ requirements.txt               # Python dependencies
└── README.md                       # Project documentation

Supported LLM Models

All models are served through the Groq API platform:

Model Context Window Best For
llama3-8b-8192 8,192 tokens ⚑ Recommended β€” Fast inference, balanced performance
llama3-70b-8192 8,192 tokens High-quality reasoning responses
mixtral-8x7b-32768 32,768 tokens Long-context processing
gemma-7b-it 8,192 tokens Lightweight alternative

Technology Stack

Backend

Technology Role
Flask Web framework & API server
LangChain RAG orchestration & chaining
FAISS Vector similarity search
PyPDF2 PDF text extraction
Groq API LLM inference backend
HuggingFace Embeddings Sentence-level text embedding

Frontend

Technology Role
HTML / CSS / JavaScript Web interface
Font Awesome Iconography
Google Fonts Typography

Getting Started

Prerequisites

  • Python 3.9 or higher
  • A Groq API key (available at groq.com)

1. Clone the Repository

git clone https://github.com/Intelli2Byte/Automated-Research-Synthesis-ARS-.git
cd Automated-Research-Synthesis-ARS-

2. Create a Virtual Environment

python -m venv venv

# Activate on Windows
venv\Scripts\activate

# Activate on macOS / Linux
source venv/bin/activate

3. Install Dependencies

pip install -r requirements.txt

4. Configure Environment Variables

Create a .env file in the project root:

GROQ_API_KEY=your_api_key_here

5. Launch the Application

python app.py

The web interface will be available at http://127.0.0.1:5000 by default.


Challenges & Solutions

Challenge Approach
Noisy PDF text extraction LLM-based refinement pipeline to clean and normalize extracted text
Accurate section identification Regex-based detection supplemented with LLM processing
Reducing LLM hallucination RAG architecture grounds responses in retrieved source context
Optimizing retrieval performance FAISS vector index with HuggingFace embeddings for efficient semantic search

Roadmap

  • Multi-paper comparison β€” Analyze and compare multiple research papers side by side
  • Citation extraction β€” Automatically extract and format citations from papers
  • UI enhancements β€” Improved user experience and responsive design
  • Cloud deployment β€” Production-ready cloud hosting and scaling

Contributing

Contributions are welcome! Please follow these steps:

  1. Fork the repository

  2. Create a feature branch

    git checkout -b feature/your-feature-name
  3. Commit your changes

    git commit -m "Add: feature description"
  4. Push to your branch

    git push origin feature/your-feature-name
  5. Open a Pull Request β€” Describe your changes and their purpose


Acknowledgements

This project builds upon the following open-source tools and platforms:

  • Groq β€” Fast LLM inference
  • LangChain β€” LLM application framework
  • HuggingFace β€” Embedding models
  • FAISS β€” Vector similarity search
  • PyPDF2 β€” PDF processing

Author

Neha Maurya

Platform Link
πŸ“§ Email mauryaneha2006@gmail.com
πŸ”— LinkedIn linkedin.com/in/neha-maurya-644a1a290

πŸ’Ό Internship Opportunities

I'm actively seeking internship opportunities in AI, Machine Learning, and Full-Stack Development. If you feel my work on this project demonstrates a good fit for your team, I'd love to hear from you!

I'm eager to contribute, learn, and grow β€” if I'd be a good fit, let me in! πŸš€


Built with ❀️ for academic research workflows.

About

This project implements an end-to-end Retrieval-Augmented Generation (RAG) pipeline to extract, analyze, and answer questions from research papers with minimal hallucination.

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages