- Overview
- Key Features
- Architecture & Workflow
- Project Structure
- Supported LLM Models
- Technology Stack
- Getting Started
- Challenges & Solutions
- Roadmap
- Contributing
- Acknowledgements
- Author
ResearchPortal is an end-to-end Retrieval-Augmented Generation (RAG) system designed to extract, analyze, and answer questions from research paper PDFs with minimal hallucination.
By combining semantic search powered by FAISS with LLM reasoning via Groq and LLaMA 3, the platform delivers accurate, context-aware responses for academic research workflows.
Upload PDF β Select Section β View Summary β Ask Questions
| Feature | Description |
|---|---|
| π€ PDF Upload | Upload research papers in PDF format with a simple web interface |
| π Section Detection | Automatically identifies and splits structured academic sections using regex-based parsing |
| π€ LLM Refinement | Cleans and normalizes noisy text extracted from PDFs using Groq-powered LLM processing |
| π Section-wise Summaries | Generates detailed, AI-driven summaries for each detected section of the paper |
| π¬ RAG Chat System | Interactive Q&A interface to query the contents of the uploaded paper |
| π§ Semantic Search | FAISS vector store combined with HuggingFace sentence embeddings for precise retrieval |
| π Flask Web UI | Lightweight, interactive web interface built with Flask |
The system follows a multi-stage pipeline from PDF ingestion to answer generation:
βββββββββββββββββββ
β PDF Upload β
ββββββββββ¬βββββββββ
βΌ
βββββββββββββββββββ
β Text Extraction β β PyPDF2
ββββββββββ¬βββββββββ
βΌ
βββββββββββββββββββ
β Section Detectionβ β Regex-based parsing
ββββββββββ¬βββββββββ
βΌ
βββββββββββββββββββ
β LLM Refinement β β Groq (text cleanup)
ββββββββββ¬βββββββββ
βΌ
βββββββββββββββββββ
β Text Chunking β
ββββββββββ¬βββββββββ
βΌ
βββββββββββββββββββ
β Embedding β β HuggingFace Embeddings
ββββββββββ¬βββββββββ
βΌ
βββββββββββββββββββ
β FAISS Vector β β Semantic search store
β Search β
ββββββββββ¬βββββββββ
βΌ
βββββββββββββββββββ
β LLM Answer β β Groq + LLaMA 3
β Generation β
βββββββββββββββββββ
ResearchPortal/
βββ app.py # Flask application entry point
βββ src/
β βββ load_and_extract_text.py # PDF text extraction (PyPDF2)
β βββ detect_and_split_sections.py # Regex-based section detection
β βββ get_summary.py # LLM-powered section summarization
β βββ create_vector_db.py # FAISS vector store creation
β βββ RAG_retrival_chain.py # RAG retrieval & answer chain
βββ templates/
β βββ index.html # Web UI template
βββ uploads/ # Uploaded PDFs (git-ignored)
βββ .env # Environment variables (git-ignored)
βββ requirements.txt # Python dependencies
βββ README.md # Project documentation
All models are served through the Groq API platform:
| Model | Context Window | Best For |
|---|---|---|
llama3-8b-8192 |
8,192 tokens | β‘ Recommended β Fast inference, balanced performance |
llama3-70b-8192 |
8,192 tokens | High-quality reasoning responses |
mixtral-8x7b-32768 |
32,768 tokens | Long-context processing |
gemma-7b-it |
8,192 tokens | Lightweight alternative |
| Technology | Role |
|---|---|
| Flask | Web framework & API server |
| LangChain | RAG orchestration & chaining |
| FAISS | Vector similarity search |
| PyPDF2 | PDF text extraction |
| Groq API | LLM inference backend |
| HuggingFace Embeddings | Sentence-level text embedding |
| Technology | Role |
|---|---|
| HTML / CSS / JavaScript | Web interface |
| Font Awesome | Iconography |
| Google Fonts | Typography |
- Python 3.9 or higher
- A Groq API key (available at groq.com)
git clone https://github.com/Intelli2Byte/Automated-Research-Synthesis-ARS-.git
cd Automated-Research-Synthesis-ARS-python -m venv venv
# Activate on Windows
venv\Scripts\activate
# Activate on macOS / Linux
source venv/bin/activatepip install -r requirements.txtCreate a .env file in the project root:
GROQ_API_KEY=your_api_key_herepython app.pyThe web interface will be available at http://127.0.0.1:5000 by default.
| Challenge | Approach |
|---|---|
| Noisy PDF text extraction | LLM-based refinement pipeline to clean and normalize extracted text |
| Accurate section identification | Regex-based detection supplemented with LLM processing |
| Reducing LLM hallucination | RAG architecture grounds responses in retrieved source context |
| Optimizing retrieval performance | FAISS vector index with HuggingFace embeddings for efficient semantic search |
- Multi-paper comparison β Analyze and compare multiple research papers side by side
- Citation extraction β Automatically extract and format citations from papers
- UI enhancements β Improved user experience and responsive design
- Cloud deployment β Production-ready cloud hosting and scaling
Contributions are welcome! Please follow these steps:
-
Fork the repository
-
Create a feature branch
git checkout -b feature/your-feature-name
-
Commit your changes
git commit -m "Add: feature description" -
Push to your branch
git push origin feature/your-feature-name
-
Open a Pull Request β Describe your changes and their purpose
This project builds upon the following open-source tools and platforms:
- Groq β Fast LLM inference
- LangChain β LLM application framework
- HuggingFace β Embedding models
- FAISS β Vector similarity search
- PyPDF2 β PDF processing
Neha Maurya
| Platform | Link |
|---|---|
| π§ Email | mauryaneha2006@gmail.com |
| π LinkedIn | linkedin.com/in/neha-maurya-644a1a290 |
I'm actively seeking internship opportunities in AI, Machine Learning, and Full-Stack Development. If you feel my work on this project demonstrates a good fit for your team, I'd love to hear from you!
- π§ Reach out: mauryaneha2006@gmail.com
- π Connect on LinkedIn: linkedin.com/in/neha-maurya-644a1a290
I'm eager to contribute, learn, and grow β if I'd be a good fit, let me in! π
Built with β€οΈ for academic research workflows.