An intelligent art gallery tour guide system that leverages Large Language Models (LLMs) and computer vision capabilities to generate detailed, contextual explanations for artworks. The system combines structured metadata with visual analysis using the Metropolitan Museum of Art's collection data.
-
Data Processing & Preparation
- Processed Met Museum dataset (50 sample artworks from Drawings and Prints)
- Optimized feature selection (35 columns)
- Maintained image URLs for visual analysis
-
Baseline Implementation
- Prompt engineering framework using Gemini API
- Structured prompts with artwork metadata
- Comprehensive evaluation system with quality ratings
-
Generation Control & Optimization
- Token length control (200 tokens)
- Response length analysis and quality correlation
- Generation quality evaluation across constraints
-
Multimodal Integration
- Image processing with Gemini's multimodal capabilities
- Combined textual and visual analysis
- Image-aware evaluation framework
- Python 3.x
- Key Libraries:
- python-dotenv (environment variables)
- pandas (data processing)
- google-generativeai (Gemini API)
- requests (HTTP requests)
- tqdm (progress tracking)
- aiohttp (async HTTP)
/Capstone Project
├── data/
│ └── sample50_df_drawings_and_prints.csv
├── demo.ipynb
├── results/
│ ├── baseline/
│ ├── output_length_control/
│ └── image_input/
└── README.md
The following large data files have been excluded from this repository due to GitHub's file size limitations:
Capstone Project/data/met_with_images_cleaned.csv(166.78 MB)Capstone Project/data/met_with_images.csv(167.20 MB)Capstone Project/data/met_public_domain.csv(151.83 MB)Capstone Project/data/Met/MetObjects.txt(302.94 MB)
These files can be obtained from the Metropolitan Museum of Art's Open Access dataset:
After downloading, place the files in the appropriate directories as listed above.
For detailed setup and usage instructions, refer to the project's dedicated README in the Capstone Project directory.
- Prompting Fundamentals:
- Introduction to prompt engineering
- Best practices for working with LLMs
- Techniques for effective prompt design
- Evaluation and Structured Output:
- Techniques for evaluating LLM outputs
- Automated evaluation methods
- Pointwise and pairwise evaluation approaches
- Working with structured outputs
- Practical evaluation challenges and solutions
- Reference Materials:
- Foundational Large Language Models & Text Generation
- Prompt Engineering Best Practices
- NeurIPS Evaluation Guidelines
- Embeddings and Similarity:
- Working with embeddings from Gemini API
- Calculating and utilizing similarity scores
- Understanding vector representations
- Classification with Keras:
- Using embeddings for classification tasks
- Implementing neural networks with Keras
- Building classification models
- Document Q&A with RAG:
- Implementation of Retrieval-Augmented Generation (RAG)
- Using Chroma for document storage and retrieval
- Building Q&A systems with RAG architecture
- Reference Materials:
- Embeddings and Vector Stores Guide
- Task-Aware Embedding Implementation
- Function Calling with Gemini API:
- Understanding function calling capabilities
- Building chat interfaces with automatic function calling
- Practical implementation of function-enabled conversations
- Building Agents with LangGraph:
- Introduction to LangGraph framework
- Creating and deploying AI agents
- Implementing agent-based workflows and interactions
- Reference Materials:
- AI Agents Implementation Guide
- Agents Companion Guide
- Fine-tuning Custom Models:
- Understanding model fine-tuning concepts
- Implementing task-specific model customization
- Best practices for fine-tuning Gemini models
- Google Search Integration:
- Incorporating Google Search results with Gemini API
- Implementing search-based grounding
- Building real-time information retrieval systems
- Reference Materials:
- Solving Domain-Specific Problems using LLMs
No practical sessions. Please review the materials from previous days and complete any pending exercises.
- Reference Materials:
- Operationalizing Generative AI on Vertex AI