You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Star0 (0)You must be signed in to star a repository
About
A production ready /mutli-cloud GenAI Data ingestion pipeline for document processing, LLM-based tagging and RAG(Retrieval Augmented Generation) implementation. This Project demonstrate enterprise-grade infra as code deployments both on AWS and GCP
A production-ready, multi-cloud GenAI data ingestion pipeline for document processing, LLM-based tagging, and RAG (Retrieval-Augmented Generation) implementation. This project demonstrates enterprise-grade infrastructure as code with deployments on both AWS and Google Cloud Platform (GCP).
This pipeline automates the ingestion, processing, and vectorization of documents for use in RAG-based AI applications. It supports multiple document types (PDF, DOCX, TXT, etc.) and provides:
Document Processing: Text extraction, standardization, and semantic chunking
LLM Tagging: Automated metadata extraction using large language models
Vector Storage: Embeddings generation and storage for similarity search
RAG Queries: Natural language querying over your document corpus
The GCP deployment uses a hybrid architecture that combines Cloud Run for low-traffic services and GKE for high-compute services, optimizing for both cost and performance.
Service Mapping
Component
GCP Service
Low-Traffic Compute
Cloud Run
High-Compute
GKE (Kubernetes)
Orchestration
Cloud Workflows
Storage
Cloud Storage
Queue
Pub/Sub
Events
Eventarc
Vector DB
Vertex AI Vector Search
LLM
Vertex AI (Gemini 1.5 Pro)
Embeddings
Vertex AI (text-embedding-004)
Database
Cloud SQL (PostgreSQL)
API
API Gateway
Service Classification
Cloud Run Services (Low-Traffic, Scale-to-Zero)
detect-file-type - File type detection
text-standardize - Text standardization
identify-distinct-process - Process identification
create-process-docs - Document creation
read-from-storage - Storage operations
add-llm-tags - Metadata tagging
GKE Services (High-Compute, Always-On)
text-extraction - Heavy PDF/document processing
semantic-chunking - ML-based chunking
llm-tagging - LLM inference
chunk-sop - SOP chunking
generate-embedding - Embedding generation
store-to-vector-db - Vector storage
Prerequisites
Google Cloud SDK (gcloud) installed
Terraform >= 1.4.0
kubectl (for GKE management)
Docker (for local testing)
Deployment Steps
cd iac-gcp
# Authenticate with GCP
gcloud auth login
gcloud auth application-default login
# Set project
gcloud config set project YOUR_PROJECT_ID
# Enable required APIs
gcloud services enable \
compute.googleapis.com \
container.googleapis.com \
run.googleapis.com \
workflows.googleapis.com \
eventarc.googleapis.com \
pubsub.googleapis.com \
storage.googleapis.com \
aiplatform.googleapis.com \
sqladmin.googleapis.com \
secretmanager.googleapis.com
# Create Terraform state bucket
gsutil mb -l us-central1 gs://tsh-industries-terraform-state
gsutil versioning set on gs://tsh-industries-terraform-state
# Update backend.tf to uncomment the GCS backend# Initialize Terraform
terraform init
# Select workspace
terraform workspace selectdev# Update dev.tfvars: set skip_gcp_auth = false# Review plan
terraform plan -var-file="environments/dev.tfvars"# Deploy
terraform apply -var-file="environments/dev.tfvars"# Deploy Kubernetes manifests (after GKE is ready)
gcloud container clusters get-credentials $(terraform output -raw gke_cluster_name) --region us-central1
kubectl apply -k k8s/
Scale to Zero: Cloud Run services cost nothing when idle
Right-size instances: Start small, scale based on metrics
Use committed use discounts: For predictable workloads
Optimize LLM calls: Use Gemini Flash for simple tasks
Batch embeddings: Process in batches to reduce API calls
Break-Even Analysis
Workload
Cloud Run Only
Hybrid
Recommendation
< 1,000 docs/month
$200
$350
Cloud Run
1,000-5,000 docs/month
$500
$450
Hybrid
5,000-20,000 docs/month
$1,200
$900
Hybrid
> 20,000 docs/month
$2,500+
$1,500
Hybrid
Local Development
Quick Start with Docker Compose
cd iac-gcp
# Start local environment (emulators + databases)
./scripts/local-setup.sh start
# Access services:# - MinIO Console: http://localhost:9001 (minioadmin/minioadmin)# - PostgreSQL: localhost:5432 (tsh-industries/localdevpassword)# - Adminer (DB UI): http://localhost:8088# - Pub/Sub Emulator: localhost:8085# - Qdrant (Vector DB): http://localhost:6333# - Redis: localhost:6379# Run a test through the pipeline
./scripts/local-setup.sh test# View logs
./scripts/local-setup.sh logs
# Stop all services
./scripts/local-setup.sh stop
# Clean up (remove all data)
./scripts/local-setup.sh cleanup
Terraform Validation (Without GCP Account)
cd iac-gcp
# Initialize without backend
terraform init -backend=false
# Validate configuration
terraform validate
# Plan with local validation mode
terraform plan -var-file="environments/dev.tfvars"# Note: dev.tfvars has skip_gcp_auth=true for local validation
Environment Variables
AWS
Variable
Description
AWS_REGION
AWS region (default: us-east-1)
AWS_PROFILE
AWS CLI profile name
TF_VAR_environment
Environment name
GCP
Variable
Description
GOOGLE_PROJECT
GCP project ID
GOOGLE_REGION
GCP region (default: us-central1)
TF_VAR_skip_gcp_auth
Skip GCP auth for local validation
Monitoring & Observability
AWS
CloudWatch Logs: Centralized logging
CloudWatch Metrics: Performance metrics
X-Ray: Distributed tracing
CloudWatch Alarms: Alerting
GCP
Cloud Logging: Centralized logging
Cloud Monitoring: Metrics and dashboards
Cloud Trace: Distributed tracing
Error Reporting: Error aggregation
Security Considerations
Network Isolation: VPC with private subnets
IAM: Least-privilege service accounts
Encryption: KMS for data at rest
Secrets: Secret Manager for credentials
API Security: JWT authentication
Audit Logging: All API calls logged
Troubleshooting
Common Issues
Terraform State Lock
# Force unlock (use with caution)
terraform force-unlock LOCK_ID
GKE Authentication
# Refresh credentials
gcloud container clusters get-credentials CLUSTER_NAME --region REGION
Cloud Run Cold Starts
Increase min_instances for critical services
Use Cloud Run "always on CPU" feature
Contributing
Fork the repository
Create a feature branch (git checkout -b feature/amazing-feature)
Commit your changes (git commit -m 'Add amazing feature')
Push to the branch (git push origin feature/amazing-feature)
Open a Pull Request
License
This project is licensed under the MIT License - see the LICENSE file for details.
Acknowledgments
AWS and GCP documentation
Terraform community modules
Open-source LLM and embedding models
Contact
For questions or support, please open an issue in this repository.
About
A production ready /mutli-cloud GenAI Data ingestion pipeline for document processing, LLM-based tagging and RAG(Retrieval Augmented Generation) implementation. This Project demonstrate enterprise-grade infra as code deployments both on AWS and GCP