A curated list of awesome open source tools and commercial products to catalog, version, and manage data 🚀
-
Updated
Apr 20, 2022
A curated list of awesome open source tools and commercial products to catalog, version, and manage data 🚀
🦀 A smart, persistent key-value store in Rust for managing data with a lifecycle. Features atomic TTLs, frequency counting, and is built on sled for performance and durability.
A spectral ClickHouse analyzer that tracks which tables are actually used and by whom. Maps queries to IPs/K8s services, flags dead tables and orphaned MVs, and helps keep your cluster clean and intentional.
KVKK & GDPR compliant data lifecycle management system with automated retention tracking, controlled destruction approval, PDF record generation, and audit logging. Built with FastAPI and React.
A coursera course on Machine Learning in Production
universal data decay engine that automatically compresses, summarizes, and purges stale files over time.
Alchemist is an intelligent data transformation engine that turns raw, messy data into clean, curated, and governed datasets. It automates and simplifies the data engineering lifecycle with a focus on quality, lineage, and operational excellence across the modern data ecosystem.
📊 Analyze ClickHouse usage to uncover table relationships, discover safe cleanup options, and generate insightful visual reports effortlessly.
The objective of this project was to develop a revenue forecasting solution for a digital wallet company. The primary goal was to analyze historical transaction data, create a predictive model to forecast future revenue, and present these insights through an interactive dashboard.
Practical, code-first guides for automating InfluxDB at IoT scale: task scheduling & orchestration, downsampling pipelines, retention and data-lifecycle management, and external workflow engines (Airflow, Prefect, Dagster, Kubernetes).
End-to-end Heart Disease Predictive Model using a Five-Layer Cloud Architecture (Bronze-Silver-Gold) on GCP with Dataproc, BigQuery ML, and Vertex AI.
Declarative, Terraform-style manager for InfluxDB 2.x bucket retention and downsampling tasks: plan/apply a YAML spec, generate Flux rollups, no state file
RDM course with tactical solutions for researchers and data professionals alike
Multi-tier storage lifecycle system for regulatory audit logs — hot Postgres to S3/Glacier, compliance-aware retention, crash-safe archival state machine
“A full-stack data lifecycle project for stock market data using Python, MySQL, Feature Engineering, and EDA, focused on FAANG companies.”
Práctica 1. Tipología y Ciclo de Vida de los Datos. Caso práctico de Web Scraping orientado a aprender a identificar los datos relevantes por un proyecto analítico y usar las herramientas de extracción de datos.
Storage tiering CLI for InfluxDB: archive aged time-series windows to compressed Parquet on S3/R2/MinIO with checksummed manifests, deep verification, guarded expiry and lossless rehydrate.
Enterprise IAM governance pipeline — maps 75 applications to IAM products, classifies risk, detects governance gaps, and produces management dashboards with remediation queues and department scorecards.
Free, hands-on data-literacy lab for non-coders — 68 interactive demos across the full data lifecycle: cleaning, analysis, machine learning, generative AI, and the statistical traps that fool decision-makers. Bilingual EN/العربية. No signup, open-source, citable (DOI).
Add a description, image, and links to the data-lifecycle topic page so that developers can more easily learn about it.
To associate your repository with the data-lifecycle topic, visit your repo's landing page and select "manage topics."