Skip to content
View simones99's full-sized avatar

Block or report simones99

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
simones99/README.md

Simone Mezzabotta

Data analyst based in Bologna, working on data warehousing, ETL pipelines and data quality at UNI, the Italian standardisation body. Former Blue Book trainee at the European Commission (2024–2025).

I am interested in public-sector and official statistics: turning administrative and open data into reliable, documented indicators.

Tools: SQL · Python (pandas, scikit-learn) · R (sdcMicro, sdcTable) · dbt · DuckDB · PostgreSQL · Power BI · data validation and data quality checks · Git and CI

Selected projects

Project What it shows
eu-stats-pipeline Eurostat SDMX ingestion with file-level lineage, dbt models on DuckDB, 27 validation rules including regional–national coherence, an ESS-structured quality report, a generated data dictionary and a NUTS 2 dashboard
marche-regional-stats Marche region: business-register checks on constant boundaries, NUTS 3 indicators and maps from Eurostat, SARIMA forecast evaluated out of sample, Power BI reports
eu-silc-sdc Statistical disclosure control in R (sdcMicro, sdcTable) on EU-SILC public microdata: re-identification risk, three protection scenarios, information loss on poverty and income indicators, audited cell suppression for a magnitude table
juve-momentum-index Football data platform (FastAPI, PostgreSQL, Next.js): scheduled ingestion, Elo ratings, calibrated win/draw/loss probabilities backtested against baselines
bloodio Local-first web app that parses lab-report PDFs in the browser, validates and normalises values and charts trends; privacy by design, live demo
csv_ai_agent Querying CSV files in natural language with a local LLM; data validation and outlier checks; tested with CI
goal-line-calibration Walk-forward evaluation of Elo, Dixon–Coles, logistic regression and XGBoost forecasts against de-margined bookmaker odds on 45,000 matches: frozen hyperparameters, paired bootstrap, calibration, published report
employee-attrition-prediction Capstone case study (Salifort Motors): EDA, logistic regression and random forest on HR data, with an executive summary

Languages

Italian (native) · English (C2) · French (C1) · Spanish (C1)

Pinned Loading

  1. csv_ai_agent csv_ai_agent Public

    Local tool to query CSV files using natural language, no APIs, full privacy. Optimized for Apple Silicon with interactive dashboard and advanced AI analysis.

    Python

  2. employee-attrition-prediction employee-attrition-prediction Public

    Predicting Employee Attrition at Salifort Motors using Machine Learning. This project analyses HR data to identify factors influencing employee turnover and builds a predictive model to support ret…

    Jupyter Notebook

  3. eu-silc-sdc eu-silc-sdc Public

    Statistical disclosure control on EU-SILC microdata in R: re-identification risk, sdcMicro protection scenarios, information loss, and audited cell suppression with sdcTable

    R

  4. eu-stats-pipeline eu-stats-pipeline Public

    Reproducible Eurostat regional statistics pipeline: SDMX ingestion with lineage, dbt + DuckDB, validation rules, ESS quality report, data dictionary and dashboard.

    Python