Sign inSign up

kchetan/co2-dashboard

By kchetan

•Updated 10 months ago

Interactive Global CO₂ Dashboard implementing a rigorous, reproducible Data Engineering pipeline.

Image
Data science
0

85

kchetan/co2-dashboard repository overview

Global CO₂ Intelligence Platform 🌍

A Lakehouse-Lite implementation of modern Data Engineering principles. Stateless. Declarative. Strictly Typed. The Global CO₂ Intelligence Platform is an interactive analytics dashboard visualizing global emissions data from 1750 to the present.

Unlike standard dashboards that simply read a static CSV, this image contains a robust, production-grade Data Engineering Pipeline scaled down for a portfolio application. It implements a "Lakehouse-Lite" architecture where:

Ingestion is stateless and reproducible.

Transformations are handled by an in-process OLAP engine (DuckDB).

Quality is enforced via runtime schema validation (Pandera).

Visualization is decoupled from data logic (Streamlit).

🚀 Quick Start

Because this application follows a strict "Lakehouse" architecture, data is generated on the fly and persisted to your local machine via volume mapping.

  1. Pull the Image Bash

docker pull kchetan/co2-dashboard:v1.0

  1. Run the Data Pipeline (ETL)

This step fetches the latest raw data from the cloud, transforms it using SQL, and validates it against the schema. The processed data is saved to your local data/ folder. Bash

⁠Create a local folder for the data

mkdir -p data

⁠Run Ingestion & Transformation

docker run --rm -v $(pwd)/data:/app/data --entrypoint python kchetan/co2-dashboard:v1.0 src/ingest.py docker run --rm -v $(pwd)/data:/app/data --entrypoint python kchetan/co2-dashboard:v1.0 src/transform.py

  1. Launch the Application

Start the dashboard, mounting the data you just generated. Bash

docker run -p 8501:8501 -v $(pwd)/data:/app/data kchetan/co2-dashboard:v1.0

Access the App: Open your browser to http://localhost:8501

🏗️ Architecture

The system follows a strict separation of concerns:

Ingest: Python scripts fetch raw CSV data from the OWID GitHub repository.

Processing: DuckDB serves as an in-process OLAP engine. It executes SQL transformations to clean, cast, and aggregate data into Parquet files.

Validation: Pandera enforces strict schema contracts (types, ranges, null-checks) before data is allowed into the visualization layer.

Presentation: Streamlit renders the frontend, reading directly from the validated Parquet output.

🛠️ Tech Stack Component Technology Purpose Compute DuckDB 🦆 In-process OLAP & SQL transformation Quality Pandera 🛡️ Statistical typing & data validation Frontend Streamlit 👑 Interactive web application Viz Plotly Express Interactive charts Base Image python:3.9-slim Lightweight, secure runtime 🔗 Resources

GitHub Repository: k-chetan/CO2-Dashboard

Data Source: Our World in Data

License: MIT

Tag summary

Content type

Image

Digest

sha256:f69c960ae…

Size

553.8 MB

Last updated

10 months ago

docker pull kchetan/co2-dashboard:v1.0