Marketing ML Lakehouse

Reusable data template

Turn campaign files into checked reporting tables. Explore the public sample, or rebuild the published data pipeline on your computer.

Browser snapshot needs no setup. Python is required to rebuild the template.

Before using this

The public console is a fixed sample snapshot, not a live view of the current Python output. Rebuild and inspect your own output in the local Streamlit dashboard. Neither route establishes results for a connected advertising account.

The console, pipeline and next-day evaluation are published and rebuilt in CI. The optional GA4 public-data profile needs your own BigQuery access and is not part of the published results.

Published lakehouse browser snapshot showing sample quality checks, not a live Python pipeline run

Public console built from the pipeline's generated evidence file.

My contribution Data and ML engineer
Status Public template with a local ML extension
Focus Lakehouse and ML delivery

Why I built it

Before a marketing team acts on a forecast, it needs to know whether the source data is usable and whether the model beats a simple alternative. This reusable local template checks campaign files, keeps a trace of their source, builds analysis tables and tests next-day bookings predictions against carrying today’s figure forward. It is an engineering template, not a connected advertising service.

What I chose

I made the data rebuildable from files through checked table layers to a dashboard. The prediction target is explicitly in the future and the model is compared with carrying today's figure forward.

What the example shows

The public console is rebuilt from the pipeline's generated evidence on every push: data-quality warnings, lineage and the next-day model check. On the 24-row holdout the model's error intervals overlap the prior-day baseline's, so it shows no reliable skill yet.

What I learned

A polished forecast is not evidence that a model adds value. Data available only after the prediction time must stay out of training features, and a model needs to beat a simple baseline on later dates before its extra complexity is justified.