Automated Cloud Forecast Pipeline

Automated Cloud Forecast Pipeline

Project Type
🟧 Predictive MLOps Automation
Tools
PythonGoogle Cloud PlatformTerraformPower BIBigQuery-MLDataform
Description

A production-grade, automated serverless ML forecasting pipeline with trust-gated predictions and an explainable revenue model.

GitHub Project Repository: automated-predictive-pipeline

The Bottleneck

Marketing budget decisions ran without attribution. While historical revenue shows which channels and products have more revenue, it doesn't show which ones drive it. Nothing in the data distinguished the two.

  • Budget allocation followed historical revenue size, with no way to tell which channels or products actually drove it.
  • No deterministic, reproducible method existed to evaluate a product or channel before committing spend.

The Solution

A trust-gated forecast that links weekly marketing spend to explained revenue drivers, so allocation decisions are backed by attribution instead of momentum.

image

Pipeline Design: Medallion to ML

The Dataform pipeline transforms external and core data into ML-ready datasets through a medallion architecture (source → contract → ml_data → ml_model → published). Each layer enforces structural guarantees before data advances, and the model lifecycle runs on a fixed frequency with a trust gate at every stage.

image
  • Contract assertions: malformed data is blocked before feature engineering.
  • Data-loss gate: row loss above 10% aborts the run before predictions are published.
  • Grain enforcement: one row per (week_start, traffic source, product category).
  • Monthly retrain: the model rebuilds from scratch each cycle on pre-2026 training data.
  • Weekly predict: next-week revenue forecasts published every week.
  • Holdout evaluate: performance is checked against weeks the model never saw (R², RMSE, MAE, MedAE) before the forecast is used.

Forecast Decision Support

The dashboard follows a Trust → Explain → Forecast sequence: the reader verifies model health, understands driver contributions, and only then sees the live prediction.

image
  • Trust (model confidence): performance metrics, residual distribution, data-loss headroom, and source freshness must all pass before proceeding. Decision makers see data quality first, which removes "blame the data" as a reflex when a number disagrees with expectation.
  • Explain (revenue drivers): features ranked by relative impact; numeric weights show dollar impact per unit, categorical weights show which levels lift or drag revenue.
  • Forecast (live prediction): next-week revenue with lower and upper bounds, broken down by the same categories.

What This Unblocks

Marketing leadership makes weekly allocation decisions backed by attribution, with no manual forecast work and no bad number slipping through unchecked.

  • Attribution-backed allocation: budget reallocates across channels and products weekly, backed by attribution instead of momentum.
  • Evidence-based product decisions: products and channels get prioritized, deprioritized, or discontinued with evidence.
  • Trust before acting: quality gates show before insights, so a forecast disagreement with expectation can't be waved off as a data problem.
  • Zero manual forecast work: the forecast lands every Monday on an automated cadence.
  • Traceable reasoning: forecast, drivers, and quality gates live in one system.

Codified Observability

The system health suite is defined as Terraform code. pipeline_health tracks row loss and data-quality metrics shown on the dashboard's Trust page. CRITICAL email alerts cover every orchestration layer (scheduler, extractor, workflow, ML pipeline).

image

Evidence and Artifacts

SuperMade with Super