All projects
PROJECT 01 / INCREMENTAL PIPELINES / SPORTS ANALYTICSWorking implementation

Rugby Analytics Lakehouse

Rugby results can change. A useful analytics pipeline needs to remember what changed, preserve the source and keep its models consistent. This lakehouse brings that thinking to United Rugby Championship data.

PythonSQLPySparkDatabricksDeltadbt
Explore the analytics Explore the source ↗
01

The question

A final score tells you who won. I wanted to follow the longer story: how do fixtures, teams, player scoring and team performance change across United Rugby Championship seasons?

The current dataset spans 2021–22 to 2025–26: five seasons and 755 completed matches, reconciled after a backfill and a source repair. These are season snapshots, so a result can arrive late or be corrected. Keeping that history is part of the question, too.

A little context before kick-off: team wins include playoffs, so these are team records, not official league standings. Player analysis covers lineups and listed scoring events, not a complete picture of player performance. A missing result means “result unavailable”, not necessarily an upcoming fixture.

READ THE DETAILSProject overview Dashboard definitions
02

The engineering

THE BIG PICTURE / HIGH-LEVEL PIPELINE

From the final whistle to a useful answer.

  1. SOURCESeason JSONRugby match feeds, saved as local snapshots
  2. ORCHESTRATEAirflow in DockerUpload selected file → run ingestion → run dbt
  3. LANDDatabricks VolumeScan the folder for new snapshots
DATABRICKS / DELTA TABLES
  1. PYSPARKBronzeKeep raw versions
  2. PYSPARKSilverModel fixtures & completed matches
  3. DBT COREGoldBuild analytics models & run quality tests
Dashboard queriesExplore curated Gold models
Website JSON exportVersioned snapshot → browser explorer

The job is to turn revisable source files into a consistent view of the game. Each layer has a specific responsibility: keep the evidence, model the matches, then make the results useful.

FOLLOW THE DATAOpen a stage for the detail
  1. LOCAL JSONStart with a seasonA snapshot of the source, ready to load.

    The input is a local season JSON file. It represents the season as the source knows it at that moment, rather than a stream of individual changes. Later results and corrections arrive as revised snapshots.

  2. AIRFLOW · DOCKERCoordinate the runUpload a file. Run ingestion. Build and test.

    The local, manually triggered DAG selects 2025–26 by default, with other seasons configurable. It validates and transforms the selected file into JSONL, uploads it, waits for the Databricks job, then runs dbt Core. A failed task stops the later steps.

  3. DATABRICKS · MANAGED VOLUMELand it, then discover what’s newOne selected upload; a scan of the whole folder.

    Landing filenames include the season, transform version and a content hash: a fingerprint of the snapshot. The Databricks job scans all matching files and checks the loaded-snapshot manifest. Known snapshots are skipped; changed content creates a new filename to ingest.

  4. BRONZE · PYSPARK + DELTAKeep the original evidencePreserve raw versions when a match changes.

    Match-level content hashes distinguish a correction from a repeat load. Identical content adds no new match version. A changed score or lineup creates a new version while the previous raw record stays in Bronze, making corrections traceable.

  5. SILVER · PYSPARK + DELTAGive fixtures and results their own placeA fixture can exist before its result does.

    Silver separates versioned fixtures from completed matches, player details and scoring events. Current views follow the latest published season snapshot, so late results and corrected records can replace the current view without erasing history. The snapshot manifest is published after its rows are written.

  6. GOLD · DBT COREBuild useful models. Check the numbers.Facts, dimensions and analytics with explicit grains.

    Nine dbt models turn current Silver data into reusable team, match and scoring views. The 46 tests check unique keys, relationships, fixture status, two team appearances per completed match and agreement between team and match score totals.

  7. DASHBOARD + WEBSITE EXPORTPut the work in front of someoneTwo views of curated Gold data.

    Read-only dashboard queries use Gold models. A separate exporter produces versioned JSON for the website, checks reconciliation and writes its manifest last so readers see a complete snapshot. The website reads that saved export, independently of Databricks. Refreshing the public snapshot remains a separate release step.

Repeat loads should be uneventful. Changed records should leave a trail.

READ THE DETAILSAirflow Databricks ingestion dbt models Website export
03

What works today

All five seasons have been ingested and reconciled in Databricks Free Edition. The documented Gold build contains nine dbt models and 46 passing tests, covering keys, relationships, fixture status, team appearances and score totals.

seasons reconciled
5
dbt models built
9
dbt tests passing
46

On 2 October 2026, a manually triggered local Airflow run in Docker uploaded the unchanged 2025–26 file, started the Databricks folder-ingestion job and completed dbt. The existing season snapshot was not duplicated.

That distinction matters: Airflow selects one local season file for upload by default; the Databricks job scans the whole landing folder for new snapshots. Dashboard queries and the website explorer then use saved exports of the curated Gold data.

Still to verify or deliver: a changed-source run through Airflow, automatic scheduling, Power BI reporting and automatic public website data refreshes. The successful orchestration run used an unchanged source file.

READ THE DETAILSVerified Airflow run Reconciliation & tests
04 / THE NEXT CHAPTER

Beyond my laptop

The next test is a less glamorous one: can it keep working when I’m not there to press Run? Here’s how I’d take this working project towards a production platform.

  1. Give the data a reliable way in

    Automate source delivery and move orchestration to a hosted service, with a deliberate refresh schedule, retries and checks for late or missing files.

  2. Make change safe to ship

    Add CI/CD checks and promote tested changes through development, staging and production. Use managed identities or service credentials in a secret store, with access scoped to each environment.

  3. Make problems easy to find

    Monitor freshness, quality and failed runs, with actionable alerts. Extend lineage from source snapshot through Gold to each export, so a surprising number has a traceable explanation.

  4. Plan for the awkward matchday

    Document and rehearse recovery: replay a snapshot, backfill a season and roll back a bad release. Publish a governed, versioned website feed only after quality checks and reuse rights are settled.

These are deliberate next steps, not implemented features. Public redistribution of the source rugby data still needs appropriate reuse permission, or a source licensed for that use.

READ THE DETAILSCurrent deployment boundary Public release boundary

Source: Project README and implementation status. Project descriptions reflect documented implementation, not independently rerun results.

05 / THE EXPLORER

The game, from
another angle.

SAVED DATA SNAPSHOT

Pick a season. Follow a team.
See what the scoreline leaves out.

The explorer loads as you reach it.

KEEP EXPLORING / WORKING IMPLEMENTATIONCentra vs SPARBack to selected work