Skip to content

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Latest commit

 

History

2 Commits

Folders and files

Repository files navigation

Lakehouse Framework

End-to-end, modular, reusable Data Lakehouse framework built on: Azure Databricks · Delta Lake · Unity Catalog · dbt · Azure DevOps · Azure Key Vault · Azure Monitor


Architecture

Source Systems  →  Bronze (Delta)  →  Silver (DLT + dbt)  →  Gold (dbt)
                   [pre-loaded]       [CDC / SCD1]            [Aggregations]

Three environments: Dev → UAT → Prod, each fully isolated with its own Unity Catalog, Key Vault scope, and Azure Monitor log scope.


Framework Layers

Layer Purpose Key Files
0 — Bundle DAB root config + env overrides databricks.yml, conf/
2 — Silver DLT APPLY CHANGES (SCD1) pipelines/silver/
3 — Gold dbt fact + dimension models dbt/models/gold/
4 — Orchestration Databricks Workflows pipelines/workflows/
5 — CI/CD Azure DevOps pipelines azure-devops/
6 — Security Key Vault integration libs/lakehouse_utils/keyvault_utils.py
7 — Observability Azure Monitor + audit table monitoring/, libs/lakehouse_utils/logging_utils.py
8 — Shared libs pip-installable Python package libs/lakehouse_utils/

Folder Structure

lakehouse-framework/
├── databricks.yml                        # DAB root bundle
├── conf/
│   ├── dev.yml                           # Dev env variables
│   ├── uat.yml                           # UAT env variables
│   └── prod.yml                          # Prod env variables
├── pipelines/
│   ├── silver/
│   │   ├── dlt_pipeline.yml              # DLT pipeline DAB config
│   │   ├── silver_transform.py           # DLT notebook (config-driven)
│   │   └── silver_config.yml             # Table definitions (add tables here)
│   └── workflows/
│       └── workflow_silver_gold.yml      # Databricks Workflow DAB config
├── dbt/
│   ├── dbt_project.yml
│   ├── profiles.yml
│   ├── packages.yml
│   ├── macros/helpers.sql
│   └── models/
│       ├── silver/
│       │   ├── sources.yml               # Silver DLT tables as dbt sources
│       │   ├── stg_orders.sql
│       │   └── stg_customers.sql
│       └── gold/
│           ├── schema.yml                # dbt tests (only DQ framework used)
│           ├── fact_orders.sql
│           ├── agg_orders_daily.sql
│           └── dim_customers.sql
├── libs/
│   ├── setup.py
│   └── lakehouse_utils/
│       ├── __init__.py
│       ├── spark_utils.py                # Delta merge, optimize, audit cols
│       ├── schema_utils.py               # YAML schema → PySpark / DDL / dbt
│       ├── keyvault_utils.py             # Key Vault secret retrieval
│       └── logging_utils.py             # Structured JSON + pipeline_audit table
├── scripts/
│   ├── optimize_tables.py                # Post-run OPTIMIZE maintenance
│   └── observability_bootstrap.py       # One-time env setup
├── monitoring/
│   └── alert_rules.yml                  # Azure Monitor alert rules as code
└── azure-devops/
    ├── deploy-dev.yml                    # CI on feature/* + develop
    ├── deploy-uat.yml                    # CD on main (+ manual approval)
    └── deploy-prod.yml                  # CD on release tag (+ 2-person approval)

Quick Start

1. Install the shared library

pip install -e libs/

2. Install dbt dependencies

cd dbt && dbt deps

3. Configure environment variables (local dev)

export DATABRICKS_HOST="https://<workspace>.azuredatabricks.net"
export DATABRICKS_TOKEN="<pat-token>"
export DATABRICKS_HTTP_PATH="/sql/1.0/warehouses/<warehouse-id>"
export DBT_TARGET="dev"
export DBT_CATALOG="lakehouse_dev"

4. Deploy to Dev via DAB

databricks bundle deploy --target dev

5. Run dbt (dev)

cd dbt
dbt run --target dev --vars '{"catalog": "lakehouse_dev", "env": "dev"}'
dbt test --target dev --vars '{"catalog": "lakehouse_dev", "env": "dev"}'

6. Bootstrap observability (once per env)

databricks bundle run observability_bootstrap --target dev

Adding a New Silver Table

  1. Add an entry to pipelines/silver/silver_config.yml:
- name: my_new_table
  source_table: my_new_table_raw      # must exist in {catalog}.bronze
  primary_keys:
    - my_primary_key
  sequence_by: _ingestion_timestamp
  partition_by:
    - partition_col                    # optional
  1. Add it as a dbt source in dbt/models/silver/sources.yml.

  2. Create a silver staging model in dbt/models/silver/stg_my_new_table.sql.

  3. Deploy: databricks bundle deploy --target dev

No Python code changes required.


Adding a New Gold Model

  1. Create dbt/models/gold/my_gold_model.sql using {{ ref() }} or {{ source() }}.
  2. Add tests to dbt/models/gold/schema.yml.
  3. Run: dbt run --select my_gold_model --target dev

Branching & Promotion Strategy

feature/*  →  develop  →  main  →  release/v*.*.*
   ↓              ↓          ↓            ↓
  Dev CI       Dev CD     UAT CD      Prod CD
  • Dev: auto-deploys on push to feature/* or develop
  • UAT: auto-deploys on merge to main (manual approval gate)
  • Prod: deploys on release tag v*.*.* (two-person approval)

Security Model

  • All secrets are stored in Azure Key Vault, one scope per environment
  • Databricks secret scopes are backed by Key Vault (no secrets in code or YAML)
  • Service principal credentials retrieved at runtime via KeyVaultUtils
  • No PAT tokens or connection strings committed to the repository

Observability

Asset Description
{catalog}.observability.pipeline_audit Structured run log for all pipelines
{catalog}.observability.dq_failures dbt test failure tracking
{catalog}.observability.v_pipeline_health Last-7-days health view
Azure Monitor alert rules Pipeline failures, DQ failures, staleness, slow runs

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages