Data platforms where pipelines fail loudly — not silently
I design and build cloud data platforms that survive partial failures, schema drift, and operational scale — incremental ingestion with checkpoint recovery, contract-driven data quality, medallion lakehouse patterns, and CI-validated platform automation on AWS, Azure, Databricks, and Snowflake.
Senior data engineering, to me, is judgment under constraint: intentional trade-offs, failure-aware design, and systems the next team can actually operate.
Design for failure.
Automate repeatable work.
Measure data quality.
Document decisions.
Keep systems understandable.
| Signal | How it shows up |
|---|---|
| System design | Architecture docs, ADRs, and clear boundaries between ingestion, transform, quality, and platform layers |
| Failure handling | Checkpoint recovery, idempotent loads, quality gates before bronze, operations runbooks |
| Operational proof | pytest, GitHub Actions CI, smoke tests, alert routing, and honest trade-off documentation |
| ⭐ Lakehouse |
lakehouse-platform-starter — Runnable reference architecture: Airflow + Cosmos · dbt · Iceberg · Trino · OpenLineage · Great Expectations
|
Connected layers — not isolated demo repos:
| Layer | Project | Focus |
|---|---|---|
| Flagship · Lakehouse | lakehouse-platform-starter | Iceberg + Trino + Cosmos dbt + Airflow + Marquez + GE · Docker stack · CI · hosted docs |
| Ingest & transform | production-data-pipeline | Incremental API · PostgreSQL bronze · dbt · Airflow · quarantine/DLQ · v0.3.0 |
| Quality & observability | data-quality-observability | YAML contracts · schema/freshness checks · run history · alerts |
| Platform & governance | cloud-lakehouse-blueprint | Medallion manifests · Terraform · IAM · lineage · CI validation |
| Domain | Technologies & practices |
|---|---|
| Data engineering | Python · SQL · incremental ingestion · ETL/ELT · API pipelines · checkpointing · idempotent loads |
| Data architecture | Medallion lakehouse · Iceberg · bronze/silver/gold · lineage · governance · cost modeling |
| Orchestration | Apache Airflow · Cosmos · dbt · Prefect · Spark |
| Cloud platforms | AWS · Azure · Databricks · Snowflake |
| Quality & observability | Data contracts · Great Expectations · OpenLineage · schema validation · CI/CD |
Production operations knowledge contributed upstream:
| Project | PR | Change |
|---|---|---|
| dbt docs | #9781 ✓ merged | Fusion telemetry: use duration_ms for slowest-nodes ranking (#9717) |
| dbt docs | #9960 | Macro arg types: bool not boolean (#9891) |
| dbt docs | #9961 | Clarify which behavior flags Fusion removes (#8972) |
| Meltano | #10253 ✓ merged | elt vs run decision guide for replication workloads (#6289) |
| Airflow | #71158 ✓ merged | Clarify metrics vs traces otel_* config options (#43366) |
| Airflow | #70171 | Surface dbt Cloud failure details in Airflow task logs (CI green) |
| Prefect | #22500 ✓ merged | Kubernetes readiness vs liveness probes |
| Prefect | #22533 | Global concurrency limit setup docs (re-review requested) |
| dbt docs | #9606 ✓ merged | Prefixed custom schema troubleshooting |
Since June 2026 — building a credible public data-engineering profile: portfolio releases, upstream merges, and technical writing in the open.
| Period | Highlights |
|---|---|
| Sep 2026 | Dev.to cover assets; GSC checklist; 7 upstream merges; Airflow #70171 CI green; dbt docs #9960 + #9961 |
| Aug 2026 | pipeline v0.3.0 · article #4 (contract versioning) · dbt + Meltano merges · dqo wired into Airflow DAG |
| Jul 2026 | lakehouse-platform-starter v1.0.0 · articles #1–#3 · Prefect + dbt docs merges |
| Jun 2026 | Portfolio stack started · 90-day public commit plan · pipeline v0.1.0 |
Full timeline: docs/work-history.md · Daily rhythm: contribution-log.md · Activity graph: github.com/br413
Since 2026 — I publish production-style data platform work in the open: portfolio repos, upstream contributions, and technical writing. Prior employer work lived in private GitLab/Azure DevOps; this GitHub profile is my public proof of craft, not a full career timeline.
| What you'll find here | Where |
|---|---|
| Lakehouse + pipeline + quality + platform stack | Pinned repos |
| Upstream OSS (Prefect, dbt, Airflow) | OSS table above |
| Architecture write-ups | Dev.to · br413.github.io |
| 90-day contribution plan | docs/90-day-contribution-plan.md |
| 90-day retrospective · next quarter | Discussion #34 · Q4 plan · Nov–Jan plan |
| Public work history (milestones) | docs/work-history.md |
| Daily activity log (automated) | docs/contribution-log.md |
Each flagship repo includes ADRs, pytest coverage, GitHub Actions CI, operations runbooks, and documented trade-offs — not toy demos.
Current focus (Sep 2026 – Jan 2027): Land Airflow #70171, Prefect #22533, and dbt docs #9960 / #9961; GSC indexing for articles #2–#4 and portfolio site. See nov-jan-contribution-plan.md.
| Article | Topic |
|---|---|
| Building a Production Data Pipeline with Incremental Loading and dbt | Incremental checkpoints, idempotent loads, medallion layering, Airflow orchestration, failure modes |
| Data Quality Contracts in Production Pipelines (Without a Separate Platform Team) | Row-level quarantine at ingestion, YAML dataset contracts, alert routing, CI enforcement |
| What I Learned Contributing to Prefect, dbt, and Airflow | Honest OSS retrospective — three merges, four open PRs, building in public |
| Contract Versioning in Production Pipelines | Registry → CLI → run history → CI guards — ADR 0002 stack |
More at br413.github.io · series: Cloud Data Platform Patterns
Open to senior data engineering roles, data platform architecture discussions, and technical collaboration.
Website: br413.github.io · Flagship: lakehouse-platform-starter · dbt docs: live
data engineering · lakehouse · dbt · Airflow · Iceberg · Trino · OpenLineage · Terraform



