Protective is looking for Data Engineers to design, build, and operate production data pipelines on Voyager, our Databricks lakehouse on Azure. You will own the end-to-end flow of data through a medallion architecture - ingesting source data into the Bronze layer, applying cleansing, validation, and conformance in Silver, and publishing trusted, consumption-ready data products in Gold.
This role sits at the intersection of engineering, analytics, and platform operations. You will write production-grade Python and SQL, enforce data contracts, and help Protective treat data as a product with named owners and real consumers. You will be embedded on a delivery pod that owns its data products end to end, rather than servicing tickets from a queue.
On Voyager, the medallion layers are named Raw, Prep, and Prod. They map directly to Bronze, Silver, and Gold and are used interchangeably in this description.
Key Responsibilities
- Design, develop, and maintain production data pipelines on Databricks using Python, SQL, Apache Spark, and Delta Lake.
- Build Bronze-layer ingestion that reliably captures data from APIs, relational databases, flat files, cloud storage, and SaaS platforms - using dlt (dltHub) and Databricks-native ingestion where each fits - including incremental loading, pagination, watermarking, state management, and replay after failure.
- Develop Silver-layer transformations in dbt and Python over Delta Lake that cleanse, standardize, type, deduplicate, validate, conform, and enrich data so that it is reusable across domains. A meaningful share of this role is making messy source data trustworthy.
- Create Gold-layer data products: dimensional models, slowly changing dimensions, fact and bridge tables, aggregates, and serving tables aligned to how consumers actually query.
- Produce and maintain the curated datasets ML engineering trains and serves models from - feature and training tables that are versioned and reproducible, not one-off extracts.
- Author and maintain data contracts using the Open Data Contract Standard (ODCS) - schema with real semantics, named owner, known consumers, quality rules, and freshness expectations - and assess backward compatibility before every change.
- Implement data quality as code: uniqueness and not-null on keys at minimum, plus referential, accepted-value, freshness, and custom business-rule tests, surfaced to producers and consumers rather than buried in logs.
- Orchestrate ingestion and transformation as assets in Dagster, deployed to Dagster Cloud, and operate what you build across development, branch, and production deployments - schedules and sensors, asset dependencies, backfills, and run observability.
- Apply governance through Unity Catalog - catalogs, schemas, external locations, grants, row- and column-level security, and lineage - and handle credentials through Azure Key Vault rather than in code.
- Implement incremental and merge-based processing with Delta Lake (MERGE, schema evolution, time travel, OPTIMIZE) and tune Spark jobs, table layouts, and compute for performance and cost.
- Troubleshoot production failures, data-quality issues, source-system changes, and late-arriving or duplicate data - including backfills and recovery - and take part in the pod’s on-call rotation for the pipelines it owns, with root-cause analysis that closes the gap rather than reopening the ticket.
- Build and maintain CI/CD for data assets in Azure DevOps - automated tests and CI checks on dlt, dbt, and Dagster changes, promotion from development through branch deployments to production, and releases that are repeatable and auditable.
- Instrument what you own for observability: freshness, volume, quality, latency, and cost, with alerting tied to the SLAs and SLOs your contract commits to instead of depending on someone noticing.
- Work inside the platform’s control expectations - least-privilege access, secrets in Azure Key Vault, change management through pull request and pipeline, and audit evidence that falls out of the deployment path rather than being reconstructed later.
- Participate in code review and document architecture, runbooks, and data products so others can discover, trust, and reuse them.
- Work with data architects, analysts, product owners, and business stakeholders to translate requirements into maintainable data solutions.
Qualifications
- Bachelor’s degree in Computer Science, Information Systems, Engineering, or a related field; equivalent practical experience considered.
- 3+ years building and supporting production data pipelines in a cloud data platform environment.
- Strong hands-on Python and SQL. Both are used daily and neither substitutes for the other.
- Hands-on experience with Databricks or a comparable Spark-based lakehouse, including Delta Lake tables, MERGE, and incremental load patterns.
- Practical understanding of medallion / multi-layer lakehouse design, and the judgment to say what belongs in Bronze versus Silver versus Gold.
- Experience ingesting data from APIs, relational databases, files, or SaaS applications, including the incremental and state-management problems that come with it.
- Working knowledge of dimensional modeling - grain, keys, facts and dimensions, slowly changing dimensions - and of ELT design patterns and data quality practice.
- Experience with orchestration and scheduling using Dagster, Databricks Workflows, Airflow, Azure Data Factory, or similar.
- Git-based source control, pull request review, automated testing, and CI/CD as normal practice - Azure DevOps or comparable.
- Experience troubleshooting production data failures, performance bottlenecks, and source-system changes.
- Experience with pipeline monitoring and alerting, and a working understanding of what a freshness or quality SLA means once real consumers depend on it.
- Ability to explain technical designs and trade-offs to both technical and non-technical partners.
- Databricks certification (Data Engineer Associate or Professional) or equivalent demonstrated depth.
- Unity Catalog experience: catalogs, schemas, volumes, external locations, storage credentials, permissions, and lineage.
- dbt on Databricks, or another transformation framework used alongside Spark.
- Python-based modeling frameworks over Delta Lake, and experience implementing Type 2 history, surrogate keys, and merge strategies in code.
- Experience with a declarative Python ingestion framework such as dlt (dltHub), Airbyte, Meltano, or Fivetran.
- Dagster experience specifically, including assets, asset checks, sensors, schedules, and branch deployments.
Required Qualifications
Preferred Qualifications

