Job Description
We are seeking a Data Engineering Lead to help build and evolve a high-quality measurement data foundation that enables trusted analytics and decision-making at scale. This role focuses on designing and delivering resilient datasets, pipelines, and reusable metrics that support hypothesis-driven analyses and experiments across the product development lifecycle (PDLC).
You’ll be hands-on where needed, drive engineering standards, and help teams move faster by improving reliability, observability, and usability across the data lifecycle-so leaders can clearly see what’s driving value, what’s creating friction, and what operating-model shifts materially improve outcomes as teams become more agentic.
Job Responsibilities
- Design, build, and operate scalable data pipelines (batch and/or streaming) with clear SLAs, monitoring, and incident response practices.
- Develop and curate trusted data products (e.g., conformed dimensions, event models, marts) with strong documentation and clear ownership.
- Build and maintain well-defined metrics and feature-ready datasets that enable measurement of AI adoption and productivity outcomes (e.g., reusable aggregates, cohorting, time-windowed measures), including change control as definitions evolve.
- Drive data quality and governance through validations, reconciliations, lineage, access controls, retention, and auditability aligned to requirements.
- Develop and operate workflow orchestration (e.g., Apache Airflow) to schedule, monitor, and manage data movement and transformations.
- Model and transform data for analytics using SQL/dbt to support trusted reporting and repeatable measurement.
- Write production-grade Python/PySpark with disciplined testing, performance tuning, and maintainable design.
- Partner with analytics, product, and engineering stakeholders to define requirements, success criteria, and consistent interpretation of key measures-particularly where inputs span finance business cases, PDLC/SDLC tools, and AI tool logs.
- Establish and enforce engineering best practices (version control, code review, testing strategy, deployment processes, runbooks) and continuously improve observability and cost/performance (freshness, completeness, timeliness, scalability, spend).
- Mentor and develop a team of 2, influencing technical direction through standards, reviews, and knowledge sharing.
Required Qualifications
- Bachelor’s degree in Computer Science, Engineering, or equivalent practical experience.
- 5+ years of hands-on experience delivering production data solutions in a fast-paced engineering environment (actively coding and owning outcomes).
- Strong software engineering fundamentals (system design, data structures, object-oriented programming, testing strategies, and end-to-end development lifecycle).
- Strong understanding of data modeling (conceptual, logical, physical), including dimensional, normalized, and event-based approaches.
- Hands-on experience with Databricks and large-scale distributed data processing/performance tuning (Spark/PySpark).
- Strong SQL skills and experience with modern transformation tooling (e.g., dbt), including building maintainable, testable data codebases.
- Experience designing and operating orchestration pipelines using Airflow (or equivalent), including backfills, retries, and operational monitoring.
- Demonstrated rigor building and maintaining trusted metrics (definitions, edge cases, validation/testing, documentation) and keeping them reliable as upstream sources change.
- Demonstrated ability to lead delivery in complex environments with multiple stakeholders and ambiguous requirements.
Preferred Qualifications
- Experience with modern lakehouse/warehouse patterns and broader cloud data platforms (e.g., Databricks, Snowflake).
- Experience with BI/semantic layers and metrics management practices.
- Exposure to experimentation or hypothesis-driven analytics approaches (e.g., measurement design to support tests, rollouts, and pre/post evaluation); deep causal specialization not required.

