Job Description
We are seeking a Data Science Senior Associate focused on building and operating resilient datasets, pipelines, and reusable metrics that support hypothesis-driven analyses and experiments across the product development lifecycle (PDLC)
In this role, you will be hands-on in designing, developing, and maintaining data products that are reliable, observable, and well-documented-enabling partners across product, engineering, and analytics to measure what’s driving value, where friction exists, and how operating-model changes impact outcomes as teams adopt more agentic ways of working. You’ll contribute to engineering standards and help raise the quality bar through strong delivery and collaboration.
Job Responsibilities
- Build and operate scalable batch/streaming pipelines with SLAs, monitoring, and incident response participation (as needed).
- Create and maintain trusted data products (dimensions, event models, marts) with clear ownership and documentation.
- Deliver metrics and feature-ready datasets for AI adoption/productivity measurement; manage definition changes over time.
- Implement data quality and governance controls (validation, reconciliation, lineage, access, retention, auditability).
- Orchestrate workflows in Airflow (or equivalent), including backfills and retries.
- Model/transform data using SQL and dbt (or equivalent) for trusted reporting and repeatable measurement.
- Write production-grade Python/PySpark with testing, performance tuning, and maintainable design.
- Partner with cross-functional stakeholders to define requirements, success criteria, and metric interpretation across finance, PDLC/SDLC, and AI tool logs.
- Contribute to engineering best practices (version control, code review, CI/CD, runbooks) and improve observability and cost/performance.
- Mentor peers through reviews, documentation, and knowledge sharing (no formal people management).
Required Qualifications
- Bachelor’s degree in Computer Science, Engineering, or equivalent practical experience.
- 3+ years building production data solutions; strong ownership and delivery.
- Strong engineering fundamentals (OOP, testing, development lifecycle).
- Strong data modeling skills (dimensional, normalized, event-based).
- Experience with Databricks and/or Spark/PySpark.
- Strong SQL; experience with dbt (or equivalent) and building testable data codebases.
- Experience operating orchestration pipelines (Airflow or equivalent).
- Proven ability to build and maintain reliable metrics as sources/definitions evolve.
- Effective delivery in ambiguous, multi-stakeholder environments.
Preferred Qualifications
- Experience with modern lakehouse/warehouse patterns and broader cloud data platforms (e.g., Databricks, Snowflake).
- Experience with BI/semantic layers and metrics management practices.
- Exposure to experimentation or hypothesis-driven analytics approaches (e.g., measurement design to support tests, rollouts, and pre/post evaluation); deep causal specialization not required.
- Experience improving observability (data freshness/SLA monitoring, lineage, alerting) and contributing to operational maturity (runbooks, incident follow-ups).

