659,077open jobs
38,403companies
95,422added this week
Browse all
Salary
$110k – $135k per year
Location
In office (New York)
Seniority
Middle · 3+ years exp
Employment
Full-Time
Overview
Company
Impact
Profile match

H1

H1 is a healthcare data company headquartered in New York City and founded in 2017. The company maintains a global database of healthcare professionals, their publications, trial history, and affiliations, which pharmaceutical companies use for clinical trial site selection, medical affairs, and commercial targeting. It sells to life sciences organizations and combines licensed and public data sources into a single provider graph.

At H1, we believe access to the best healthcare information is a basic human right. Our mission is to provide a platform that can optimally inform every doctor interaction globally. This promotes health equity and builds needed trust in healthcare systems. To accomplish this, our teams harness the power of data and AI-technology to unlock groundbreaking medical insights and convert those insights into action that result in optimal patient outcomes and accelerates an equitable and inclusive drug development lifecycle. Visit h1.com to learn more about us.

H1's Data Network (H1DN) team is the client-data mastering network at the core of how H1's products get their data. We run production ingestion for major enterprise customers. Clinical trial data is one of our highest-visibility streams: it feeds decisions about where trials run and who runs them, and the people who depend on it are as often clinical experts as they are engineers. SLAs and customer expectations drive how we work, and we're looking for engineers who are energized by that.

WHAT YOU'LL DO AT H1

As a Data Engineer II on the H1DN team, you will build and operate the pipelines behind H1's clinical trials data. You'll work primarily in Python, PySpark, and SQL, and you'll work directly with clinical subject matter experts and Customer Success Managers to turn their domain knowledge into pipeline logic that holds up in production.

You will:

- Build and maintain the Python and PySpark pipelines behind the CTMS trial data pipeline intake workflows, including scoring and status logic.

- Develop the transformation logic that maps raw trial and customer data to H1's internal data models, handling diverse source formats including CSV, JSON, Parquet, and APIs.

- Write and tune SQL against large datasets to investigate data questions, validate pipeline output, and support analysis that clinical SMEs and customer-facing teams depend on.

- Turn around customer-driven changes quickly, scoping requests as they arrive, shipping changes that hold up under enterprise SLAs, and reworking logic as customer needs shift mid-flight.

- Partner with clinical SMEs to translate domain expertise into concrete data rules, then walk them through the results, explain what the pipeline did and why, and fold their feedback back into the logic.

- Build the data quality checks, validation logic, and reconciliation that let non-engineers trust pipeline output without reading the code.

- Participate in code reviews, maintaining a high bar for quality and adherence to engineering standards.

- Monitor and improve pipeline observability, contributing to alerting and dashboards that surface job health and data anomalies for both the team and internal users.

ABOUT YOU

You are a data engineer with a strong Python foundation and real distributed-processing experience. You're drawn to high-impact teams where the work is tangible: pipelines running, enterprise customers getting their data on time, clinical data that people make real decisions from. You're comfortable in an environment where recurring production runs and customer SLAs shape day-to-day priorities, and where a customer request can reorder your week. You'd rather sit down with a domain expert and understand why the data looks the way it does than build to a spec handed to you secondhand.

You bring experience:

- Building and shipping production data pipelines in Python, with an understanding of what makes them reliable and maintainable under real load

- Working with PySpark or a comparable distributed processing framework on datasets too large for a single machine

- Writing SQL well enough to answer hard questions about data, not just retrieve it

- Working in an operationally-driven environment where reliability and on-time delivery matter as much as new feature work

- Working directly with non-engineering partners, subject matter experts, analysts, or customer-facing teams, and communicating clearly about data with people who don't read code

- Holding a high bar in code review and expecting the same from those who review your work

- Identifying data problems early and seeing work through to resolution rather than handing it off

REQUIREMENTS

- 3+ years of experience in software or data engineering, with meaningful Python in your background

- Demonstrated experience building and maintaining production-grade data pipelines in Python

- Hands-on experience with PySpark or a similar distributed data processing framework

- Strong SQL skills, including working with large, messy, multi-source datasets

- Strong understanding of software quality practices: testing, code review, documentation, and CI/CD

- Experience working with cross-functional and non-technical stakeholders

- Experience with pipeline orchestration tooling (Argo, Airflow, Databricks, dbt, or similar) preferred

- Familiarity with clinical trial data, healthcare data, or another regulated data domain a plus

- Familiarity with entity matching or data mastering a plus

- Familiarity with AWS services (S3, Lambda, ECS, or similar) a plus

COMPENSATION

This role pays $110,000 to $135,000 per year, based on experience, in addition to stock options.

Anticipated role close date: 10/20/2026

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
659,077 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
New York
$195k – $361k per year • Equity • Remote • Full-Time • 7+ years exp • PhD • United States
Python
SQL
Python
pySpark
Databases
Snowflake
Databricks
Google BigQuery
BigQuery
AI/ML
Spark
dbt
Prefect
Feast
Anomaly Detection
Time Series Forecasting
Amazon SageMaker
Feature Store
Edge AI
DevOps
CI/CD
Self-Healing
Cybersecurity
GDPR
HIPAA
Apply
$36k – $92k per year (Estimated) • Remote/Hybrid • Full-Time • 5+ years exp • Bachelor's Degree • Athens
Python
SQL
C#
C#
.NET
Databases
Databricks
Microsoft Fabric
Mobile
Clean Architecture
DevOps
Rest API
CI/CD
Git
Analytics
ETL/ELT
Azure Data Factory
Management
Power Automate
Power Apps
Apply
$88k – $118k per year • Remote/Hybrid • Full-Time • 5+ years exp • Bachelor's Degree • The Hague
Python
AI/ML
RAG
EU AI Act
DevOps
Azure
CI/CD
Git
Management
WhatsApp
Apply
$97k – $209k per year (Estimated) • Remote/Hybrid • Full-Time • The Hague
Python
DevOps
AWS CDK
CI/CD
AWS
Kubernetes
Apply
$88k – $196k per year (Estimated) • Remote/Hybrid • Full-Time • United Kingdom
DevOps
Azure DevOps
Azure
CI/CD
Bicep
Cybersecurity
Least Privilege
Apply
$105k – $135k per year • Equity • Remote/Hybrid • Full-Time • 3+ years exp • New York
Python
SQL
Apply
$190k – $265k per year • Equity • In office • Full-Time • 10+ years exp • New York
AI/ML
Agentic Workflows
Apply
$125k – $160k per year • In office • Full-Time • 2+ years exp • Bachelor's Degree • New York
AI/ML
Claude
ChatGPT
Apply
$190k – $220k per year • Equity • In office • Full-Time • 10+ years exp • New York
SQL
AI/ML
Claude
dbt
OpenAI
DevOps
SLI/SLO/SLA
Apply
$130k – $160k per year • In office • Full-Time • 5+ years exp • New York
Apply
$161k – $296k per year (Estimated) • In office • Full-Time • 10+ years exp • Bachelor's Degree • New York • Madison
Python
Java
Apply
$60k per year • Remote/Hybrid • Full-Time • 3+ years exp • Bachelor's Degree • New York
Management
Google Workspace
Apply
$172k – $346k per year (Estimated) • In office • Full-Time • New York
AI/ML
AI Agents
Management
Notion
Apply
$195k – $370k per year (Estimated) • In office • Full-Time • New York • Philadelphia
Apply
$96k – $153k per year • In office • Full-Time • 3+ years exp • Bachelor's Degree • Milpitas • Charlotte • Raleigh • Tampa • Baltimore
Apply
See all jobs
This is one of many
659,077 more open roles from verified company boards, updated every day.