368,657open jobs
9,442companies
50,883added this week
Browse all
Salary
$190k – $220k per year
Location
Remote/Hybrid (New York, United States)
Seniority
Staff · 8+ years exp
Employment
Full-Time
Overview
Company
Impact
Profile match

H1

H1 is a healthcare data company headquartered in New York City and founded in 2017. The company maintains a global database of healthcare professionals, their publications, trial history, and affiliations, which pharmaceutical companies use for clinical trial site selection, medical affairs, and commercial targeting. It sells to life sciences organizations and combines licensed and public data sources into a single provider graph.

At H1, we believe access to the best healthcare information is a basic human right. Our mission is to provide a platform that can optimally inform every doctor interaction globally. This promotes health equity and builds needed trust in healthcare systems. To accomplish this, our teams harness the power of data and AI technology to unlock groundbreaking medical insights and convert those insights into actions that result in optimal patient outcomes and accelerate an equitable and inclusive drug development lifecycle. Visit h1.co to learn more about us.

Data Engineering is responsible for the development and delivery of our most important asset-our data. With thousands of data sources from around the world, the team ensures that data is accurate, normalized, and delivered at a velocity that keeps up with real-world changes. As we expand our markets and the scope of data we provide to our customers, our team must scale to meet that demand.

WHAT YOU'LL DO AT H1

As a Staff Data Engineer on the Data Lake team at H1, you will play a critical role in shaping the architecture, scalability, reliability, and long-term direction of our core data platform. This role is designed for a highly technical engineer who is excited to grow into an Engineering Manager track while remaining deeply hands-on technically.

The Data Lake is the foundation of H1’s platform, responsible for the validation, accuracy, standardization, and quality of the data powering every downstream product and team across the organization. You will help lead the evolution of this platform while supporting and mentoring a growing team of engineers.

You will:

- Architect, build, and scale distributed ETL/ELT pipelines and large-scale ingestion frameworks across structured and unstructured healthcare datasets.

- Lead the evolution of H1’s Data Lake architecture with a focus on scalability, observability, reliability, and cost optimization.

- Own and improve data quality, validation, normalization, and standardization workflows across thousands of global data sources.

- Design and optimize batch and near real-time data processing frameworks using cloud-native distributed systems.

- Optimize distributed compute and storage systems, including Spark workloads, query performance, partitioning strategies, and infrastructure efficiency.

- Drive improvements in monitoring, governance, operational excellence, and production reliability across the platform.

- Troubleshoot complex production data and infrastructure issues across distributed systems.

- Partner closely with Product, Infrastructure, Security, Compliance, and downstream engineering teams to support scalable and secure data delivery.

- Mentor engineers through technical leadership, architecture reviews, and engineering best practices.

- Help define technical roadmap priorities and contribute to long-term platform strategy and execution planning.

- Support production operations, incident response, and platform health as part of overall ownership of the Data Lake ecosystem.

ABOUT YOU

You are a highly technical data engineer who thrives in lean, high-ownership environments and enjoys solving complex distributed systems challenges. You are excited by the opportunity to influence technical direction, mentor engineers, and grow into broader engineering leadership responsibilities while remaining hands-on.

- You have deep experience designing and scaling distributed data platforms and large-scale pipelines in cloud-native environments.

- You excel at building reliable, observable, and maintainable data systems supporting critical business and analytics workloads.

- You have strong expertise in distributed processing, performance optimization, and modern data architecture patterns.

- You are comfortable leading technical initiatives and influencing architecture decisions across teams.

- You communicate effectively with both technical and non-technical stakeholders.

- You enjoy mentoring engineers and helping raise the engineering bar across teams.

- You are energized by ownership, autonomy, and solving ambiguous technical challenges.

REQUIREMENTS

- 8+ years of experience in data engineering, software engineering, or related fields with significant experience building and scaling distributed data platforms.

- Demonstrated technical leadership experience with interest in or experience mentoring and leading engineers.

- Strong proficiency in Python (PySpark), Java, Scala, or similar programming languages.

Advanced SQL expertise, including performance tuning and optimization across large datasets.

- Deep experience with Apache Spark and cloud-native big data platforms, preferably within AWS environments (EMR, Glue, S3, Athena, Redshift, or similar).

- Experience designing and scaling modern cloud-native data lake architectures and large-scale ingestion frameworks.

- Experience with orchestration and workflow management tools such as Argo, Airflow, or similar technologies.

- Strong understanding of distributed storage systems, partitioning strategies, and file formats such as Parquet, Avro, and ORC.

- Experience with Docker, Kubernetes, and modern containerization technologies.

- Experience implementing monitoring, observability, and data quality frameworks within production environments.

- Experience with large-scale data cleaning, parsing, normalization, and validation workflows preferred.

- Experience working with healthcare, life sciences, publication, or large-scale entity-resolution datasets preferred.

- Exposure to ML/AI-driven data enrichment, parsing, or validation workflows is a plus.

- Experience using AI-assisted coding tools (e.g., GitHub Copilot, Claude Code) to accelerate development while maintaining quality is encouraged

COMPENSATION

This role pays $190,000 to $220,000 per year, based on experience, in addition to stock options.

Anticipated role close date: 9/15/2026

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
368,657 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
New York
$84k – $178k per year (Estimated) • In office • Full-Time • 10+ years exp • Wellington
Java
Python
DevOps
Ansible
AWS
Azure
CI/CD
Docker
GCP
Helm
Kubernetes
Platform Engineering
Prometheus
Service Mesh
Terraform
GitLab
IAM
Apply
Platform Engineer 1 day ago
$87k – $140k per year • In office • Full-Time • 3+ years exp • Berlin
Databases
PostgreSQL
Redis
DevOps
AWS
Azure
Bicep
CI/CD
Docker
GCP
GitHub Actions
Kubernetes
OpenShift
Terraform
GitHub
Apply
Founding Engineer 1 day ago
$81k – $116k per year • In office • Full-Time • Bachelor's Degree • Munich
JavaScript
Python
TypeScript
Databases
MySQL
PostgreSQL
Frontend
Next.js
React.js
Tailwind CSS
DevOps
AWS
Azure
CI/CD
Docker
GCP
Grafana
Kubernetes
OpenTelemetry
Prometheus
Apply
Staff Engineer 1 day ago
$105k – $233k per year • In office • Full-Time • 8+ years exp • Munich
Python
TypeScript
JavaScript
AI/ML
LLM
Frontend
React.js
DevOps
AWS
Azure
Docker
GCP
Terraform
Apply
$98k – $195k per year (Estimated) • In office • Full-Time • 7+ years exp • Wellington
Java
Python
SQL
Java
Spring Boot
Databases
Apache Kafka
Databricks
Neo4j
AI/ML
Flink
Spark
Frontend
GraphQL
DevOps
Azure
CI/CD
Datadog
Dynatrace
Kibana
Kubernetes
OpenShift
Platform Engineering
Splunk
Amazon ECS
Apply
$190k – $220k per year • Equity • Remote/Hybrid • Full-Time • 10+ years exp • New York
SQL
AI/ML
Claude
dbt
OpenAI
DevOps
SLI/SLO/SLA
Apply
$28k – $68k per year (Estimated) • Remote • Full-Time • 6+ years exp
Go
Node JS
Python
TypeScript
JavaScript
Node JS
Nest.JS
Python
pySpark
Databases
Apache Kafka
ElasticSearch
PostgreSQL
AI/ML
LangChain
LLM
RAG
Spark
Anthropic
OpenAI
Frontend
React.js
DevOps
AWS
CI/CD
Helm
Kubernetes
Vercel
Apply
$130k – $160k per year • In office • Full-Time • 5+ years exp • New York
Apply
$27k – $67k per year (Estimated) • Remote • Full-Time
Node JS
Python
JavaScript
Databases
ElasticSearch
PostgreSQL
AI/ML
Airflow
Claude
Copilot
Knowledge Graph
DevOps
CI/CD
CircleCI
Docker
Git
Kubernetes
Terraform
Analytics
ETL/ELT
Management
Jira
Apply
$170k – $190k per year • Equity • Remote • 8+ years exp • Bachelor's Degree
C++
Go
Java
Python
Scala
Databases
Apache Kafka
AI/ML
Spark
DevOps
AWS
Apply
$197k – $374k per year (Estimated) • In office • Full-Time • 8+ years exp • PhD • New York
AI/ML
AI Agents
Claude
LangChain
OpenAI
Vertex AI
Management
n8n
Zapier
Apply
Senior Data Analyst 2 hours ago
$86k – $171k per year (Estimated) • Equity • Remote • Full-Time • 5+ years exp • Bachelor's Degree • New York
Python
SQL
Databases
Snowflake
AI/ML
Anomaly Detection
Claude
Copilot
Cursor
dbt
Edge AI
DevOps
AWS
Analytics
Tableau
Marketing
Salesforce
Apply
$96k – $134k per year • Remote/Hybrid • Full-Time • Bachelor's Degree • New York
JavaScript
Swift
TypeScript
Java
Java
Spring Framework
Databases
Apache Kafka
PostgreSQL
AI/ML
AI Agents
Claude
Copilot
Fine-tuning
Flink
LangChain
LangGraph
Llama
LlamaIndex
Prompt Engineering
PyTorch
RAG
TensorFlow
Transformers
Devin
Hugging Face
OpenAI
Frontend
Angular
React.js
Mobile
MVC
DevOps
AWS
CI/CD
Docker
Kubernetes
OpenShift
Splunk
Vector
GitHub
Analytics
Tableau
Apply
$150k – $180k per year • In office • Full-Time • PhD • New York
Python
AI/ML
Anthropic
Anthropic SDK
Computer Vision
Fine-tuning
LangChain
LlamaIndex
LLM
OpenAI
OpenAI SDK
RAG
DevOps
AWS
Azure
GCP
Apply
Senior AI Architect 2 hours ago
$131k – $136k per year • In office • Full-Time • 4+ years exp • Master's Degree • New York
Python
Databases
Databricks
AI/ML
Anthropic
Computer Vision
EU AI Act
LLMOps
OpenAI
DevOps
AWS
Azure
GCP
Terraform
Cybersecurity
GDPR
Apply
See all jobs
This is one of many
368,657 more open roles from verified company boards, updated every day.