368,634open jobs
9,437companies
50,578added this week
Browse all
Salary
$150k – $200k per year
Location
Remote (United States)
Seniority
Senior · 20+ years exp
Employment
Full-Time
Overview
Company
Impact
Profile match
novellia is the first and only company that lets anyone in the us gain access to nearly a decade of their health data in under 30 seconds, 100% free. all your health records, across doctors, all in one place, always up to date. we work with the...

Our Story

Since 2023, our mission has been clear: to be the north star of patient equity.

Every day, we strive to bridge the gaps in healthcare access and outcomes to ensure that every patient, regardless of background or circumstance, receives the care they deserve. As a member of our team, you'll be at the forefront of innovation, working alongside passionate individuals who share your dedication to creating best-in-class, patient-centric products in healthcare. Together, we're revolutionizing the way people understand their health and working with the world's top researchers to accelerate innovation.

About Novellia

Novellia is the first and only company that lets anyone in the U.S. gain access to nearly a decade of their health data in under 30 seconds - 100% free. All your health records, across every doctor, in one place, always up to date.

We are the only patient-powered real-world data platform delivering comprehensive, patient-authorized longitudinal health insights to accelerate biopharma innovation. Unlike traditional RWD providers who deliver fragmented institutional data, we empower patients to access 20+ years of their health records, then transform these complete health journeys into fit-for-purpose datasets for evidence generation, regulatory submissions, and market access. We are growing 5x year over year, have raised close to $30M in funding, and are backed by tier-1 investors including Spark Capital, Khosla Ventures, and Bling Capital.

Working with the world's top researchers, we turn health insights into life-changing action for millions of people around the world.

About the role

Most of what matters in a health record isn't in a structured field - it's in the note, the discharge summary, the pathology report, the scanned fax. Turning that unstructured clinical text into trustworthy, structured features is what makes a longitudinal health history usable for research, and it's one of the highest-leverage capabilities Novellia can own.

You'll be our first ML hire, joining Platform Engineering and reporting to the Head of Platform Engineering, as technical owner of this multi-quarter effort. The interesting decisions are still open - what we extract first, how we know we're right, what a mature extraction pipeline looks like at our scale. There's no existing approach to inherit or defend.

The work draws on two toolkits. Roughly 70% is applied ML on clinical text: entity extraction, classification, sequence labelling, annotation strategy, error analysis, calibration, and the evaluation discipline that tells you whether your numbers mean anything. Roughly 30% is LLM-based: prompt development, structured output, retrieval, and the evals and observability that keep generative approaches honest. Deciding which approach a given problem calls for is the most interesting part of the job, and that call is yours.

We're looking for a leader in this seat: setting technical direction rather than waiting to be handed a problem. If this grows the way we think it will, leading the team we build around it is on the table.

What you'll do

  • Own the full lifecycle of extraction models - framing, data/annotation strategy, model selection, training/fine-tuning, evaluation, deployment, monitoring, retraining. Not a research seat, not a hand-off seat.

  • Define what "accurate enough" means with clinical and customer-facing stakeholders, and build the evaluation harness that makes the answer defensible - the first deliverable, not a follow-up.

  • Partner with Clinical Data Managers on curation design and own the technical half of QA/QC alongside them: which variables are extractable, how an instruction becomes a model spec, and the tooling/sampling/error analysis behind human-in-the-loop review.

  • Build clinical NLP pipelines against messy real-world data and work with backend engineers to productionize what you build.

  • Use LLMs with the same rigor you'd apply anywhere: versioned prompts, real evals, tracked cost/latency, known failure modes.

  • Make extraction quality legible to non-ML colleagues, and treat de-identification, PHI handling, audit trails, and access controls as part of the modelling problem, not someone else's checklist.

  • Help shape the roadmap around the problems you see - a mission and a close working partner, not a backlog.

What we're looking for

  • Healthcare or life sciences experience with real clinical data - clinical notes, EHR data, claims, registries, or similar. This one is not negotiable for us.

  • 6+ years in applied ML, with models you personally took from problem statement to production and kept working - you know what degraded, how you found out, and what you did.

  • Depth in applied ML on text: information extraction, NER, classification, sequence labelling, weak supervision, and the evaluation practice around them, including annotation guidelines and inter-annotator agreement you've had to act on.

  • Practical, current experience with LLM-based approaches: prompt development, structured output, retrieval, fine-tuning where warranted, evals and observability for generative systems - enough to know where they help, and where they quietly don't.

  • Strong engineering fundamentals in Python. Your work runs in production, not only in a notebook.

  • Strong collaboration instincts across the ML boundary: you define problems with stakeholders before solving them, write clearly, and bring people along.

  • A track record of solving problems rather than closing tickets. Self-directed, comfortable without a playbook, and comfortable being wrong in public when the evidence says so.

Nice to have

  • Fluency with clinical terminologies and standards: SNOMED CT, ICD-10, LOINC, RxNorm, CPT, FHIR

  • Experience with HIPAA, SOC 2, de-identification methodology, or IRB and regulatory-grade data work

  • Experience as an early or first ML hire

  • Experience building or running human-in-the-loop annotation and QC operations at scale

  • OCR and document-understanding experience on low-quality real-world documents

  • Experience mentoring or leading ML engineers, or interest in growing that way

What this role is not

  • Not a research role. The bar is extraction quality in production, not publications.

  • Not an LLM-wrapper role. If your instinct is that every problem is a prompt away from being solved, we'll frustrate each other.

  • Not a large-team role yet. You'd be the first ML engineer in a small Platform Engineering function - breadth and influence, and fewer specialists to lean on.

  • Not a role where someone hands you a clean labelled dataset. Building it is the job.

Why this role is a good bet

  • Ground-floor ownership of a capability with direct commercial weight, with influence over architecture, roadmap, and eventually hiring.

  • Both halves of the modern ML toolkit in one seat, on a problem where the choice between them genuinely matters.

  • A manager who treats process and people work as legitimate engineering work, and intends for this seat to grow.

  • Health tech means the work has stakes - better extraction means higher quality research

Benefits & Perks

  • Equity in Novellia

  • Medical, dental, and vision coverage

  • 401(k)

  • Flexible time off

  • Wellness stipend

  • Up to 12 weeks of parental leave

Don't meet every requirement? Studies show women and people of color are less likely to apply unless they meet every qualification. If you're excited about this role but your experience doesn't align perfectly, we encourage you to apply anyway - you may be the right fit for this or another role.

U.S. Applicants Only

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
368,634 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
In your city
Lead AI Engineer 10 hours ago
$30k – $73k per year (Estimated) • In office • Full-Time • Pune
Python
AI/ML
Fine-tuning
LLM
Reinforcement Learning
LLM Guardrails
AI Agents
DevOps
CI/CD
Docker
GitOps
Helm
Kubernetes
OpenShift
Platform Engineering
Vector
Apply
$33k – $78k per year (Estimated) • Equity • Remote • Full-Time • 8+ years exp • Bachelor's Degree • India
Apex
JavaScript
Python
TypeScript
Apex
Copado
Lightning Web Components
AI/ML
AutoGen
CrewAI
Fine-tuning
Hallucination
LangChain
LangGraph
LlamaIndex
LLM
RAG
Semantic Kernel
Semantic Search
Synthetic Data
Vertex AI
Agentforce
AWS Bedrock AgentCore
Semantic Search
AI Agents
Model Context Protocol
DevOps
AWS
CI/CD
GitHub Actions
Jenkins
Vector
GitHub
Cybersecurity
Crowdstrike
Management
Slack
Marketing
Salesforce
Apply
$35k – $86k per year (Estimated) • In office • Moscow
Python
SQL
Databases
Apache Kafka
AI/ML
LLM
Model Context Protocol
RAG
DevOps
Grafana
Apply
$18k – $51k per year (Estimated) • Remote • Moscow
Bash
Python
Databases
ClickHouse
PostgreSQL
AI/ML
Feature Store
Hadoop
LLM
DevOps
Ansible
CI/CD
Docker
GitLab
Kubernetes
Apply
$71k – $154k per year (Estimated) • In office • Full-Time • Dublin
Java
Databases
Apache Kafka
AI/ML
Copilot
LLM
LLM Guardrails
DevOps
AWS
CI/CD
Kubernetes
GitHub
Apply
$130k – $223k per year (Estimated) • Remote • Full-Time • 15+ years exp
SQL
Marketing
Amplitude
Mixpanel
Apply
$150k – $200k per year • Remote • Full-Time • 15+ years exp • Bachelor's Degree
Java
Node JS
Python
JavaScript
AI/ML
LangChain
LlamaIndex
Prompt Engineering
RAG
Spark
DevOps
AWS
Azure
GCP
Vector
Apply
$150k – $200k per year • Remote • Full-Time • 15+ years exp • Bachelor's Degree • New York
Node JS
JavaScript
Databases
PostgreSQL
Frontend
Next.js
React.js
DevOps
GCP
Apply
See all jobs
This is one of many
368,634 more open roles from verified company boards, updated every day.