702,154open jobs
41,623companies
98,791added this week
Browse all
Salary
$135k – $165k per year
Location
Remote (New York, United States)
Seniority
Senior · 5+ years exp
Employment
Full-Time
Overview
Company
Impact
Profile match
We give innovators the data they need to build healthcare AI.

Our Team

Dandelion Health was founded in 2020 by experts in health tech, hospital systems, academia, and clinical AI. We are building the world’s largest AI training and clinical development platform. Today, we pride ourselves on our ability to make data access as easy as possible for AI developers, pharma, and medical devices, while raising the bar for patient safety and data quality. Tomorrow, we will be the place where any healthcare organization can go to build a responsible clinical AI product. Our culture is all about learning from data and improving, so we can help our clients improve health through AI. Meet the rest of our team here.

Our Data

We partner with health systems to safely and ethically make their de-identified patient data available to AI developers. Currently, the data is acquired from Sharp HealthCare, Sanford Health, and Texas Health Resources - with two additional U.S. health systems joining soon.

We have clinical data dating back to July 1, 2016. This data represents over 10 million patients and includes but is not limited to:

  • Structured data (e.g., 100% of the EMR, including some claims)

  • Unstructured text (e.g., clinical notes, radiology reports)

  • Images (e.g., DICOM, pathology)

  • Video

  • Waveforms

  • Continuous streaming monitoring data

Your Role

You are a healthcare data scientist who knows your way around clinical and text-based data. Your primary responsibility will be to leverage and build upon large language models (LLM) and other ML-based approaches for meaningful data abstraction from both unstructured and structured healthcare data. You will join a team of data scientists who own the creation and maintenance of AI-ready datasets for clients in our environment. You will work across the organization to curate datasets by identifying patient subpopulations or disease cohorts, developing methodology to abstract information from multimodal healthcare data sources to support patient phenotyping, and pooling data to enable rapid exploratory AI/ML analyses, model experimentation and/or model validation. You will use your data expertise, programming abilities, and critical thinking skills to support our technical product team, and develop your own analyses to derive insights and enhance our datasets based on the use case. Your team’s ultimate goal is to deliver the highest-quality data possible to our clients, who are building products that improve patient health. You will report to the Data Science Manager, under the Head of Data.

Responsibilities

Your day-to-day responsibilities will include the following:

  • Develop Natural Language Processing (NLP), Large Language Model (LLM) and other ML-based pipelines to abstract relevant labels from text-based healthcare data and store them in scalable data models;

  • Query complex source systems in a range of health data sources (e.g., EMRs, semi-structured reports, free-text clinical provider notes) to identify key data elements and create and enrich high-quality datasets for real-world evidence analyses and training AI algorithms;

  • Own data extraction, wrangling, labeling and QC tasks to create analytical datasets that include abstracted clinical concepts and provide a range of solutions to support customers’ AI activities;

  • Stay current on the latest in applied NLP and generative AI methods and proactively leverage these technologies where applicable;

  • Support the design, testing, validation, analysis, and merging of multimodal data structures from a wide variety of source systems;

  • Develop code and documentation to deliver high-quality and HIPAA-compliant data products on time to customers;

  • Identify and resolve problems using your knowledge, background, and troubleshooting skills;

  • Ensure accuracy, data integrity, and validity of data and analysis in all work;

  • Provide support for technical product team to advance development of the suite of data-related product offerings;

  • Summarize the complexity of abstraction methods, findings and recommendations into clear explanations and presentations for internal and external audiences that have a varying range of technical and clinical experience;

You are not afraid to dig into massive, confusing, disorganized new datasets and get them under control. You are excited to learn new environments, languages, and skills. This is a small, early stage company with enormous ambitions and everyone pitches in across the team.

Qualifications

  • Advanced degree in a quantitative field (ex. Data Science, Biomedical Informatics, Computer Science, Biostatistics), or B.S. with at least 5 years of professional experience

  • At least 2 years of data science and machine learning experience, including building pipelines to extract and curate unstructured and semi-structured data by applying advanced machine learning and AI techniques. Prior experience with clinical and healthcare data is a strong bonus.

  • Fluency in Python and SQL, including fluency with ML/NLP libraries (PyTorch, Tensorflow, HuggingFace, etc.)

  • Familiarity with using modern applied LLM techniques on real-world data

  • Strong technical writing, editing, and communication skills, along with a collaborative mindset

  • Excellent organizational skills with an ability to embrace change and effectively manage multiple projects and consistently plan work to meet deadlines

  • Experience working in or with startups is a plus

Technical Experiences and Skills

We don’t expect anyone to have all of the following skills or experiences, but we do seek candidates who are interested in growing their skill sets and working with healthcare data in all its glorious complexity. The Data Team works closely with our Engineering Team to put our work into production and meet client needs.

  • Git and version control

  • Familiarity with encryption methods

  • Prior experience querying EDWs or databases and creating reports or analytics for healthcare data

  • Familiarity with the data aspects of electronic medical records, ex. Epic, Cerner, Allscripts

  • Any medical ontology experience

  • Any experience working with DICOM or other imaging modalities

  • Experience with AWS

  • Experience with publishing work in peer-reviewed journals

Nature of our work

Our work is fast paced and iterative. We are growing, and we want to support our team members to grow in their skills as well. We are building a team that approaches problems with a diversity of perspectives, values experimentation, and refining our approach based on that experimentation. We work with the full spectrum of healthcare data from tabular data, videos, images, waveforms, etc. If a health system collects it, we might work with it!

If this looks like a partial fit, please reach out, we would love to share more about the work we do for you to understand if it would be a good fit for you.

There is occasional travel for in-person company working days on roughly a quarterly basis.

Team Benefits

  • Remote work and flexible hours. Availability needed for meetings, which we try to keep to a healthy minimum

  • Complete wellness benefits including healthcare, dental, vision, PTO, sick days and more. Ask for details

  • Professional development days to build your skills

  • Collegial work environment

  • Academic bent towards inquiry and problem solving but start-up speed and flexibility

  • Great balance of focus time to work on projects but easy to access team members to discuss issues and work collaboratively

  • Dandelion is a mission-driven company that is focused on improving patient care

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
702,154 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
New York
$149k – $287k per year (Estimated) • Remote/Hybrid • 5+ years exp • Master's Degree • San Francisco
Python
SQL
AI/ML
Copilot
Claude
Fine-tuning
SciPy
Computer Vision
AI Agents
PyTorch
Statsmodels
Machine Learning
DevOps
AWS
Analytics
Seaborn
Matplotlib
Apply
Senior Data Engineer 8 hours ago
$87k – $212k per year (Estimated) • Remote • 5+ years exp • Bogotá
Python
SQL
Python
FastAPI
pySpark
Databases
Snowflake
Databricks
Google BigQuery
BigQuery
AI/ML
Cursor
LangGraph
LangChain
Spark
Claude Code
Dagster
dbt
Scikit-learn
PyTorch
Gemini
LLM
Structured Outputs
DevOps
Terraform
GCP
CloudFormation
Pulumi
CI/CD
AWS
Docker
Platform Engineering
Google Cloud Run
Apply
$335k – $385k per year • Remote/Hybrid • 10+ years exp • Bachelor's Degree • San Francisco
AI/ML
Multimodal AI
Anthropic
Interpretability
Machine Learning
Apply
$120k – $165k per year • Remote/Hybrid • Public Trust • 5+ years exp • Bachelor's Degree • Ashburn
Python
JavaScript
Java
TypeScript
COBOL
Java
Spring Boot
Hibernate
COBOL
IBM MQ
Databases
PostgreSQL
ActiveMQ
Apache Kafka
Frontend
Angular
Mobile
JUnit
DevOps
CI/CD
Jenkins
Git
AWS
Docker
Kubernetes
Bamboo
Linux
Cybersecurity
SonarQube
Management
Jira
Agile
Apply
$170k – $250k per year • In office • Bachelor's Degree • Seattle
Python
Apply
$155k – $170k per year • Remote • Full-Time • 7+ years exp • Master's Degree
Python
SQL
AI/ML
Unstructured.io
Multimodal AI
NLP
Machine Learning
DevOps
Git
AWS
Cybersecurity
HIPAA
Apply
$140k – $160k per year • Remote • Full-Time • 5+ years exp • New York
AI/ML
Multimodal AI
Marketing
LinkedIn
Apply
$160k – $180k per year • Remote • Full-Time • 7+ years exp • Bachelor's Degree • New York
Apply
Analytics Engineer 2 months ago
$140k – $150k per year • Remote • Full-Time • 3+ years exp • Bachelor's Degree
Python
SQL
Databases
Snowflake
AI/ML
Unstructured.io
dbt
Multimodal AI
NLP
AWS Bedrock
LLM
RAG
OpenAI
DevOps
Prometheus
Azure
CI/CD
Git
AWS
Cortex
Analytics
ETL/ELT
Management
Linear
Jira
Agile
Apply
$135k – $150k per year • Remote • Full-Time • 4+ years exp
Management
Jira
Marketing
Salesforce
Apply
$135k – $245k per year • In office • Full-Time • 7+ years exp • PhD • San Francisco • Chicago • New York • Boston • Washington
AI/ML
AI Agents
Agentforce
Agentic Workflows
Management
ServiceNow
Apply
$197k – $314k per year • In office • Full-Time • 10+ years exp • PhD • San Francisco • Boston • Chicago • New York • Austin
AI/ML
AI Agents
LLM
Agentforce
LLM Guardrails
Apply
$130k per year • In office • Full-Time • 8+ years exp • New York • Charlotte
Management
Agile
Apply
$90k per year • In office • Full-Time • Bachelor's Degree • San Francisco • New York • Dallas
Apply
$106k – $203k per year (Estimated) • In office • Full-Time • 10+ years exp • Bachelor's Degree • Minneapolis • Columbus • Philadelphia • New York • Chicago
Analytics
Power BI
Marketing
Salesforce
Apply
See all jobs
This is one of many
702,154 more open roles from verified company boards, updated every day.