663,523open jobs
38,718companies
99,026added this week
Browse all
Salary
$180k – $220k per year
Location
In office (San Francisco)
Seniority
Staff · 5+ years exp
Employment
Full-Time
Overview
Company
Impact
Profile match
About us Our mission is to help R&D teams go from idea to IND, 50% faster. Mithrl is built by scientists, for scientists.

ABOUT MITHRL

We envision a world where novel drugs and therapies reach patients in months, not years, accelerating breakthroughs that save lives.

Mithrl is building the world’s first commercially available AI Co-Scientist-a discovery engine that empowers life science teams to go from messy biological data to novel insights in minutes. Scientists ask questions in natural language, and Mithrl answers with real analysis, novel targets, and patent-ready reports.

Our traction speaks for itself:

  • 12X year-over-year revenue growth

  • Trusted by leading biotechs and big pharma across three continents

  • Driving real breakthroughs from target discovery to patient outcomes.

WHAT YOU WILL DO

Build and own an AI-powered ingestion & normalization pipeline to import data from a wide variety of sources - unprocessed Excel/CSV uploads, lab and instrument exports, as well as processed data from internal pipelines.

Develop robust schema mapping, coercion, and conversion logic (think: units normalization, metadata standardization, variable-name harmonization, vendor-instrument quirks, plate-reader formats, reference-genome or annotation updates, batch-effect correction, etc.).

Use LLM-driven and classical data-engineering tools to structure “semi-structured” or messy tabular data - extracting metadata, inferring column roles/types, cleaning free-text headers, fixing inconsistencies, and preparing final clean datasets.

Ensure all transformations that should only happen once (normalization, coercion, batch-correction) execute during ingestion - so downstream analytics / the AI “Co-Scientist” always works with clean, canonical data.

Build validation, verification, and quality-control layers to catch ambiguous, inconsistent, or corrupt data before it enters the platform.

Collaborate with product teams, data science / bioinformatics colleagues, and infrastructure engineers to define and enforce data standards, and ensure pipeline outputs integrate cleanly into downstream analysis and storage systems.

WHAT YOU BRING

Must-have

  • 5+ years of experience in data engineering / data wrangling with real-world tabular or semi-structured data.

  • Strong fluency in Python, and data processing tools (Pandas, Polars, PyArrow, or similar).

  • Excellent experience dealing with messy Excel / CSV / spreadsheet-style data - inconsistent headers, multiple sheets, mixed formats, free-text fields - and normalizing it into clean structures.

  • Comfort designing and maintaining robust ETL/ELT pipelines, ideally for scientific or lab-derived data.

  • Ability to combine classical data engineering with LLM-powered data normalization / metadata extraction / cleaning.

  • Strong desire and ability to own the ingestion & normalization layer end-to-end - from raw upload → final clean dataset - with an eye for maintainability, reproducibility, and scalability.

  • Good communication skills; able to collaborate across teams (product, bioinformatics, infra) and translate real-world messy data problems into robust engineering solutions.

Nice-to-have

  • Familiarity with scientific data types and “modalities” (e.g. plate-readers, genomics metadata, time-series, batch-info, instrumentation outputs).

  • Experience with workflow orchestration tools (e.g. Nextflow, Prefect, Airflow, Dagster), or building pipeline abstractions.

  • Experience with cloud infrastructure and data storage (AWS S3, data lakes/warehouses, database schemas) to support multi-tenant ingestion.

  • Past exposure to LLM-based data transformation or cleansing agents - building or integrating tools that clean or structure messy data automatically.

  • Any background in computational biology / lab-data / bioinformatics is a bonus - though not required.

WHAT YOU WILL LOVE AT MITHRL

  • Mission-driven impact: you’ll be the gatekeeper of data quality - ensuring that all scientific data entering Mithrl becomes clean, consistent, and analysis-ready. You’ll have outsized influence over the reliability and trustworthiness of our entire data + AI stack.

  • High ownership & autonomy: this role is yours to shape. You decide how ingestion works, define the standards, build the pipelines. You’ll work closely with our product, data science, and infrastructure teams - shaping how data is ingested, stored, and exposed to end users or AI agents.

  • Team: Join a tight-knit, talent-dense team of engineers, scientists, and builders

  • Culture: We value consistency, clarity, and hard work. We solve hard problems through focused daily execution

  • Speed: We ship fast (2x/week) and improve continuously based on real user feedback

  • Location: Beautiful SF office with a high-energy, in-person culture

  • Benefits: Comprehensive PPO health coverage through Anthem (medical, dental, and vision) + 401(k) with top-tier plans

We encourage you to apply even if you do not believe you meet every single qualification. Not all strong candidates will meet every single qualification as listed. Research shows that people who identify as being from underrepresented groups are more prone to experiencing imposter syndrome and doubting the strength of their candidacy, so we urge you not to exclude yourself prematurely and to submit an application if you're interested in this work. We think AI systems like the ones we're building have enormous social and ethical implications. We think this makes representation even more important, and we strive to include a range of diverse perspectives on our team.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
663,523 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
San Francisco
$39k – $99k per year (Estimated) • Remote/Hybrid • Full-Time • 3+ years exp • Bachelor's Degree • Bengaluru
Python
AI/ML
LangGraph
AutoGen
LangChain
Model Context Protocol
Prompt Engineering
Function Calling
AI Agents
Anthropic SDK
Langfuse
LangSmith
OpenAI SDK
AgentOps
CrewAI
LLM
RAG
Reranking
Semantic Search
Hybrid Search
LLMOps
A2A
Human-in-the-Loop
Context Engineering
Semantic Search
LLM Evaluation
LLM Guardrails
Multi-Agent Systems
Frontend
GraphQL
DevOps
GCP
Azure
CI/CD
Git
AWS
Vector
Apply
Data Engineer 1 day ago
$28k – $68k per year (Estimated) • Remote/Hybrid • Full-Time • 3+ years exp • Bachelor's Degree • Mumbai • Bengaluru
Python
SQL
Databases
SAP HANA
Snowflake
Databricks
AI/ML
LangGraph
LangChain
Spark
dbt
AI Agents
Semantic Kernel
LLM
RAG
DevOps
GCP
Azure
AWS
Apply
$22k – $44k per year (Estimated) • Remote • 2+ years exp • Moscow
Python
Python
Django
Celery
Databases
MySQL
PostgreSQL
Redis
AI/ML
YandexGPT
GigaChat
LLM
OpenAI
DevOps
Rest API
Git
Docker
Gitflow
Trunk-Based Development
Management
Jira
Apply
$17k – $39k per year (Estimated) • Equity • Remote/Hybrid • Full-Time • 1+ year exp • Master's Degree • Hyderabad
Python
SQL
Analytics
Tableau
Power BI
Apply
$229k – $345k per year • Equity • Remote • Full-Time • 8+ years exp • Austin • Dallas • Houston • Phoenix
Python
JavaScript
Perl
Apply
$300k – $500k per year • In office • Full-Time • 8+ years exp • San Francisco
Apply
$180k – $220k per year • In office • Full-Time • San Francisco
Python
AI/ML
Multimodal AI
AI Agents
LLM
Knowledge Graph
Agentic Workflows
Multi-Agent Systems
Apply
$180k – $220k per year • In office • Full-Time • San Francisco
Python
AI/ML
Multimodal AI
AI Agents
LLM
Apply
$180k – $220k per year • In office • Full-Time • 7+ years exp • San Francisco
Python
JavaScript
Python
Django
Frontend
GraphQL
React.js
DevOps
GCP
Azure
AWS
Apply
$180k – $220k per year • In office • Full-Time • San Francisco
DevOps
Terraform
CloudFormation
AWS
Docker
Kubernetes
Platform Engineering
Configuration Management
Amazon EKS
AWS Lambda
Amazon EC2
Amazon S3
IAM
Amazon ECS
Apply
$163k – $204k per year • Remote • 7+ years exp • San Francisco
Apply
$128k – $192k per year • Equity • In office • Full-Time • 1+ year exp • San Francisco • Pleasanton
Apply
Investment Officer 1 day ago
$129k – $223k per year • In office • Full-Time • 5+ years exp • Bachelor's Degree • San Francisco
Apply
$152k – $327k per year (Estimated) • In office • 13+ years exp • San Francisco
AI/ML
OpenAI
Anthropic
Management
Stripe
Apply
$152k – $279k per year (Estimated) • Remote/Hybrid • Full-Time • 12+ years exp • Associate's Degree • Chicago • Milwaukee • Dallas • Columbus • Kirkland
Python
JavaScript
Node JS
AI/ML
LLM
Frontend
React.js
DevOps
Terraform
AWS CDK
Azure
AWS
Docker
Kubernetes
Platform Engineering
Amazon EKS
AWS Fargate
AWS Lambda
Amazon EC2
HPC
Apply
See all jobs
This is one of many
663,523 more open roles from verified company boards, updated every day.