404,711open jobs
14,073companies
78,231added this week
Browse all
Salary
$94k – $198k per year (Estimated)
Location
Remote (Germany)
Seniority
Middle · 3+ years exp
Employment
Full-Time
Overview
Company
Impact
Profile match
Apheris enables federated machine learning so that models can be trained across organisations without moving the data. Founded in 2019 in Berlin, it focuses on pharmaceutical and industrial consortia that cannot pool proprietary datasets. Its governance layer defines exactly what computations each party permits.

About Apheris

At Apheris, we are building the future of how AI is applied in pharmaceutical R&D. We enable leading pharmaceutical teams to discover and develop drugs faster. We host the industry’s largest federated data networks for drug discovery AI, spanning co-folding, ADMET, and antibody developability.

Across these networks, models are trained on proprietary industry datasets to achieve higher performance and broader applicability while keeping data control and IP protected. We deliver these superior models through drug discovery applications that enable teams to run them at scale, further customize them, and integrate them into existing R&D workflows.

  • AI Structural Biology (AISB) Network: Pharmaceutical companies collaborate in the field of co-folding, structure-based binding affinity predictions and antibody design.
  • ADMET Network: Pharmaceutical and biotech companies collaborate to improve small-molecule property prediction and expand into further drug modalities.
  • Antibody Developability Network: Pharma partners collaborate to federate historical and purpose-built antibody developability data sets for secure ML training, without data leaving each partner’s environment.

About the role

We are looking for a Forward-Deployed Cheminformatician to own how binding data is prepared across our co-folding focused networks and initiatives. Binding data is the input that decides whether our co-folding and binding-affinity models perform in real drug programs. It arrives from pharma partners in heterogeneous shapes - different assay registries, different metadata, different chemical-representation standards, different choices on qualifiers, replicates and censoring.

We need someone who turns this into a repeatable, well-documented preparation pipeline that pharma representatives can run alongside us, and that scales to the public-data corpus we build for our own model training.

This is half engineering, half forward-deployed work. You will define the protocol, harden it with validators and scripts, integrate it into the Apheris products, run it with each new partner, and own the equivalent pipeline for the public binding-data corpus.

What you will do

  • Define and own the binding-data preparation protocol - data schema, small-molecule standardization, assay metadata model, value handling (KD, Ki, IC50, pIC50), qualifier and censored-value handling, duplicate and replicate aggregation.
  • Build the tooling that runs it - modular scripts, validators with actionable errors, and reusable pipelines that survive different pharma upstream systems (Dotmatics, Spot fire, in-house registries).
  • Work forward-deployed with pharma. Sit with their biologists and medicinal chemists, walk them through the protocol, sense-check what an assay column actually measures, and unblock retrieval.
  • Maintain the small-molecule representation pipeline - RDK it standardization, tautomer and ionization handling, stereochemistry preservation, and PAINS / frequent-hitter filtering.
  • Curate the public binding-data foundation - ChEMBL,BindingDB, PubChemBioAssay -prepared to the same standard, so our models train on the strongest public baseline anyone can assemble.
  • Hand the productized pipeline cleanly to engineering for scaling, and partner with ML to keep the data contract valid as models and networks evolve.

What we expect from you

You should apply if:

  • You have a BSc, MSc, PhD or equivalent in cheminformatics, computational chemistry, or a related field, plus 3+ years preparing biological assay data in a discovery setting.
  • You are fluent in Python andRDKit. SMILES normalization, tautomer / ionization / stereochemistry handling, and scaffold extraction are second nature, and you understand why each matters for activity cliffs and model training.
  • You have hands-on experience curating quantitative binding assay data (KD, Ki, IC50, pIC50) and HTS data - censored values, qualifiers, duplicates, replicate aggregation, and assay metadata interpretation.
  • You write good engineering code -version control, tested modular scripts, validators that return useful errors.
  • You are comfortable forward-deployed with pharma medicinal chemists and biologists. You can sit in a sense-check meeting, pull out what is actually meant by a column label, and encode that back into the protocol.
  • You enjoy turning a messy ad-hoc cleaning job into a repeatable protocol others can run.
Bonus points if:

  • You have practical familiarity with public binding-data sources (ChEMBL,BindingDB, PubChemBioAssay) and the gotchas in each.
  • You have applied LLM tooling (Claude, Codex, Cursor) to accelerate data cleaning or metadata harmonization.
  • You have worked across institutional data boundaries - federated, multi-party, or otherwise - where the data-preparation contract has to hold under partial visibility.
  • You have a publication record or open-source contributions in cheminformatics or quantitative pharmacology.

What we offer you

  • Industry-competitive compensation, including early-stage virtual share options
  • Remote-first work - work where you work best
  • Wellbeing budget, mental health support, work-from-home budget, co-working stipend, and learning budget
  • Generous holiday allowance
  • Office Days at our Berlin HQ or a different European location (3x per year)
  • A high-calibre, execution-focused team with experience from leading organizations
Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
404,711 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
In your city
Data Audit Manager 1 day ago
$104k – $156k per year • In office • Full-Time • 5+ years exp • Richmond
Python
SQL
Apply
$14k – $38k per year (Estimated) • In office • Full-Time • 2+ years exp • Bachelor's Degree • Hyderabad
JavaScript
Python
SQL
TypeScript
Frontend
npm
React.js
DevOps
AWS
Azure
CI/CD
Docker
Git
Grafana
Kubernetes
Prometheus
Rest API
QA
Cypress
Jest
Apply
In office • 3+ years exp
Go
Python
TypeScript
Databases
BigQuery
Google BigQuery
DevOps
CI/CD
Docker
GCP
Git
Google Cloud Run
Grafana
Prometheus
Terraform
Apply
Remote
Python
Python
Django
DevOps
Bitbucket
CI/CD
GitHub
Analytics
ETL/ELT
Marketing
LinkedIn
Apply
Founding Engineer 1 day ago
$137k – $190k per year • In office • Full-Time • 3+ years exp • New York
Python
TypeScript
AI/ML
AI Agents
LLM
Apply
Remote • Master's Degree
Apply
Remote • Full-Time • PhD
Apply
See all jobs
This is one of many
404,711 more open roles from verified company boards, updated every day.