385,264open jobs
10,055companies
49,525added this week
Browse all
Salary
$26k – $50k per year (Estimated)
Location
In office (Hyderabad)
Seniority
Senior · 5+ years exp
Employment
Full-Time
Overview
Company
Impact
Profile match
Roche is a Swiss healthcare group founded in Basel in 1896 and is unusual in operating two large divisions of comparable importance: prescription medicines and in vitro diagnostics. The pharmaceutical business is built on oncology, neurology, immunology, ophthalmology and haemophilia, with products such as Ocrevus, Hemlibra, Perjeta and Vabysmo, while the diagnostics division supplies laboratory analysers, molecular tests and sequencing systems to hospitals worldwide. The group owns the American biotechnology company Genentech outright and the Japanese firm Chugai in majority, is controlled by a family shareholder pool, and spends among the largest research budgets in the industry.

At Roche you can show up as yourself, embraced for the unique qualities you bring. Our culture encourages personal expression, open dialogue, and genuine connections, where you are valued, accepted and respected for who you are, allowing you to thrive both personally and professionally. This is how we aim to prevent, stop and cure diseases and ensure everyone has access to healthcare today and for generations to come. Join Roche, where every voice matters.

The Position

Job Description:

The Data Engineer - Clinical Study Design sits at the intersection of data architecture, clinical science, and technology delivery, helping build the data foundation behind Study Designer, a digital product transforming how studies are designed. This role combines strong data engineering skills, an understanding of clinical study data and workflows, and technical curiosity to translate complex clinical data structures into reliable, scalable pipelines and models that power the platform's insights. The ideal candidate is curious, collaborative, AI-minded, and passionate about using data to enable smarter, faster, and more effective clinical study design.

Key Responsibilities

  • Build ingestion pipelines for clinical trial protocols, ICF documents, SmPCs, CSRs, and published articles (PubMed, CTIS, ClinicalTrials.gov) - handling PDF parsing, text extraction, and structured data normalization

  • Design and implement data models in Amazon Aurora (relational) and GraphDB (knowledge graph) to represent trial design entities: endpoints, eligibility criteria, study arms, interventions, therapeutic areas, and their relationships

  • Develop embedding and vectorization pipelines to prepare extracted clinical text for RAG-based retrieval in LangGraph agentic workflows - chunking strategies, metadata enrichment, and vector store population

  • Build and maintain ETL/ELT workflows that transform unstructured clinical content into queryable, linked data across both relational and graph stores

  • Implement data quality validation specific to clinical data - protocol section classification accuracy, entity extraction completeness, cross-reference integrity (NCT IDs, EudraCT numbers, MeSH terms)

  • Build data serving APIs (Python/FastAPI) that expose curated datasets to the Angular frontend and LangGraph agent layer

  • Set up data lineage tracking and audit trails to support regulatory traceability of AI-generated trial design recommendations

Preferred Qualifications:

  • Education: Bachelor's degree in Computer Science, Data Engineering, or a related discipline.

  • 5-8 years of experience building production grade data platforms and pipelines.

  • Experience with biomedical knowledge graphs (e.g., linking drugs -> targets -> diseases -> trials)

  • Prior work with PubMed/MEDLINE data, ClinicalTrials.gov API, or EMA/CTIS data

  • Apache Spark or Databricks for batch processing of large document

Required Skills

  • Python - Primary language; experience with PDF/document parsing libraries (PyMuPDF, pdfplumber, unstructured.io, or similar)

  • SQL - Advanced PostgreSQL-compatible SQL (Aurora); schema design, migrations, query optimization, indexing strategies for clinical data volumes

  • Graph Databases - Hands-on with Neptune, Neo4j, or similar; SPARQL or Cypher query language; ontology/knowledge graph modeling for biomedical entities

  • AWS - Aurora (PostgreSQL), S3, Lambda, Step Functions, SQS/SNS, IAM; infrastructure for data pipeline orchestration

  • NLP / Document Processing - Text extraction from PDFs, section classification, named entity recognition for clinical/biomedical text; familiarity with embedding models and vector stores (OpenSearch, pgvector, or Pinecone)

  • FastAPI - Building data serving endpoints; async patterns; integration with the application backend

  • AI/ML Data Infrastructure - Preparing data for LangChain/LangGraph consumption; RAG pipeline design (chunking, retrieval, reranking); prompt-data integration patterns

  • Pipeline Orchestration - Experience with workflow orchestration tools (Airflow, Prefect, Step Functions, or Temporal); designing DAGs for multi-stage data pipelines with dependency management, retry logic, and monitoring

  • CI/CD & IaC - Terraform or CDK, Docker, Git; automated pipeline testing and deployment on AWS

Who we are

A healthier future drives us to innovate. Together, more than 100’000 employees across the globe are dedicated to advance science, ensuring everyone has access to healthcare today and for generations to come. Our efforts result in more than 26 million people treated with our medicines and over 30 billion tests conducted using our Diagnostics products. We empower each other to explore new possibilities, foster creativity, and keep our ambitions high, so we can deliver life-changing healthcare solutions that make a global impact.

Let’s build a healthier future, together.

Roche is an Equal Opportunity Employer.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
385,264 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
Hyderabad
$69k – $172k per year (Estimated) • In office • Full-Time • 8+ years exp • Master's Degree • Melbourne • Sydney
AI/ML
AI Agents
Copilot
DevOps
Azure
GitHub
Apply
$245k – $384k per year • Equity • In office • Full-Time • 7+ years exp • Bachelor's Degree • Bucharest
C#
JavaScript
Node JS
PowerShell
SQL
TypeScript
C#
.NET
Databases
Redis
Frontend
Angular
DevOps
Azure
Azure DevOps
CI/CD
GitHub
GitHub Actions
GitOps
Terraform
Analytics
Power BI
Apply
Remote/Hybrid • Full-Time • Singapore
C#
Java
Python
TypeScript
JavaScript
C#
.NET
Databases
Apache Kafka
Kafka
Frontend
Angular
React.js
Vue.js
DevOps
CI/CD
Kubernetes
WebSockets
Cybersecurity
Volatility
Apply
$181k – $326k per year • In office • Full-Time • 8+ years exp • Bachelor's Degree • United States
Python
TypeScript
AI/ML
AI Agents
Apply
$170k – $220k per year • Equity • In office • Full-Time • 3+ years exp • New York
AI/ML
AI Agents
Design
Figma
Marketing
LinkedIn
Apply
$61k – $114k per year (gross) • Remote/Hybrid • Full-Time • Warsaw
DevOps
CI/CD
Platform Engineering
Apply
In office • Full-Time • Master's Degree • Basel • South San Francisco
Apply
In office • Full-Time • 4+ years exp • Bachelor's Degree • San José
PowerShell
Cybersecurity
Microsoft Entra ID
Apply
$29k – $69k per year (Estimated) • In office • Full-Time • 12+ years exp • Hyderabad
Apply
$16k – $39k per year (Estimated) • In office • Full-Time • 4+ years exp • Bachelor's Degree • Hyderabad
ABAP
Apex
Apex
MuleSoft
Management
Google Workspace
Jira
ServiceNow
Apply
AI / ML Engineer 2 days ago
$31k – $79k per year (Estimated) • In office • Full-Time • 5+ years exp • Pune • Bengaluru • Hyderabad
AI/ML
PyTorch
TensorFlow
Apply
Security Architect 2 days ago
$34k – $84k per year (Estimated) • In office • Full-Time • 5+ years exp • Hyderabad
Apply
Full Stack Engineer 2 days ago
$26k – $69k per year (Estimated) • In office • Full-Time • 8+ years exp • Bachelor's Degree • Hyderabad • Gurgaon
Java
Java
Spring Boot
Databases
Apache Kafka
Azure Cosmos DB
Kafka
PostgreSQL
AI/ML
LLM
DevOps
Azure
Azure AKS
CI/CD
Docker
Kubernetes
Rest API
Apply
Full Stack Engineer 2 days ago
$47k – $102k per year (Estimated) • In office • Full-Time • 10+ years exp • Bachelor's Degree • Chennai • Hyderabad
Java
Java
Spring Boot
Databases
Apache Kafka
DynamoDB
Kafka
Oracle
Redis
DevOps
Amazon EC2
Amazon S3
API Gateway
AWS
AWS Lambda
CI/CD
Docker
IAM
Kubernetes
Rest API
Apply
$30k – $81k per year (Estimated) • In office • Full-Time • 5+ years exp • Hyderabad • Pune • Mumbai
Databases
Databricks
DevOps
Azure
Analytics
ETL/ELT
Apply
See all jobs
This is one of many
385,264 more open roles from verified company boards, updated every day.