429,021open jobs
14,503companies
63,779added this week
Browse all
Salary
$22k – $46k per year (Estimated)
Location
Remote/Hybrid (Hyderabad, India)
Seniority
Senior · 5+ years exp
Employment
Full-Time
Overview
Company
Impact
Profile match
Merck & Co. is an American pharmaceutical company founded in 1891 as the United States arm of the German firm Merck and made independent after the First World War, trading as MSD outside the United States and Canada. Its portfolio is dominated by the immuno-oncology drug Keytruda, alongside vaccines such as Gardasil and Vaxneuvance, hospital acute care products, and a large animal health business covering livestock and companion animals. The company is headquartered in Rahway, New Jersey, listed on the New York Stock Exchange, and spends a very large share of revenue on research into oncology, infectious disease and cardiometabolic therapies.

Job Description

Specialist: Data Engineering

The Opportunity:

Join a global biopharma company with a 130-year legacy and mission to achieve new milestones in healthcare. Be part of a technology-driven, data-led organization supporting a diversified portfolio of medicines, vaccines, and animal health products. Work alongside passionate teams that use data, analytics, and insights to drive decisions and tackle some of the world’s greatest health threats.

Our Technology Centers are globally distributed hubs that enable our digital transformation and business outcomes across IT. They bring together diverse teams to collaborate, share best practices, and deliver solutions that save and improve lives.

This role is based at our Hyderabad Tech Center and follows a hybrid working model (3 days onsite, 2 days remote). Candidates are expected to reside within commuting distance of the Hyderabad office.

Role Overview

We are hiring a hands-on Data Engineer who can design, build, and operate production-grade data platforms and pipelines end to end. You will deliver reliable, governed, secure, and analytics-ready data by implementing modern data warehousing and Lakehouse patterns on AWS and Databricks, with strong focus on data quality, dimensional modeling, and scalable ETL/ELT. This role partners closely with analytics, data science, and business stakeholders to translate requirements into robust datasets, while applying engineering best practices such as testing, code reviews, CI/CD, and observability.

What will you do in this role

  • Design, build, and operate batch and streaming data pipelines to ingest data from multiple sources into an AWS data lake / lakehouse and data warehouse.

  • Develop and maintain ETL/ELT transformations using Python, PySpark, and SQL; optimize jobs for performance, cost, and reliability.

  • Partner with Data Analysts, Data Scientists, and business stakeholders to understand use cases and deliver curated, analytics-ready datasets and features.

  • Implement data quality controls (validation rules, reconciliation, anomaly checks), define SLAs/SLOs, and contribute to metadata, lineage, and data catalog practices.

  • Use orchestration and observability to run pipelines reliably (e.g., Databricks Workflows, AWS Step Functions, scheduling, logging, monitoring, alerting).

  • Apply engineering best practices: unit/integration testing, automated data tests, code reviews, and quality gates within CI/CD.

  • Model and publish data for BI/analytics using dimensional modeling (star/snowflake), facts & dimensions, and slowly changing dimensions (SCD).

  • Write and tune advanced SQL for profiling, transformations, and performance troubleshooting across large datasets.

  • Build on AWS using services such as S3, Glue, Lambda, Step Functions, EMR, and CloudWatch; follow security best practices (IAM, encryption, least privilege).

  • Provision and manage cloud resources using Infrastructure as Code (e.g., Terraform) across dev/test/prod environments.

  • Package and deploy workloads using Docker (and where applicable ECS/Fargate); manage dependencies and runtime configurations.

  • Use GitHub for version control (branching strategies, pull requests, code reviews) and set up CI/CD for automated build, test, and deployment.

  • Develop scalable processing on Databricks / Apache Spark using PySpark and lakehouse concepts (e.g., Delta Lake, ACID, schema evolution).

  • Use notebooks (e.g., Jupyter/Databricks) for exploration and PoCs, then productionize solutions with reusable modules, tests, and deployment pipelines.

  • Work in an Agile delivery model (planning, daily sync, reviews, retros), providing accurate estimates and proactively managing risks/dependencies.

  • Create and maintain technical documentation (data contracts, pipeline specs, runbooks) and support operational handoffs.

What Should you have:

  • 5+ years of hands-on experience in data engineering building production pipelines and data 5+ years of hands-on experience in data engineering building production pipelines and data platforms.

  • Strong AWS experience: S3, Glue, Lambda, Step Functions, EMR (and/or ECS/Fargate), plus CloudWatch; solid grasp of IAM and encryption.

  • Nice to have: AWS certification (Developer/Architect) or equivalent demonstrated expertise.

  • Experience working in Agile teams; strong collaboration, communication, and stakeholder management skills.

  • Experience with Databricks and lakehouse capabilities (e.g., Delta Lake, job/workflow orchestration, cluster tuning) is strongly preferred.

  • Strong SQL skills including complex joins/window functions, data profiling, and performance tuning; understanding of dimensional modeling concepts.

  • Proficient in Python and PySpark with solid Spark fundamentals (partitioning, shuffle, caching, file formats) and ability to debug/optimize.

  • Strong with GitHub, CI/CD concepts, and engineering practices (code reviews, branching, release management); working knowledge of Docker and Terraform.

  • Demonstrated ability to work across teams, drive alignment, and take ownership to deliver outcomes (including production support/on call as needed).

  • Nice to have experience in ETL tools such as DBT, Matillion and data quality/testing frameworks (e.g., Collibra) and data governance tools such as Immuta

  • Nice to have experience with orchestration tools (e.g., Airflow), streaming (Kafka/Kinesis), and modern table formats (Delta/Iceberg/Hudi).

  • Bachelor’s degree in computer science, Engineering, or related field (or equivalent practical experience).

Primary Skills:

Python, PySpark, SQL, AWS, Databricks, GitHub, Data Lake, ETL/ELT and CI/CD

Secondary Skills:

Dimensional modeling, Docker and Terraform

Who we are

We are known as well-known org Inc., Rahway, New Jersey, USA in the United States and Canada and MSD everywhere else. For more than a century, bringing forward medicines and vaccines for many of the world's most challenging diseases. Today, our company continues to be at the forefront of research to deliver innovative health solutions and advance the prevention and treatment of diseases that threaten people and animals around the world.

What we look for

Imagine getting up in the morning for a job as important as helping to save and improve lives around the world. Here, you have that opportunity. You can put your empathy, creativity, digital mastery, or scientific genius to work in collaboration with a diverse group of colleagues who pursue and bring hope to countless people who are battling some of the most challenging diseases of our time. Our team is constantly evolving, so if you are among the intellectually curious, join us-and start making your impact today.

Required Skills:

Amazon Web Services (AWS), CI/CD, Databricks Platform, Data ETL, Data Lake, GitHub, PySpark, Python (Programming Language), Structured Query Language (SQL)

Preferred Skills:

Dimensional Modeling, Docker (Software), Terraform

Current Employees apply HERE

Current Contingent Workers apply HERE

Secondary Language(s) Job Description:

#MSDHYDIT

Search Firm Representatives Please Read Carefully

Merck & Co., Inc., Rahway, NJ, USA, also known as Merck Sharp & Dohme LLC, Rahway, NJ, USA, does not accept unsolicited assistance from search firms for employment opportunities. All CVs / resumes submitted by search firms to any employee at our company without a valid written search agreement in place for this position will be deemed the sole property of our company. No fee will be paid in the event a candidate is hired by our company as a result of an agency referral where no pre-existing agreement is in place. Where agency agreements are in place, introductions are position specific. Please, no phone calls or emails.

Employee Status:

Regular

Relocation:

Domestic

VISA Sponsorship:

No

Travel Requirements:

No Travel Required

Flexible Work Arrangements:

Hybrid

Shift:

Not Indicated

Valid Driving License:

No

Hazardous Material(s):

n/a

Job Posting End Date:

09/16/2026

*A job posting is effective until 11:59:59PM on the day BEFORE the listed job posting end date. Please ensure you apply to a job posting no later than the day BEFORE the job posting end date.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
429,021 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
Hyderabad
$22k – $46k per year (Estimated) • Remote/Hybrid • Full-Time • 5+ years exp • Bachelor's Degree • Hyderabad
Python
SQL
Python
pySpark
Databases
Databricks
Apache Iceberg
Delta Lake
Apache Kafka
Apache Hudi
AI/ML
Spark
Airflow
DevOps
Terraform
CI/CD
AWS
Docker
AWS Fargate
AWS Lambda
GitHub
Amazon S3
IAM
Amazon ECS
Amazon CloudWatch
Amazon Kinesis
AWS Step Functions
Cybersecurity
Least Privilege
Shuffle
Analytics
ETL/ELT
Matillion
Dimensional Modeling
Collibra
Apply
$15k – $42k per year (Estimated) • Remote/Hybrid • Full-Time • 2+ years exp • Bachelor's Degree • Hyderabad
Python
SQL
Python
pySpark
Databases
Databricks
Apache Iceberg
Delta Lake
AI/ML
Spark
Airflow
DevOps
Terraform
CI/CD
AWS
Docker
AWS Fargate
AWS Lambda
GitHub
Amazon S3
IAM
Amazon ECS
Amazon CloudWatch
AWS Step Functions
Cybersecurity
Least Privilege
Analytics
ETL/ELT
Dimensional Modeling
Collibra
Apply
$15k – $42k per year (Estimated) • Remote/Hybrid • Full-Time • 2+ years exp • Bachelor's Degree • Hyderabad
Python
SQL
Python
pySpark
Databases
Databricks
Apache Iceberg
Delta Lake
AI/ML
Spark
Airflow
DevOps
Terraform
CI/CD
AWS
Docker
AWS Fargate
AWS Lambda
GitHub
Amazon S3
IAM
Amazon ECS
Amazon CloudWatch
AWS Step Functions
Cybersecurity
Least Privilege
Analytics
ETL/ELT
Dimensional Modeling
Collibra
Apply
$23k – $49k per year (Estimated) • Remote/Hybrid • Full-Time • 5+ years exp • Master's Degree • Mumbai
Python
SQL
Cybersecurity
GDPR
Analytics
Tableau
Power BI
Apply
$87k – $137k per year • Remote/Hybrid • Full-Time • Bachelor's Degree • Rahway
Python
SQL
Analytics
Power BI
Alteryx
Microsoft Excel
Management
UiPath
Apply
$40k – $111k per year • In office • Internship • 2+ years exp • Wilson
Apply
$40k – $111k per year • Remote/Hybrid • Internship • Bachelor's Degree • Wilmington
Apply
$40k – $111k per year • In office • Full-Time • 2+ years exp • Bachelor's Degree • Boston • South San Francisco
Apply
$23k – $49k per year (Estimated) • Remote/Hybrid • Full-Time • 5+ years exp • Master's Degree • Mumbai
Python
SQL
Cybersecurity
GDPR
Analytics
Tableau
Power BI
Apply
$20k – $58k per year (Estimated) • Remote/Hybrid • Full-Time • 5+ years exp • Mumbai
Apply
$25k – $62k per year (Estimated) • In office • Full-Time • 5+ years exp • Hyderabad
Apply
$9k – $22k per year (Estimated) • In office • Full-Time • Hyderabad
DevOps
Incident Management
SLI/SLO/SLA
Apply
$25k – $62k per year (Estimated) • In office • Full-Time • 5+ years exp • Hyderabad
Apply
$28k – $67k per year (Estimated) • In office • Full-Time • 5+ years exp • Jaipur • Hyderabad • Pune
Apply
$25k – $62k per year (Estimated) • In office • Full-Time • 12+ years exp • Hyderabad
Apply
See all jobs
This is one of many
429,021 more open roles from verified company boards, updated every day.