{"id":1024726,"url":"https://alion.io/job/takeda-pharmaceutical-data-engineering-professional-ii","title":"Data Engineering Professional II","company":{"id":5984,"name":"Takeda Pharmaceutical","domain":"takeda.com","url":"https://alion.io/company/takeda","size_band":"5000+","is_staffing_agency":false,"employer_type":"direct","is_intermediary":false,"listed_via":null,"ats_vendor":"Workday","truth_index":{"grade":"A","score":86,"open_postings":26,"ghost_share":0,"stale_share":0.577,"repost_share":0,"time_to_fill_p50_days":22,"computed_at":"2026-10-03T05:45:00Z"}},"role":"Data Science","role_family":"Data Science","seniority":"middle","employment_type":"full_time","work_mode":"on_site","remote_scope":null,"remote_scope_basis":null,"remote_working_hours":null,"hiring_geo_confidence":"structured","locations":["Bengaluru, India"],"countries":["IN"],"hiring_countries":[],"hiring_countries_total":0,"salary":null,"salary_estimate":{"min_usd":17500,"max_usd":42000,"period":"year","method":"role_seniority_country_remote_cell","sample_n":13},"experience_years_min":4,"visa_sponsorship":false,"relocation_package":false,"has_equity":false,"technologies":[{"name":"Amazon CloudWatch","optional":false},{"name":"Amazon ECS","optional":false},{"name":"Amazon EKS","optional":false},{"name":"Amazon Kinesis","optional":false},{"name":"Amazon S3","optional":false},{"name":"Amazon SageMaker","optional":false},{"name":"Apache Kafka","optional":false},{"name":"AWS","optional":false},{"name":"AWS Bedrock","optional":false},{"name":"AWS Lambda","optional":false},{"name":"AWS Step Functions","optional":false},{"name":"CI/CD","optional":false},{"name":"CloudFormation","optional":false},{"name":"Databricks","optional":false},{"name":"Delta Lake","optional":false},{"name":"Docker","optional":false},{"name":"Feature Store","optional":false},{"name":"GDPR","optional":false},{"name":"Git","optional":false},{"name":"GitHub Actions","optional":false},{"name":"GitLab CI","optional":false},{"name":"HIPAA","optional":false},{"name":"IAM","optional":false},{"name":"Jenkins","optional":false},{"name":"Kubernetes","optional":false},{"name":"LLM","optional":false},{"name":"LLM Guardrails","optional":false},{"name":"Machine Learning","optional":false},{"name":"MLFlow","optional":false},{"name":"Platform Engineering","optional":false},{"name":"pySpark","optional":false},{"name":"Python","optional":false},{"name":"RAG","optional":false},{"name":"SLI/SLO/SLA","optional":false},{"name":"Spark","optional":false},{"name":"SQL","optional":false},{"name":"Terraform","optional":false}],"status":"live","first_seen_at":"2026-07-27T00:00:00Z","employer_posted_date":"2026-07-27","last_verified_at":"2026-10-03T22:38:02Z","board_verified":true,"closed_at":null,"days_open":69,"trust":{"level":"ok","repost_count":null,"flags":[],"days_open":69},"description":"By clicking the “Apply” button, I understand that my employment application process with Takeda will commence and that the information I provide in my application will be processed in line with Takeda’s Privacy Notice and Terms of Use. I further attest that all information I submit in my employment application is true to the best of my knowledge.\nJob Description\nData Engineering Professional II\nDigital, Data & Technology (DD&T) - R&D MLOPs\nAbout the Role\nWe are seeking an MLOps Engineer to operationalize machine learning and generative AI across our R&D and enterprise data ecosystem. You will build and maintain the platforms, pipelines, and controls that move models from notebook experiments into validated, production-grade, GxP-compliant services: supporting use cases that span clinical development, regulatory operations, pharmacovigilance, real-world evidence, and translational/biomarker research.\nThis is a hands-on engineering role at the intersection of data engineering, ML lifecycle automation, and regulated-systems discipline.\nYou will work primarily in Databricks and AWS, partnering with data scientists, platform/cloud engineering, quality, and regulatory teams to ship models that are reproducible, monitored, auditable, and trustworthy.\nKey Responsibilities\nML Lifecycle & Pipeline Automation\n Design, build, and operate end-to-end ML pipelines (data ingestion → feature engineering → training → validation → deployment → monitoring) using Databricks (Delta Lake, MLflow, Unity Catalog, Feature Store, Workflows/Jobs) and AWS services.\n Implement CI/CD for ML and data assets (e.g., GitHub Actions, GitLab CI, or Jenkins), including automated testing, environment promotion (dev → test → prod), and reproducible builds.\nStand up and maintain model registries, model versioning, and artifact lineage so every deployed model is traceable to its data, code, and configuration.\nCloud & Platform Engineering (AWS)\n Build and manage ML infrastructure on AWS - e.g., SageMaker, Bedrock, S3, Lambda, ECS/EKS, Step Functions, ECR, IAM, CloudWatch - using Infrastructure as Code (Terraform or CloudFormation/CDK).\n Integrate Databricks with AWS securely (Unity Catalog governance, cross-account access, VPC/networking, KMS encryption, secrets management).\n Optimize compute and cost (cluster policies, autoscaling, spot strategy, job orchestration) without compromising performance or compliance.\nProduction Monitoring & Reliability\nImplement model and data monitoring: drift detection, data-quality checks, performance/SLA tracking, and automated alerting/retraining triggers.\nEstablish observability and incident-response practices for ML services; participate in on-call/runbook ownership as needed.\nMaintain feature stores and data contracts to ensure consistency between training and serving.\nRegulated-Environment & Compliance Engineering\n Build ML systems that meet GxP expectations and support Computer System Validation (CSV) / Computer Software Assurance (CSA), GAMP 5, 21 CFR Part 11, and data-integrity (ALCOA+) requirements.\nImplement audit trails, electronic records/signatures controls, access controls, and change-management workflows suitable for validated environments.\n Handle PII/PHI and sensitive R&D data in line with HIPAA, GDPR, and internal privacy/data-governance policies (de-identification, anonymization, role-based access).\nAuthor and maintain technical documentation, validation deliverables, and SOP-aligned procedures; partner with Quality/QA and Regulatory on audits and inspections.\nCollaboration & Enablement\nWork under the guidance of Director, Solution Engineering/Solution Architect to produce artifacts and deliverables that adhere to best practices at Takeda.\n Partner with data scientists to productionize models (including LLM/GenAI and RAG applications) and to translate research code into robust, maintainable services.\nContribute reusable templates, accelerators, and self-service tooling that raise the engineering bar across teams.\nPromote MLOps best practices, mentor peers, and document standards.\nRequired Qualifications\n Bachelor’s degree in Computer Science, Engineering, Data Science, or a related field (or equivalent practical experience).\n4+ years of hands-on experience in MLOps, ML engineering, data engineering, or DevOps for data/ML systems.\n Strong Databricks experience: Delta Lake, MLflow, Unity Catalog, Jobs/Workflows, and Spark (PySpark).\n Strong AWS experience across compute, storage, and IAM, plus at least one ML service (SageMaker and/or Bedrock).\n Proficiency in Python for production code (packaging, testing, typing), plus solid SQL.\nExperience building CI/CD pipelines and using Git-based workflows.\n Working knowledge of containerization (Docker) and orchestration (Kubernetes/EKS or ECS).\n Experience with Infrastructure as Code (Terraform, CloudFormation, or CDK).\nUnderstanding of ML lifecycle concepts: experiment tracking, model registry, feature stores, and model monitoring/drift.\nPreferred / Pharma-Specific Qualifications\n Experience delivering software or ML in a GxP / regulated (FDA, EMA) life-sciences environment; familiarity with CSV/CSA, GAMP 5, 21 CFR Part 11, ALCOA+.\n Exposure to pharma/biotech data domains: clinical trial data (CDISC/SDTM/ADaM), regulatory submissions, pharmacovigilance/safety, real-world data (RWD/RWE), or omics/biomarker datasets.\n Experience operationalizing LLM/GenAI workloads (e.g., AWS Bedrock), including RAG, prompt/version management, evaluation, and guardrails.\nFamiliarity with handling PHI/PII under HIPAA/GDPR and with data-governance tooling.\n Streaming/event-driven data (Kafka/Kinesis), data-observability tooling, and feature-store frameworks.\n Relevant certifications: AWS (ML Specialty, Solutions Architect, or DevOps Engineer) and/or Databricks (Data Engineer, ML Engineer).\nWhat Success Looks Like (First 12 Months)\nProduction ML/GenAI pipelines run reproducibly with full lineage, monitoring, and automated promotion across validated environments.\nDeployment lead time and manual handoffs are measurably reduced through reusable templates and CI/CD.\nModels in production are monitored for drift and quality, with documented retraining and rollback procedures.\nEngineering artifacts meet inspection-readiness standards and pass internal QA review.\nLocations\nIND - BengaluruWorker Type\nEmployeeWorker Sub-Type\nRegularTime Type\nFull time","description_format":"text","description_chars":6391,"description_truncated":false,"requirements":{"experience_years_min":4,"management_years_min":null,"team_size_min":null,"manages_managers":false,"education":{"level":"bachelor","optional":false},"security_clearance":false,"languages":[]},"benefits":[],"hiring_locations":[],"hiring_excludes":[],"relocation_offered":false,"industries":["Prescription Drugs","Oncology Therapeutics","Neurology & CNS Therapeutics","Rare Disease Therapeutics"],"lifecycle":[{"event":"open","at":"2026-09-18T10:02:36Z"}],"visa":[],"liveness":{"score":9,"band":"cold","label":"Long shot","p_open":1,"p_active":0.335,"p_room":0.28,"age_days":68,"expected_fill_days":22,"reasons":["conf:5","stale_co","velocity","win:tail","crowd:brand"],"computed_at":"2026-10-03T05:45:00Z"},"pay":null,"html_url":"https://alion.io/job/takeda-pharmaceutical-data-engineering-professional-ii","json_url":"https://alion.io/job/takeda-pharmaceutical-data-engineering-professional-ii.json","meta":{"generated_at":"2026-10-04T00:12:28Z","cache_seconds":300,"methodology":"https://alion.io/methodology","terms":"https://alion.io/terms","contact":"https://alion.io/contact","api":"https://alion.io/developers","usage":{"tier":"crawler","counted_by":"address","units_charged":1,"used_today":160,"day_limit":5000,"remaining_today":4840,"minute_limit":60,"resets_at":"2026-10-05T00:00:00Z"}}}