{"id":1172826,"url":"https://alion.io/job/amgen-senior-data-engineer-4","title":"Senior Data Engineer","company":{"id":4775,"name":"Amgen","domain":"amgen.com","url":"https://alion.io/company/amgen","size_band":"5000+","is_staffing_agency":false,"employer_type":"direct","is_intermediary":false,"listed_via":null,"ats_vendor":"Workday","truth_index":{"grade":"B","score":80,"open_postings":324,"ghost_share":0,"stale_share":0.818,"repost_share":0.003,"time_to_fill_p50_days":27,"computed_at":"2026-09-28T05:45:00Z"}},"role":"Data Science","role_family":"Data Science","seniority":"senior","employment_type":"full_time","work_mode":"on_site","remote_scope":null,"remote_scope_basis":null,"remote_working_hours":null,"hiring_geo_confidence":"structured","locations":["Hyderabad, India"],"countries":["IN"],"hiring_countries":[],"hiring_countries_total":0,"salary":null,"salary_estimate":{"min_usd":23000,"max_usd":47000,"period":"year","method":"role_seniority_country_remote_cell","sample_n":51},"experience_years_min":9,"visa_sponsorship":false,"relocation_package":false,"has_equity":false,"technologies":[{"name":"Agile","optional":false},{"name":"Amazon CloudWatch","optional":false},{"name":"Amazon EC2","optional":false},{"name":"Amazon EKS","optional":false},{"name":"Amazon S3","optional":false},{"name":"Anomaly Detection","optional":false},{"name":"AWS","optional":false},{"name":"AWS Lambda","optional":false},{"name":"CI/CD","optional":false},{"name":"Collibra","optional":false},{"name":"Databricks","optional":false},{"name":"Delta Lake","optional":false},{"name":"Dimensional Modeling","optional":false},{"name":"Feature Store","optional":false},{"name":"FinOps","optional":false},{"name":"Git","optional":false},{"name":"IAM","optional":false},{"name":"Jira","optional":false},{"name":"Least Privilege","optional":false},{"name":"LLM","optional":false},{"name":"LLM Guardrails","optional":false},{"name":"LLMOps","optional":false},{"name":"Machine Learning","optional":false},{"name":"Matter","optional":false},{"name":"MLFlow","optional":false},{"name":"Model Context Protocol","optional":false},{"name":"pySpark","optional":false},{"name":"Python","optional":false},{"name":"RAG","optional":false},{"name":"Rest API","optional":false},{"name":"Spark","optional":false},{"name":"SQL","optional":false},{"name":"Terraform","optional":false},{"name":"Apache Iceberg","optional":true},{"name":"Azure","optional":true},{"name":"Datadog","optional":true},{"name":"Docker","optional":true},{"name":"Embeddings","optional":true},{"name":"Function Calling","optional":true},{"name":"GCP","optional":true},{"name":"Kubernetes","optional":true},{"name":"Multi-Agent Systems","optional":true},{"name":"PostgreSQL","optional":true},{"name":"Red Teaming","optional":true},{"name":"Reranking","optional":true},{"name":"Splunk","optional":true},{"name":"Tool Use","optional":true}],"status":"live","first_seen_at":"2026-09-24T08:55:07Z","employer_posted_date":"2026-09-24","last_verified_at":"2026-09-28T14:44:21Z","board_verified":true,"closed_at":null,"days_open":4,"trust":{"level":"ok","repost_count":null,"flags":[],"days_open":4},"description":"Career Category\nEngineeringJob Description\nJoin Amgen’s Mission of Serving Patients\nAt Amgen, if you feel like you’re part of something bigger, it’s because you are. Our shared mission-to serve patients living with serious illnesses-drives all that we do.\nSince 1980, we’ve helped pioneer the world of biotech in our fight against the world’s toughest diseases. With our focus on four therapeutic areas -Oncology, Inflammation, General Medicine, and Rare Disease- we reach millions of patients each year. As a member of the Amgen team, you’ll help make a lasting impact on the lives of patients as we research, manufacture, and deliver innovative medicines to help people live longer, fuller happier lives.\nOur award-winning culture is collaborative, innovative, and science based. If you have a passion for challenges and the opportunities that lay within them, you’ll thrive as part of the Amgen team. Join us and transform the lives of patients while transforming your career.\nSenior Databricks Platform Engineer \nAbout Amgen\nAmgen harnesses the best of biology and technology to fight the world’s toughest diseases, and make people’s lives easier, fuller and longer. We discover, develop, manufacture and deliver innovative medicines to help millions of patients. Amgen helped establish the biotechnology industry more than 40 years ago and remains on the cutting-edge of innovation, using technology and human genetic data to push beyond what’s known today.\nAbout the Role\nAmgen’s Data Platform team is seeking a senior, hands-on Databricks Subject Matter Expert to lead the design, evaluation, engineering, governance, and enterprise enablement of capabilities on the Databricks Data Intelligence Platform.\nThe Data Platform team owns the enterprise platform architecture, governance model, policies, engineering standards, guardrails, lifecycle, observability, cost management, security controls, feature evaluations, and reusable platform services that enable data engineering, analytics, BI, data science, machine learning, and generative AI teams.\nThis is a senior individual-contributor and technical-leadership role. The successful candidate will combine deep Databricks platform expertise with strong AWS, security, governance, FinOps, observability, automation, and AI/ML knowledge. The role will translate business and technology needs into secure, scalable, reusable, observable, and cost-efficient platform capabilities.\nThis role enables delivery teams through paved-road patterns, frameworks, APIs, automation, reference implementations, architecture reviews, and expert guidance. It does not own the architecture or implementation of business-specific data pipelines.\nWhat you will do\nRoles & Responsibilities:\nServe as the Databricks platform lead, shaping the platform roadmap, reference architecture, standards, operating model, guardrails, lifecycle strategy, and capability backlog in partnership with Enterprise Data Architecture, security, governance, operations, and delivery teams.\nEvaluate Databricks features and third-party technologies through structured assessments and PoCs. Assess architectural fit, security, privacy, compliance, interoperability, performance, reliability, cost, supportability, and user experience before recommending adoption, restriction, deferral, or rejection.\nLead architecture reviews covering account, workspace, metastore, catalog, networking, storage, compute, SQL, ML/AI, dev/test/prod isolation, regional deployment, serverless and classic compute, capacity, upgrades, high availability, disaster recovery, and RTO/RPO requirements.\nBuild reusable frameworks, accelerators, platform APIs, custom templates, provisioning workflows, policy-as-code controls, and reference implementations using Terraform, Databricks SDKs and REST APIs, the Databricks CLI, and Declarative Automation Bundles-formerly Databricks Asset Bundles.\nDesign and operationalize integrations between Databricks and third party services like Collibra, enabling enterprise metadata, classification, ownership, lineage, business glossary, data-quality, AI-asset, and governed-tag capabilities. \nDefine and enforce security guardrails covering identity federation, SSO/SCIM, OAuth, service principals, least privilege, segregation of duties, Unity Catalog access controls and ABAC, secrets, encryption, private connectivity, egress controls, auditing, and data-exfiltration prevention.\nLead Databricks FinOps, including tagging, allocation, showback or chargeback, budgets, alerts, compute policies, serverless usage policies, cost-anomaly detection, and model or token cost controls. Optimize spend through rightsizing, Photon, autoscaling, auto-termination, workload isolation, and appropriate compute selection.\nEstablish platform observability through SLOs, KPIs, telemetry, dashboards, and alerts for availability, reliability, jobs, queries, compute, SQL warehouses, capacity, security, adoption, and cost, using system tables, audit logs, billing data, lineage, data-quality monitoring, inference tables, and MLflow tracing.\nEnable governed ML, GenAI, RAG, and agent capabilities, including feature and model lifecycle, Model Serving, AI Search, foundation-model access, evaluation, monitoring, and LLMOps. Define responsible-AI guardrails for models, agents, prompts, retrieval components, MCP tools, and external providers.\nImprove platform reliability through root-cause analysis, operational reviews, automated remediation, runbooks, resilience testing, and elimination of recurring incidents.\nDrive user enablement through onboarding, documentation, reference solutions, workshops, office hours, and best-practice communities. Translate strategic-program requirements into reusable platform capabilities and measure improvements in adoption, developer experience, governance, security, reliability, and cost.\nWhat we expect of you\nBasic Qualifications and Experience:\nMaster’s or Bachelor’s degree in computer science or engineering field and 9 to12 years of relevant experience\nMust-Have Skills:\nDemonstrated experience owning or technically leading an enterprise Databricks platform, extending beyond notebook or pipeline development into architecture, administration, governance, security, lifecycle management, reliability, and enablement.\nBroad hands-on Databricks knowledge, with expert depth across several areas such as account and workspace administration, serverless and classic compute, Apache Spark, Delta Lake, Unity Catalog, Lakeflow Jobs and Pipelines, Databricks SQL, Photon, AI/BI, MLflow, and Model Serving.\nProven ability to design enterprise account, workspace, metastore, catalog, dev/test/prod, workload-isolation, promotion, high-availability, disaster-recovery, and capacity strategies, and to establish enforceable platform standards and lifecycle controls.\nStrong Unity Catalog and security expertise, including privileges, ownership, managed and external storage, governed tags, lineage, workspace bindings, row filters, column masks, ABAC, SSO/SCIM, OAuth, service principals, secrets, encryption, private connectivity, auditing, and data-exfiltration controls.\nStrong AWS experience with IAM, VPC, PrivateLink, S3, EC2, KMS, CloudWatch, CloudTrail, Secrets Manager, and STS. Familiarity with EKS, Lambda, Glue, EMR, and RDS is beneficial.\nProven cost-management and observability experience, including compute rightsizing, serverless adoption, SQL warehouse optimization, Photon, budgets, tagging, showback or chargeback, system billing data, system tables, operational telemetry, dashboards, alerts, and cost-anomaly investigation.\nStrong understanding of Databricks AI/ML foundations, including MLflow, Models in Unity Catalog, feature engineering or Feature Store, batch and real-time inference, Model Serving, inference tables, model monitoring, MLOps, RAG, agent evaluation, AI security, and LLM cost and performance considerations.\nStrong Python, PySpark, SQL, automation, and platform API skills. Hands-on experience with Databricks SDKs, REST APIs, CLI, Terraform, Git, CI/CD, automated testing, controlled promotion, rollback, secrets management, and Declarative Automation Bundles or equivalent deployment patterns.\nStrong Spark and SQL troubleshooting and performance-analysis skills, including query plans, partitioning, shuffles, skew, file sizing, table optimization, concurrency, cluster utilization, and Photon-enabled workloads.\nStrong software-engineering and distributed-systems fundamentals, including modular and API design, version control, testing, secure coding, maintainability, production support, technical documentation, relational and dimensional modeling, operational readiness, and recovery patterns.\nStrong analytical problem-solving skills and experience working in Agile environments using tools such as Jira or Jira Align.\nGood-to-Have Skills:\nExperience with current Databricks agent capabilities, including Agent Bricks, custom agents, Knowledge Assistant, Supervisor Agent, AI Playground, tool calling, multi-agent systems, and Model Context Protocol integrations.\nExperience building governed RAG solutions with Databricks AI Search-formerly Vector Search-and MLflow for GenAI, including embeddings, hybrid retrieval, reranking, tracing, evaluation datasets, LLM judges, human feedback, production monitoring, and quality, latency, and cost analysis.\nExperience with Unity Gateway or AI Gateway, Foundation Model APIs, external models, AI Functions, batch inference, custom Model Serving endpoints, feature serving, provider routing, rate limits, budgets, service policies, guardrails, auditing, and token-cost attribution.\nExperience with Databricks AI/BI dashboards, Genie Agents, Genie Code, Databricks Apps, natural-language analytics, semantic metadata, AI-generated SQL controls, and responsible-AI practices such as prompt-injection defense, tool authorization, data-leakage prevention, red teaming, and human oversight.\nHands-on experience integrating Databricks with Collibra and working with data-quality or observability products such as Ataccama, Monte Carlo, Datadog, or Splunk.\nExperience developing self-service portals, Python microservices, and secure platform APIs, including collaboration with React teams and deployment using Docker, Kubernetes, or Amazon EKS.\nExperience with Lakehouse Federation, Delta Sharing or OpenSharing, Clean Rooms, Databricks Marketplace, Apache Iceberg interoperability, external lineage, SQL/NoSQL/vector databases, Lakebase/Postgres, or Azure and GCP Databricks environments.\nExperience supporting data and AI platforms in life sciences or another regulated industry. Preferred certifications include::Databricks Certified Data Engineer Professional\nDatabricks Certified Machine Learning Professional\nDatabricks Certified Generative AI Engineer Associate\nAWS Certified Data Engineer - Associate\nAWS Certified Solutions Architect\nAWS Certified Security - Specialty\nAWS Certified Machine Learning Engineer - Associate\nSAFe Agilist or another SAFe certification\n\nFunctional Skills:\nExcellent written and verbal communication, with the ability to explain complex platform concepts, architecture decisions, risks, and trade-offs in clear, business-relevant language.\nStrong influencing and consensus-building skills, including the ability to establish standards and drive adoption across teams without relying solely on formal authority.\nA platform-product mindset focused on reusable capabilities, paved roads, developer experience, measurable outcomes, and long-term platform health.\nStrong systems-thinking and structured problem-solving skills, with the ability to diagnose issues across application, platform, cloud, governance, security, and operating-model boundaries.\nHigh degree of ownership, initiative, and follow-through, with the ability to move ambiguous topics from exploration through decision, implementation, adoption, and continuous improvement.\nCollaborative and globally minded, with experience working effectively across...","description_format":"text","description_chars":14854,"description_truncated":true,"requirements":{"experience_years_min":9,"management_years_min":null,"team_size_min":null,"manages_managers":false,"education":{"level":"bachelor","optional":false},"security_clearance":false,"languages":[]},"benefits":["Continuous learning"],"hiring_locations":[],"hiring_excludes":[],"relocation_offered":false,"industries":["Biotechnology","Prescription Drugs","Immunology & Inflammation Therapeutics","Cardiometabolic Therapeutics"],"lifecycle":[{"event":"open","at":"2026-09-24T08:55:07Z"}],"liveness":{"score":63,"band":"ok","label":"Likely open","p_open":1,"p_active":0.632,"p_room":1,"age_days":3,"expected_fill_days":27,"reasons":["conf:24","stale_co","velocity","win:early","comp:brand"],"computed_at":"2026-09-28T05:45:00Z"},"pay":null,"html_url":"https://alion.io/job/amgen-senior-data-engineer-4","json_url":"https://alion.io/job/amgen-senior-data-engineer-4.json","meta":{"generated_at":"2026-09-28T23:17:36Z","cache_seconds":300,"methodology":"https://alion.io/methodology","terms":"https://alion.io/terms","contact":"https://alion.io/contact","api":"https://alion.io/developers","usage":{"tier":"crawler","counted_by":"address","units_charged":1,"used_today":863,"day_limit":5000,"remaining_today":4137,"minute_limit":60,"resets_at":"2026-09-29T00:00:00Z"}}}