1,166,878open jobs
65,935companies
208,638added this week
Browse all
Salary
≈ $16k – $37k per year (Estimated)
Location
In office (Hyderabad)
Seniority
Senior · 5+ years exp

Confirmed on the employer's own hiring board on Oct 3, 2026. First seen by Alion on Sep 25, 2026.

Overview
Company
Impact
Profile match
Cornerstone OnDemand is a global Computer Software company that offers solutions wrt Recruiting, Personalized Learning, Human Capital Management, HR Planning, etc. It is listed on NASDAQ.

Senior DevOps / Cloud Platform Engineer - ML AI Infrastructure Job Summary We are looking for a Senior DevOps / Cloud Platform Engineer with strong experience in AWS, Kubernetes, CI/CD, infrastructure automation, and ML/AI infrastructure to design, deploy, manage, and optimize cloud infrastructure supporting our machine learning and AI services. The ideal candidate will have hands-on experience with AWS EKS, SageMaker, Bedrock, Docker, Kubernetes, Terraform, Helm, GitHub Actions, Databricks, Elasticsearch, and self-hosted LLM deployments . This role will work closely with Data Engineering, Machine Learning, and Software Engineering teams to build reliable, scalable, secure, and cost-efficient platforms for ML services across development and production environments.

In this role you will... Key Responsibilities AWS Kubernetes Infrastructure Design, deploy, and administer AWS infrastructure supporting ML and AI workloads. Manage Amazon EKS clusters , including cluster provisioning, upgrades, scaling, networking, and troubleshooting.

Work with AWS SageMaker, AWS Bedrock, EKS, ECS, and related AWS services . Configure and manage Kubernetes Ingress controllers such as NGINX and AWS ALB . Manage Cloudflare Tunnels, DNS, Cloudflare configuration, networking, and security .

Troubleshoot application, networking, compute, and infrastructure issues across AWS and Kubernetes environments. Implement best practices for security, reliability, availability, and scalability. CI/CD Azure-to-AWS Migration Build and maintain CI/CD pipelines for ML and AI services.

Develop and manage GitHub Actions and Azure DevOps pipelines using YAML. Migrate repositories and CI/CD workflows from Azure DevOps to GitHub/AWS . Automate build, test, containerization, deployment, and release processes.

Establish deployment strategies across development, staging, and production environments. ML Service Deployment Deploy and manage ML services across AWS EKS/ECS and SageMaker . Build and maintain Docker containers and Kubernetes deployments.

Manage environment segregation and configuration across Dev, QA, and Production. Develop and maintain Kubernetes manifests and Helm charts . Troubleshoot ML service deployment, networking, scaling, and runtime issues.

App Runner to EKS Migration Lead migration of existing services from AWS App Runner to Amazon EKS . Containerize applications and develop Kubernetes manifests/Helm charts. Design appropriate Kubernetes architecture, networking, ingress, scaling, and deployment strategies.

Ensure minimal service disruption during migration and establish operational best practices on EKS. Self-Hosted LLM AI Infrastructure Deploy and manage self-hosted Large Language Models and inference services. Work with model serving frameworks such as vLLM .

Design containerized infrastructure for GPU-based model serving and inference. Manage model versions, deployments, configurations, and rollback strategies. Support migration of ML services from managed APIs/services to self-hosted models .

Work with engineering teams on API integration and inference infrastructure. Databricks Administration Administer Databricks workspaces, clusters, permissions, and access controls . Manage cluster configuration, policies, and resource utilization.

Support LMI Insights and related ML/AI workloads. Troubleshoot Databricks infrastructure and connectivity issues. Implement appropriate security and access-control practices.

Elasticsearch Infrastructure Design, deploy, and manage Elasticsearch clusters . Perform cluster sizing, scaling, configuration, and performance optimization. Manage indices, mappings, retention, and data lifecycle requirements.

Support Kibana configuration, dashboards, and troubleshooting. Monitor Elasticsearch health, capacity, and performance. Monitoring, Reliability Auto-Scaling Implement monitoring and observability for Kubernetes, AWS, and ML services.

Use Prometheus, Grafana, and AWS CloudWatch for monitoring and alerting. Configure Kubernetes HPA/VPA and other auto-scaling mechanisms. Establish proactive alerting for infrastructure and application health.

Perform capacity planning and resource optimization. Identify opportunities for AWS infrastructure and compute cost optimization . Infrastructure as Code Automation Build and maintain infrastructure using Terraform .

Develop reusable Terraform modules for AWS and Kubernetes infrastructure. Manage Kubernetes deployments using Helm charts . Automate infrastructure provisioning, configuration, deployments, and operational tasks.

Maintain infrastructure documentation and deployment standards. Cross-Team Collaboration Partner closely with Data Engineering, ML Engineering, Data Science, and Software Engineering teams. Understand data pipelines, SQL, APIs, and ML service architecture sufficiently to troubleshoot end-to-end workflows.

Coordinate infrastructure requirements for new ML models and services. Participate in production incident resolution, root-cause analysis, and continuous improvement. Establish engineering standards around deployment, monitoring, security, and operational ownership.

You've got what it takes if you have... Required Skills Experience 5+ years of experience in DevOps, Cloud Infrastructure, SRE, or Platform Engineering. Strong hands-on experience with AWS .

Strong experience administering Amazon EKS and Kubernetes in production. Hands-on experience with: AWS EKS AWS SageMaker AWS Bedrock AWS ECS AWS App Runner Kubernetes Docker NGINX / AWS ALB Ingress Cloudflare / Cloudflare Tunnels Strong experience with Terraform and Helm . Strong experience developing CI/CD pipelines using GitHub Actions and/or Azure DevOps.

Strong YAML scripting and Git experience. Experience migrating CI/CD pipelines and repositories from Azure to AWS/GitHub . Experience deploying and operating ML/AI services.

Experience with self-hosted LLM/model serving , preferably vLLM . Experience with GPU-based workloads is highly desirable. Experience with Databricks administration .

Experience managing Elasticsearch and Kibana . Experience with Prometheus, Grafana, and CloudWatch . Strong understanding of Kubernetes HPA/VPA, networking, ingress, DNS, and service discovery .

Strong understanding of cloud networking fundamentals. Experience with production troubleshooting, monitoring, capacity planning, and cost optimization. Strong understanding of security, IAM, secrets management, and access control.

Preferred / Nice-to-Have Skills Experience supporting Generative AI / LLM platforms . Experience with GPU infrastructure and NVIDIA/CUDA environments. Experience with model lifecycle and model version management.

Experience migrating workloads between managed cloud services and Kubernetes. Experience with AWS networking such as VPC, load balancers, security groups, and Route 53. Experience with API gateways and microservice architectures.

Experience with Python or shell scripting for infrastructure automation. Experience working with Data Engineering and ML teams in a production environment. What You'll Own AWS ML/AI infrastructure EKS cluster administration and upgrades ML service deployment and production operations CI/CD automation App Runner → EKS migration Self-hosted LLM infrastructure and vLLM Databricks platform administration Elasticsearch infrastructure Monitoring and auto-scaling Terraform and Helm-based infrastructure automation Cloud cost, reliability, and performance optimization Ideal Candidate The ideal candidate is a hands-on infrastructure engineer who can independently take an ML/AI service from containerization → CI/CD → AWS infrastructure → EKS deployment → monitoring → scaling → production support .

They should be comfortable working across both traditional DevOps infrastructure and modern AI/ML infrastructure , and should be able to collaborate closely with Data Engineering and ML teams while taking ownership of the underlying platform. #LI-Onsite

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
1,166,878 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

DevOps
Similar stack
Same company
Hyderabad
≈ $11k – $29k per year (Estimated) • In office • 3+ years exp • Hyderabad
Python
DevOps
BGP
Apply
≈ $11k – $27k per year (Estimated) • In office • 3+ years exp • Hyderabad
Python
SQL
PowerShell
Databases
MySQL
PostgreSQL
Oracle
CockroachDB
MS SQL
DevOps
GCP
Azure
CI/CD
AWS
Apply
≈ $11k – $27k per year (Estimated) • In office • 2+ years exp • Mumbai
AI/ML
Machine Learning
DevOps
CI/CD
Management
Agile
Apply
≈ $12k – $30k per year (Estimated) • In office • Full-Time • 4+ years exp • Navi Mumbai
JavaScript
SQL
AI/ML
Model Context Protocol
DevOps
Rest API
SOAP
Apply
≈ $17k – $38k per year (Estimated) • In office • Full-Time • 6+ years exp • Chennai
DevOps
SLURM
Configuration Management
OpenStack
HPC
Linux
TCP/IP
Management
Agile
Apply
Sr. Data Engineer 2 hours ago
≈ $20k – $41k per year (Estimated) • In office • Full-Time • Bachelor's Degree • Hyderabad
Python
Java
SQL
Scala
Databases
Snowflake
Databricks
Delta Lake
Apache Kafka
Microsoft Fabric
AI/ML
Machine Learning
DevOps
GCP
Azure
CI/CD
Git
AWS
Amazon Kinesis
Analytics
Power BI
ETL/ELT
Apply
Senior Java Developer 2 hours ago
≈ $16k – $43k per year (Estimated) • In office • Full-Time • 3+ years exp • Bachelor's Degree • Bengaluru
Java
SQL
Java
Spring Framework
Spring Boot
Apache Tomcat
Databases
MongoDB
PostgreSQL
Oracle
DynamoDB
Apache Kafka
AI/ML
LLM
DevOps
Rest API
Terraform
CI/CD
Jenkins
Docker
Kubernetes
Concourse
Cybersecurity
SOC 2
GDPR
Management
Outlook
Agile
Apply
Back-End Developer 1 hour ago
In office • Full-Time • Romania
SQL
DevOps
Git
Management
Slack
Apply
IT Developer 1 hour ago
Remote (likely United States, Canada, EST hours) • Full-Time
Python
PHP
TypeScript
SQL
PHP
WordPress
DevOps
GitHub
Analytics
Microsoft Excel
Management
Asana
Slack
Monday.com
Airtable
Google Workspace
Zapier
Google Sheets
Microsoft Office
Apply
≈ $124k – $239k per year (Estimated) • In office • Full-Time • 3+ years exp • Zurich
Python
JavaScript
TypeScript
SQL
Python
FastAPI
Databases
PostgreSQL
Frontend
React.js
DevOps
AWS
Amazon S3
QA
Playwright
Apply
$19k per year • Hybrid • Internship • Kraków
AI/ML
Copilot
Cursor
Claude
DevOps
Terraform
Ansible
Jenkins
Git
AWS
Docker
Linux
Management
Confluence
Jira
Apply
≈ $16k – $36k per year (Estimated) • In office • 8+ years exp • Hyderabad
SQL
Databases
Oracle
QA
Postman
SoapUI
Apply
Data Analyst 4 hours ago
≈ $16k – $32k per year (Estimated) • In office • 5+ years exp • Pune
Python
SQL
Databases
Snowflake
Analytics
Tableau
Master Data Management
Apply
Senior Data Engineer 4 hours ago
≈ $20k – $41k per year (Estimated) • In office • Hyderabad
Python
Java
SQL
C++
Scala
Databases
Redis
Snowflake
Neo4j
Databricks
Cassandra
Presto
AI/ML
Hadoop
Spark
Airflow
dbt
Machine Learning
DevOps
GCP
Azure
AWS
Analytics
ETL/ELT
Matillion
Fivetran
Apply
≈ $20k – $42k per year (Estimated) • Hybrid • 5+ years exp • Hyderabad
Java
SQL
Java
Spring Boot
Databases
MySQL
PostgreSQL
Redis
Weaviate
Milvus
Pinecone
RabbitMQ
Apache Kafka
AI/ML
Embeddings
Prompt Engineering
NLP
AWS Bedrock
LLM
RAG
Semantic Search
Hybrid Search
OpenAI
LLMOps
Semantic Search
Frontend
GraphQL
DevOps
Rest API
GCP
Azure
CI/CD
AWS
Docker
Kubernetes
Platform Engineering
Apply
≈ $18k – $39k per year (Estimated) • Hybrid • 7+ years exp • Bachelor's Degree • Pune • Hyderabad
Python
Java
SQL
Databases
Apache Kafka
Google BigQuery
BigQuery
AI/ML
Machine Learning
DevOps
GCP
CI/CD
Analytics
ETL/ELT
DataStage
Looker
Apply
≈ $15k – $33k per year (Estimated) • In office • 5+ years exp • Bachelor's Degree • Hyderabad
DevOps
Windows Server
Windows
Apply
In office • Bachelor's Degree • Hyderabad
Databases
Redis
DevOps
Azure
Kubernetes
Apply
In office • Bachelor's Degree • Hyderabad
Python
PowerShell
DevOps
Ansible
OpenShift
VMWare
Configuration Management
Incident Management
SLI/SLO/SLA
Cybersecurity
Hybrid Analysis
Apply
In office • Bachelor's Degree • Hyderabad
DevOps
Terraform
Ansible
CI/CD
AWS
Configuration Management
SLI/SLO/SLA
GitLab
API Gateway
Apply
See all jobs
This is one of many
1,166,878 more open roles from verified company boards, updated every day.