1,461,231open jobs
87,409companies
230,564added this week
Browse all
Location
In office
Seniority
Middle · 3+ years exp

Confirmed on the employer's own hiring board on Oct 11, 2026. First seen by Alion on Oct 8, 2026.

Overview
Company
Impact
Profile match
HCLTech is a major Indian multinational information technology (IT) services and consulting company headquartered in Noida, Uttar Pradesh. Spun off from the original HCL Group in 1991, it ranks as one of India's largest technology companies alongside firms like TCS, Infosys, and Wipro.

Job Summary

GenAI & ML Operations Engineer

Role Summary

The GenAI & ML Operations Engineer is responsible for ensuring the reliability, availability, observability, and operational support of Production Generative AI and Machine Learning solutions. The role focuses on monitoring AI/ML platforms, resolving incidents, implementing observability, supporting model upgrades, executing minor enhancements, and driving operational excellence across AI products and services.

The engineer works closely with Data Scientists, ML Engineers, Platform Engineers, Product Teams, and Business Stakeholders to maintain stable, scalable, and well-governed AI solutions in production.

Key Responsibilities

AI/ML Production Operations

  • Monitor and support production GenAI and ML workloads, Perform daily operational health checks and service validations.
  • Troubleshoot and resolve operational issues impacting AI services, Support L1/L2/L3 incident management activities.
  • Ensure SLA, availability, and performance objectives are met.

Observability & Monitoring

  • Design, implement, and maintain AI/ML observability solutions, Create dashboards, reports, and operational scorecards.
  • Configure alerts for failures, threshold breaches, latency, data quality, model performance, and infrastructure issues.
  • Monitor model health, drift, performance, usage, and cost metrics, continuously improve monitoring coverage and alert effectiveness.

Incident & Problem Management

  • Respond to production incidents and service disruption, Conduct root cause analysis (RCA) and document findings.
  • Drive preventive actions to reduce recurring incidents, Participate in major incident support and service restoration activities.

Model Lifecycle Support

  • Support AI/ML model upgrades, version management, Monitor production behaviour following model changes and perform post-deployment validation.
  • Support retraining, tuning, and model refresh initiatives.

Continuous Improvement & Enhancements

  • Deliver minor enhancements to AI applications and operational tooling, Automate repetitive operational tasks where feasible.
  • Identify opportunities to improve reliability, efficiency, and supportability, Contribute to AI Ops, MLOps, and operational maturity initiatives.

Documentation & Governance

  • Maintain runbooks, SOPs, knowledge articles, and support documentation, Document incidents, problems, technical debt, risks, and remediation plans.
  • Support change management and release governance processes, Ensure operational compliance with security and Responsible AI requirements.

Required Skills & Experience

Technical Skills

  • Generative AI, LLMs, RAG, and AI application support.
  • Machine Learning lifecycle and MLOps concepts.
  • Monitoring and Observability Platforms (Datadog, Dynatrace.).
  • Cloud Platforms (GCP), Kubernetes and containerized workloads, MemoryStore, Tracing, Storage, GAR.
  • Python and automation scripting.

Operational Skills

  • Incident Management, Problem Management, Change Management
  • Root Cause Analysis
  • Service Reliability Engineering (SRE)
  • Operational Reporting and Governance, Stakeholder Communication

Preferred Experience

  • 8+ years in IT Operations, DevOps, SRE, Cloud Operations, MLOps, or AI Operations.
  • 3+ years of experience with GenAI / LLM platforms / ML Models
  • Experience in working with Databricks jobs & pipelines, Catalog, compute, models, SQL warehouse.
  • GCP workloads, GCP Deployments, Vertex AI, Gemini Enterprise Agent Development Kit, BigQuery, BigTable, Redis, Postgres
  • Experience supporting production AI/ML solutions, working knowledge of ITIL-based operational processes
  • Familiarity with GenAI platforms, LLM-based applications, and AI observability practices.

Key Responsibilities

GenAI & ML Operations Engineer

Role Summary

The GenAI & ML Operations Engineer is responsible for ensuring the reliability, availability, observability, and operational support of Production Generative AI and Machine Learning solutions. The role focuses on monitoring AI/ML platforms, resolving incidents, implementing observability, supporting model upgrades, executing minor enhancements, and driving operational excellence across AI products and services.

The engineer works closely with Data Scientists, ML Engineers, Platform Engineers, Product Teams, and Business Stakeholders to maintain stable, scalable, and well-governed AI solutions in production.

Key Responsibilities

AI/ML Production Operations

  • Monitor and support production GenAI and ML workloads, Perform daily operational health checks and service validations.
  • Troubleshoot and resolve operational issues impacting AI services, Support L1/L2/L3 incident management activities.
  • Ensure SLA, availability, and performance objectives are met.

Observability & Monitoring

  • Design, implement, and maintain AI/ML observability solutions, Create dashboards, reports, and operational scorecards.
  • Configure alerts for failures, threshold breaches, latency, data quality, model performance, and infrastructure issues.
  • Monitor model health, drift, performance, usage, and cost metrics, continuously improve monitoring coverage and alert effectiveness.

Incident & Problem Management

  • Respond to production incidents and service disruption, Conduct root cause analysis (RCA) and document findings.
  • Drive preventive actions to reduce recurring incidents, Participate in major incident support and service restoration activities.

Model Lifecycle Support

  • Support AI/ML model upgrades, version management, Monitor production behaviour following model changes and perform post-deployment validation.
  • Support retraining, tuning, and model refresh initiatives.

Continuous Improvement & Enhancements

  • Deliver minor enhancements to AI applications and operational tooling, Automate repetitive operational tasks where feasible.
  • Identify opportunities to improve reliability, efficiency, and supportability, Contribute to AI Ops, MLOps, and operational maturity initiatives.

Documentation & Governance

  • Maintain runbooks, SOPs, knowledge articles, and support documentation, Document incidents, problems, technical debt, risks, and remediation plans.
  • Support change management and release governance processes, Ensure operational compliance with security and Responsible AI requirements.

Required Skills & Experience

Technical Skills

  • Generative AI, LLMs, RAG, and AI application support.
  • Machine Learning lifecycle and MLOps concepts.
  • Monitoring and Observability Platforms (Datadog, Dynatrace.).
  • Cloud Platforms (GCP), Kubernetes and containerized workloads, MemoryStore, Tracing, Storage, GAR.
  • Python and automation scripting.

Operational Skills

  • Incident Management, Problem Management, Change Management
  • Root Cause Analysis
  • Service Reliability Engineering (SRE)
  • Operational Reporting and Governance, Stakeholder Communication

Preferred Experience

  • 8+ years in IT Operations, DevOps, SRE, Cloud Operations, MLOps, or AI Operations.
  • 3+ years of experience with GenAI / LLM platforms / ML Models
  • Experience in working with Databricks jobs & pipelines, Catalog, compute, models, SQL warehouse.
  • GCP workloads, GCP Deployments, Vertex AI, Gemini Enterprise Agent Development Kit, BigQuery, BigTable, Redis, Postgres
  • Experience supporting production AI/ML solutions, working knowledge of ITIL-based operational processes
  • Familiarity with GenAI platforms, LLM-based applications, and AI observability practices.

Skill Requirements

GenAI & ML Operations Engineer

Role Summary

The GenAI & ML Operations Engineer is responsible for ensuring the reliability, availability, observability, and operational support of Production Generative AI and Machine Learning solutions. The role focuses on monitoring AI/ML platforms, resolving incidents, implementing observability, supporting model upgrades, executing minor enhancements, and driving operational excellence across AI products and services.

The engineer works closely with Data Scientists, ML Engineers, Platform Engineers, Product Teams, and Business Stakeholders to maintain stable, scalable, and well-governed AI solutions in production.

Key Responsibilities

AI/ML Production Operations

  • Monitor and support production GenAI and ML workloads, Perform daily operational health checks and service validations.
  • Troubleshoot and resolve operational issues impacting AI services, Support L1/L2/L3 incident management activities.
  • Ensure SLA, availability, and performance objectives are met.

Observability & Monitoring

  • Design, implement, and maintain AI/ML observability solutions, Create dashboards, reports, and operational scorecards.
  • Configure alerts for failures, threshold breaches, latency, data quality, model performance, and infrastructure issues.
  • Monitor model health, drift, performance, usage, and cost metrics, continuously improve monitoring coverage and alert effectiveness.

Incident & Problem Management

  • Respond to production incidents and service disruption, Conduct root cause analysis (RCA) and document findings.
  • Drive preventive actions to reduce recurring incidents, Participate in major incident support and service restoration activities.

Model Lifecycle Support

  • Support AI/ML model upgrades, version management, Monitor production behaviour following model changes and perform post-deployment validation.
  • Support retraining, tuning, and model refresh initiatives.

Continuous Improvement & Enhancements

  • Deliver minor enhancements to AI applications and operational tooling, Automate repetitive operational tasks where feasible.
  • Identify opportunities to improve reliability, efficiency, and supportability, Contribute to AI Ops, MLOps, and operational maturity initiatives.

Documentation & Governance

  • Maintain runbooks, SOPs, knowledge articles, and support documentation, Document incidents, problems, technical debt, risks, and remediation plans.
  • Support change management and release governance processes, Ensure operational compliance with security and Responsible AI requirements.

Required Skills & Experience

Technical Skills

  • Generative AI, LLMs, RAG, and AI application support.
  • Machine Learning lifecycle and MLOps concepts.
  • Monitoring and Observability Platforms (Datadog, Dynatrace.).
  • Cloud Platforms (GCP), Kubernetes and containerized workloads, MemoryStore, Tracing, Storage, GAR.
  • Python and automation scripting.

Operational Skills

  • Incident Management, Problem Management, Change Management
  • Root Cause Analysis
  • Service Reliability Engineering (SRE)
  • Operational Reporting and Governance, Stakeholder Communication

Preferred Experience

  • 8+ years in IT Operations, DevOps, SRE, Cloud Operations, MLOps, or AI Operations.
  • 3+ years of experience with GenAI / LLM platforms / ML Models
  • Experience in working with Databricks jobs & pipelines, Catalog, compute, models, SQL warehouse.
  • GCP workloads, GCP Deployments, Vertex AI, Gemini Enterprise Agent Development Kit, BigQuery, BigTable, Redis, Postgres
  • Experience supporting production AI/ML solutions, working knowledge of ITIL-based operational processes
  • Familiarity with GenAI platforms, LLM-based applications, and AI observability practices.

Other Requirements

None

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
1,461,231 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Support
Similar stack
Same company
In your city
$187k – $308k per year • Equity • Hybrid • Full-Time • 8+ years exp • Bachelor's Degree • Milpitas
Python
AI/ML
Copilot
Cursor
Cybersecurity
Threat Modeling
Management
Agile
Apply
≈ $40k – $105k per year (Estimated) • In office • Top Secret • 5+ years exp • The Hague • Leeds
DevOps
GCP
Azure
AWS
Management
Agile
ITIL
Apply
≈ $16k – $40k per year (Estimated) • In office • Full-Time • Komotini
DevOps
Windows
Apply
≈ $32k – $53k per year (Estimated) • Remote (Canada) • Full-Time • 2+ years exp • Associate's Degree • Canada
Apply
In office • Full-Time • Regensdorf
Apply
Lead SRE 2 days ago
$153k – $255k per year • Remote (likely United States) • PhD • New York
Python
Bash
Databases
OpenSearch
AI/ML
vLLM
AI Agents
Langfuse
Ollama
AgentOps
LLM
RAG
Hallucination
LLMOps
Multi-Agent Systems
DevOps
OpenShift
OpenTelemetry
CI/CD
Docker
Kubernetes
Self-Healing
Incident Management
Linux
Management
Agile
Apply
≈ $120k – $230k per year (Estimated) • Remote (United States) • Full-Time • 8+ years exp • United States
Python
Go
Java
TypeScript
Databases
Snowflake
Databricks
AI/ML
AI Agents
AgentOps
RAG
LLMOps
Human-in-the-Loop
LLM Guardrails
Tool Use
DevOps
GCP
Azure
CI/CD
AWS
FinOps
Management
ServiceNow
Apply
≈ $23k – $46k per year (Estimated) • In office • 5+ years exp • Bengaluru
Python
Databases
Databricks
AI/ML
Machine Learning
DevOps
Azure
Apply
$70k per year • In office • Full-Time • Bachelor's Degree • Australia
Python
JavaScript
DevOps
CI/CD
Git
Apply
≈ $16k – $36k per year (Estimated) • In office • 6+ years exp • Bachelor's Degree • Mumbai
Python
SQL
DevOps
CI/CD
Linux
Apply
Support Engineer 3 days ago
In office • 2+ years exp • Bachelor's Degree
Python
C++
DevOps
CI/CD
Docker
Linux
TCP/IP
Robotics
ROS
EtherCAT
Teleoperation
Design
SolidWorks
Fusion 360
Apply
In office • 3+ years exp
Python
SQL
Databases
PostgreSQL
Redis
Databricks
Google BigQuery
Google Bigtable
BigQuery
AI/ML
Vertex AI
Gemini
LLM
RAG
Machine Learning
DevOps
GCP
Datadog
Dynatrace
Kubernetes
AIOps
Incident Management
SLI/SLO/SLA
Management
ITIL
Apply
In office
DevOps
Kubernetes
Management
ITSM
Apply
In office • 5+ years exp • Bachelor's Degree
Apply
In office • 6+ years exp
Python
Go
JavaScript
TypeScript
SQL
Databases
PostgreSQL
Redis
RabbitMQ
Apache Kafka
Frontend
GraphQL
DevOps
Rest API
GCP
Azure
CI/CD
AWS
Docker
Kubernetes
Analytics
ETL/ELT
Apply
See all jobs
This is one of many
1,461,231 more open roles from verified company boards, updated every day.