Overview
Company
Profile match
Impact
Conditions
Benefits
Hiring process
Similar jobs

Mastech Digital

Mastech Digital is an IT company that accelerates business operations via IT staffing and digital transformation services like data management and analytics, Salesforce/CRM, SAP HANA, and digital learning.

We are looking for a seasoned AIOps Engineer who sits at the intersection of platform reliability, AI/ML operations, and intelligent automation. You will design and operate AI-driven observability pipelines, build self-healing infrastructure, and embed machine-learning models into DevOps toolchains, turning telemetry noise into actionable intelligence at scale. As a trusted member of the Platform Engineering team, you will own the full lifecycle of AIOps tooling: from anomaly-detection models and intelligent alerting to predictive capacity management and LLM-assisted incident response, all running on cloud-native, containerised infrastructure.

The candidate will have responsibilities across the following functions:

AIOps and Intelligent Automation:

  • Design and operate AI/ML-powered observability platforms (anomaly detection, root-cause analysis, predictive alerting) using tools such as Dynatrace, New Relic, Moogsoft, or open-source equivalents.
  • Build and maintain Python-based ML pipelines (scikit-learn, PyTorch, or TensorFlow) for log analytics, metric forecasting, and event correlation.
  • Develop intelligent auto-remediation runbooks triggered by model-detected incidents, reducing MTTR by 40% or more.
  • Integrate LLM-assisted copilots (OpenAI / Azure OpenAI) into incident management workflows for automated triage and war-room summaries.

DevOps and Platform Engineering:

  • Own CI/CD pipeline design and maintenance using GitHub Actions, GitLab CI, or Jenkins; enforce shift-left security with SAST/DAST, SBOM, and Trivy scans.
  • Provision, scale, and manage Kubernetes clusters (EKS / AKS / GKE) using Helm, Kustomize, and ArgoCD with GitOps workflows.
  • Manage Infrastructure-as-Code using Terraform and Ansible across multi-cloud environments (AWS primary, Azure secondary).
  • Champion containerisation best practices: image hardening, multi-stage builds, and supply-chain security with Cosign/Sigstore.

Site Reliability Engineering:

  • Define, implement, and operationalise SLOs, SLIs, and error-budget burn policies aligned to business objectives.
  • Architect and maintain a four-layer observability stack: metrics (Prometheus/Thanos), logs (OpenSearch/Loki), traces (Jaeger/Tempo), and dashboards (Grafana).
  • Lead blameless post-mortem processes and own the incident management lifecycle end-to-end.
  • Drive capacity planning using ML-based demand forecasting models integrated with KEDA autoscaling policies.

Python Engineering and Tooling:

  • Write production-grade Python automation: alerting integrations, CMDB reconciliation, drift-detection daemons, and Slack/PagerDuty bots.
  • Build internal CLI tools and REST APIs (FastAPI / Flask) to expose AIOps insights to engineering teams.
  • Maintain data pipelines feeding telemetry into feature stores and model retraining loops.

Requirements:

  • 5-7 years of hands-on experience in DevOps, SRE, or Platform Engineering roles.
  • 2+ years in an AIOps or MLOps capacity deploying and monitoring ML models in production.
  • Bachelor's or Master's degree in Computer Science, Information Technology, or equivalent.

Technical Skills:

  • Python (3 x): proficient in scripting, automation, REST APIs, ML libraries, and unit testing (pytest).
  • Kubernetes: workload design, RBAC, HPA/KEDA, network policies, and multi-tenancy.
  • Observability stack: Prometheus, Grafana, Alertmanager, OpenTelemetry, and distributed tracing.
  • CI/CD: GitHub Actions or GitLab CI pipeline authoring, environment promotion, and gate policies.
  • IaC: Terraform (modules, state management, remote backends) and Helm chart authoring.
  • Cloud: Strong hands-on Solution Architect experience on AWS and Azure production-grade infrastructure design, multi-cloud networking, security, and cost governance.
  • Incident Management: PagerDuty, Opsgenie, or VictorOps integrated with AIOps runbooks.

Technical Skills Strong Advantage:

  • ML frameworks: scikit-learn, Prophet, or PyTorch for time-series anomaly detection.
  • Kafka / NATS JetStream for high-throughput telemetry streaming.
  • Service mesh: Istio or Linkerd for traffic observability and canary rollouts.
  • Security tooling: Trivy, Falco, OPA/Gatekeeper, OWASP LLM Top 10 awareness.

LLM / SLM Optimisation:

  • Hands-on experience deploying and optimising Large Language Models (GPT-4o, Claude, Llama 3) and Small Language Models (Phi-3 Mistral, Gemma) in production AIOps pipelines.
  • Model quantisation techniques: GGUF/GGML, GPTQ, AWQ, reducing inference footprint by 2-4 without significant accuracy loss.
  • Prompt engineering and chain-of-thought optimisation for ops-domain tasks such as log summarisation, RCA generation, and runbook synthesis.
  • Retrieval-Augmented Generation (RAG): building domain-specific vector stores (pgvector, MongoDB Atlas Vector Search, Pinecone) backed by incident knowledge bases.
  • LLM inference serving: vLLM, Ollama, or TGI (Text Generation Inference) on Kubernetes autoscaling with KEDA based on token throughput.
  • Fine-tuning SLMs on internal ops datasets using LoRA / QLoRA for domain-specific intent classification (alert triage, ticket routing, anomaly labelling).
  • LLM observability: token usage tracking, latency p99 SLOs, hallucination guardrails, and prompt injection detection using tools such as Langfuse or LangSmith.
  • Context-window management and cost-optimisation strategies: prompt compression, semantic caching, and batching via Anthropic/OpenAI batch APIs.
  • Model routing and A/B evaluation frameworks switching between frontier LLMs and local SLMs based on latency, cost, and task complexity signals.

Behavioural Competencies

  • Ownership mindset: you treat production systems as if they were your own product.
  • Data-driven decision-making: you instrument first, optimise second.
  • Collaborative problem-solver comfortable pairing with data science, security, and product teams.
  • Strong written communication, able to author RCAs, design docs, and architecture decision records.
  • Growth orientation: curious about emerging AI tooling and proactive in knowledge sharing.

Nice to Have:

  • Experience on high-scale consumer platforms (50M+ users, sub-200ms p99 SLOs).
  • Contributions to open-source AIOps or observability projects.
  • Familiarity with MCP (Model Context Protocol) agentic architectures.
  • AWS / GCP / Azure certifications at Professional or Speciality level.
  • Knowledge of FinOps tooling (Kubecost, AWS Cost Explorer API integration).

Recommended for you based on this role

Similar stack
Same company
In your city
MLOps Engineer 15 days ago
5+ year exp • Bengaluru
Python
Python
FastAPI
Flask
Databases
Apache Kafka
NATS
OpenSearch
pgvector
Pinecone
PostgreSQL
AI/ML
Anomaly Detection
LangChain
LangGraph
LLM
Model Context Protocol
Ollama
Prompt Engineering
RAG
TGI
vLLM
DevOps
AIOps
Alertmanager
Amazon EC2
Amazon EKS
ArgoCD
AWS
Azure
Azure AKS
CI/CD
Docker
FinOps
GCP
GitHub Actions
GitLab CI
GitOps
Google GKE
Grafana
Helm
Incident Management
Istio
Jaeger
Jenkins
KEDA
Kubernetes
Linkerd
Loki
OpenTelemetry
PagerDuty
Platform Engineering
Prometheus
Rest API
Self-Healing
Terraform
Vector
Cybersecurity
Falco
Trivy
Management
Slack
QA
Pytest
Apply
Bengaluru
Python
AI/ML
AI Agents
Embeddings
Hallucination
LLM
RAG
Synthetic Data
DevOps
AWS
Azure
CI/CD
Docker
GCP
Kubernetes
Vector
Apply
Software Engineer 23 days ago
5+ year exp • Bachelor's Degree • Bengaluru
Python
AI/ML
AI Agents
Fine-tuning
Prompt Engineering
RAG
Semantic Search
DevOps
AWS
Azure
CI/CD
GCP
Apply
Fullstack Engineer 23 days ago
5+ year exp • Bengaluru
JavaScript
Node JS
Python
SQL
TypeScript
Python
FastAPI
Flask
Pydantic
AI/ML
Copilot
Function Calling
LangChain
LLM
Frontend
Material UI
Next.js
React.js
Tailwind CSS
Vue.js
DevOps
AWS
Azure
CI/CD
Rest API
WebSockets
QA
Jest
Pytest
Apply
Career impact
Discover how this job can transform your career
Get a personal career forecast for this job - salary uplift, next-level role, skill boost and a 3-year financial impact, all calculated from your profile.
Personal salary uplift vs. your current pay
Your 3-year career trajectory
Skills you will level up in this role
3-year financial impact in dollars
Create free account
Free forever • Less than a minute • No credit card

Work setup

Location
Bengaluru