Overview
Company
Profile match
Impact
Conditions
Benefits
Hiring process
Similar jobs

Mastech Digital

Mastech Digital is an IT company that accelerates business operations via IT staffing and digital transformation services like data management and analytics, Salesforce/CRM, SAP HANA, and digital learning.

We are looking for an experienced AIOps Engineer who is passionate about building intelligent, resilient, and highly automated cloud platforms. This role sits at the intersection of Platform Engineering, DevOps, Site Reliability Engineering (SRE), and AI/ML Operations, where you'll build AI-driven observability solutions, intelligent automation pipelines, and self-healing infrastructure for enterprise-scale applications. You'll work closely with Platform, AI, and Product Engineering teams to embed machine learning into operational workflows, automate incident response, improve platform reliability, and enable proactive infrastructure management using modern AI technologies.

The candidate will have responsibilities across the following functions:

AIOps and Intelligent Automation:

  • Design and implement AI-powered observability platforms for anomaly detection, predictive alerting, event correlation, and root-cause analysis.
  • Build Python-based ML pipelines for log analytics, telemetry processing, and operational intelligence.
  • Develop intelligent auto-remediation workflows to reduce manual intervention and improve platform availability.
  • Integrate LLM-powered assistants into incident management workflows for automated triage, RCA generation, and operational insights.
  • Implement predictive capacity planning using machine learning models.

Platform Engineering and DevOps:

  • Design, build, and maintain CI/CD pipelines using GitHub Actions, GitLab CI, or Jenkins.
  • Provision and manage Kubernetes environments (EKS/AKS/GKE) using GitOps methodologies.
  • Build Infrastructure-as-Code using Terraform and automate cloud provisioning.
  • Develop secure container platforms using Docker, Helm, and Kubernetes best practices.
  • Manage production deployments across AWS and Azure cloud environments.

Site Reliability Engineering:

  • Define and monitor SLIs, SLOs, and error budgets.
  • Build enterprise observability platforms using Prometheus, Grafana, OpenTelemetry, Loki/OpenSearch, and Jaeger/Tempo.
  • Lead incident management, root cause analysis, and post-incident improvements.
  • Drive platform reliability, scalability, and operational excellence.

Python Development:

  • Develop production-grade automation tools, REST APIs, internal utilities, and operational dashboards using Python.
  • Build integrations with monitoring, alerting, Slack, PagerDuty, and CMDB platforms.
  • Develop data pipelines supporting telemetry ingestion and AI model retraining workflows.

Requirements:

  • 5-7 years of experience in Platform Engineering, DevOps, or Site Reliability Engineering.
  • Strong expertise in Python, Kubernetes, Terraform, and Cloud Infrastructure.
  • Experience building production-grade AI-enabled operational platforms.
  • Passion for automation, reliability, and cloud-native engineering.
  • Excellent debugging, problem-solving, and communication skills.
  • Ability to collaborate across Platform, AI, Security, and Product Engineering teams.
  • Platform and DevOps: Kubernetes (EKS, AKS, or GKE), Docker, Helm, GitOps (ArgoCD preferred), CI/CD (GitHub Actions, GitLab CI, Jenkins), Terraform, Infrastructure Automation.
  • Programming: Python, REST API Development (FastAPI or Flask), Automation & Scripting, Unit Testing (pytest).
  • Observability: Prometheus, Grafana, Alertmanager, OpenTelemetry, Distributed Tracing, Log Aggregation Platforms.
  • Cloud: AWS (EC2 IAM, EKS, CloudWatch, VPC), Azure (preferred), Multi-cloud Infrastructure.
  • Site Reliability: Incident Management, Capacity Planning, Performance Monitoring, Reliability Engineering, Production Operations.

AI / MLOps Skills:

  • Experience deploying AI/ML models into production environments.
  • Hands-on experience with OpenAI, Azure OpenAI, or other LLM platforms.
  • Knowledge of Retrieval-Augmented Generation (RAG) architectures.
  • Experience with vector databases such as Pinecone, pgvector, or MongoDB Atlas Vector Search.
  • Prompt engineering and LLM optimisation.
  • Experience with model serving platforms such as vLLM, Ollama, or Text Generation Inference (TGI).
  • Familiarity with LangChain, LangGraph, or AI orchestration frameworks.
  • Understanding of model monitoring, prompt evaluation, and LLM observability.
  • Good to Have: MLOps experience, Kafka or NATS, Istio or Linkerd, KEDA, Security tooling (Trivy, Falco, OPA), FinOps tooling, MCP (Model Context Protocol), AWS, Azure, or GCP Certifications, Contributions to open-source DevOps or AI projects.

Recommended for you based on this role

Similar stack
Same company
In your city
AIops Engineer 14 days ago
5+ year exp • Bachelor's Degree • Bengaluru
Python
Python
FastAPI
Flask
Databases
Apache Kafka
NATS
OpenSearch
pgvector
Pinecone
PostgreSQL
AI/ML
Anomaly Detection
AWQ
Chain-of-Thought
Claude
Fine-tuning
Gemma
GGUF
GPTQ
Hallucination
Langfuse
LangSmith
Llama
LLM
LoRA
Mistral
Model Context Protocol
Ollama
Phi
Prompt Engineering
PyTorch
QLoRA
RAG
Scikit-learn
TensorFlow
TGI
vLLM
PEFT
DevOps
AIOps
Alertmanager
Amazon EKS
Ansible
ArgoCD
AWS
Azure
Azure AKS
CI/CD
Dynatrace
FinOps
GCP
GitHub Actions
GitLab CI
GitOps
Google GKE
Grafana
Helm
Incident Management
Istio
Jaeger
Jenkins
KEDA
Kubernetes
Kustomize
Linkerd
Loki
New Relic
OpenTelemetry
Opsgenie
PagerDuty
Platform Engineering
Prometheus
Rest API
Self-Healing
Service Mesh
Shift-Left
Terraform
Thanos
Vector
Cybersecurity
Cosign
Falco
SBOM
Shift-Left Security
Sigstore
Trivy
Management
Slack
QA
Pytest
Apply
Bengaluru
Python
AI/ML
AI Agents
Embeddings
Hallucination
LLM
RAG
Synthetic Data
DevOps
AWS
Azure
CI/CD
Docker
GCP
Kubernetes
Vector
Apply
Software Engineer 23 days ago
5+ year exp • Bachelor's Degree • Bengaluru
Python
AI/ML
AI Agents
Fine-tuning
Prompt Engineering
RAG
Semantic Search
DevOps
AWS
Azure
CI/CD
GCP
Apply
Fullstack Engineer 23 days ago
5+ year exp • Bengaluru
JavaScript
Node JS
Python
SQL
TypeScript
Python
FastAPI
Flask
Pydantic
AI/ML
Copilot
Function Calling
LangChain
LLM
Frontend
Material UI
Next.js
React.js
Tailwind CSS
Vue.js
DevOps
AWS
Azure
CI/CD
Rest API
WebSockets
QA
Jest
Pytest
Apply
Career impact
Discover how this job can transform your career
Get a personal career forecast for this job - salary uplift, next-level role, skill boost and a 3-year financial impact, all calculated from your profile.
Personal salary uplift vs. your current pay
Your 3-year career trajectory
Skills you will level up in this role
3-year financial impact in dollars
Create free account
Free forever • Less than a minute • No credit card

Work setup

Location
Bengaluru