Mastech Digital
We are looking for an experienced AIOps Engineer who is passionate about building intelligent, resilient, and highly automated cloud platforms. This role sits at the intersection of Platform Engineering, DevOps, Site Reliability Engineering (SRE), and AI/ML Operations, where you'll build AI-driven observability solutions, intelligent automation pipelines, and self-healing infrastructure for enterprise-scale applications. You'll work closely with Platform, AI, and Product Engineering teams to embed machine learning into operational workflows, automate incident response, improve platform reliability, and enable proactive infrastructure management using modern AI technologies.
The candidate will have responsibilities across the following functions:
AIOps and Intelligent Automation:
- Design and implement AI-powered observability platforms for anomaly detection, predictive alerting, event correlation, and root-cause analysis.
- Build Python-based ML pipelines for log analytics, telemetry processing, and operational intelligence.
- Develop intelligent auto-remediation workflows to reduce manual intervention and improve platform availability.
- Integrate LLM-powered assistants into incident management workflows for automated triage, RCA generation, and operational insights.
- Implement predictive capacity planning using machine learning models.
Platform Engineering and DevOps:
- Design, build, and maintain CI/CD pipelines using GitHub Actions, GitLab CI, or Jenkins.
- Provision and manage Kubernetes environments (EKS/AKS/GKE) using GitOps methodologies.
- Build Infrastructure-as-Code using Terraform and automate cloud provisioning.
- Develop secure container platforms using Docker, Helm, and Kubernetes best practices.
- Manage production deployments across AWS and Azure cloud environments.
Site Reliability Engineering:
- Define and monitor SLIs, SLOs, and error budgets.
- Build enterprise observability platforms using Prometheus, Grafana, OpenTelemetry, Loki/OpenSearch, and Jaeger/Tempo.
- Lead incident management, root cause analysis, and post-incident improvements.
- Drive platform reliability, scalability, and operational excellence.
Python Development:
- Develop production-grade automation tools, REST APIs, internal utilities, and operational dashboards using Python.
- Build integrations with monitoring, alerting, Slack, PagerDuty, and CMDB platforms.
- Develop data pipelines supporting telemetry ingestion and AI model retraining workflows.
Requirements:
- 5-7 years of experience in Platform Engineering, DevOps, or Site Reliability Engineering.
- Strong expertise in Python, Kubernetes, Terraform, and Cloud Infrastructure.
- Experience building production-grade AI-enabled operational platforms.
- Passion for automation, reliability, and cloud-native engineering.
- Excellent debugging, problem-solving, and communication skills.
- Ability to collaborate across Platform, AI, Security, and Product Engineering teams.
- Platform and DevOps: Kubernetes (EKS, AKS, or GKE), Docker, Helm, GitOps (ArgoCD preferred), CI/CD (GitHub Actions, GitLab CI, Jenkins), Terraform, Infrastructure Automation.
- Programming: Python, REST API Development (FastAPI or Flask), Automation & Scripting, Unit Testing (pytest).
- Observability: Prometheus, Grafana, Alertmanager, OpenTelemetry, Distributed Tracing, Log Aggregation Platforms.
- Cloud: AWS (EC2 IAM, EKS, CloudWatch, VPC), Azure (preferred), Multi-cloud Infrastructure.
- Site Reliability: Incident Management, Capacity Planning, Performance Monitoring, Reliability Engineering, Production Operations.
AI / MLOps Skills:
- Experience deploying AI/ML models into production environments.
- Hands-on experience with OpenAI, Azure OpenAI, or other LLM platforms.
- Knowledge of Retrieval-Augmented Generation (RAG) architectures.
- Experience with vector databases such as Pinecone, pgvector, or MongoDB Atlas Vector Search.
- Prompt engineering and LLM optimisation.
- Experience with model serving platforms such as vLLM, Ollama, or Text Generation Inference (TGI).
- Familiarity with LangChain, LangGraph, or AI orchestration frameworks.
- Understanding of model monitoring, prompt evaluation, and LLM observability.
- Good to Have: MLOps experience, Kafka or NATS, Istio or Linkerd, KEDA, Security tooling (Trivy, Falco, OPA), FinOps tooling, MCP (Model Context Protocol), AWS, Azure, or GCP Certifications, Contributions to open-source DevOps or AI projects.
