1,433,388open jobs
83,596companies
217,057added this week
Browse all
Salary
$28k – $36k per year (gross)
Location
In office (Hyderabad)
Seniority
Staff · 3+ years exp

First seen by Alion on Oct 5, 2026.

Overview
Company
Impact
Profile match
Job Searching Just Got So Easy. RELIABLE WAY TO FIND YOUR DREAM JOBS.

Position Overview :

We are seeking an experienced Observability Architect to join our platform engineering and reliability team focused on developing intelligent systems for AIOps and Data Science. In addition to building machine learning models for log and telemetry data, this role will contribute to product strategy, collaborate closely with engineering teams, and help shape the architecture of our observability and operational intelligence platform.

Key Responsibilities :

- Design and build machine learning models for error detection, anomaly detection, and root cause analysis from system logs and metrics.

- Develop log correlation algorithms identifying relationships between disparate log entries across distributed systems.

- Build and maintain knowledge graphs representing system dependencies, service relationships, and failure patterns.

- Create automated incident analysis pipelines ingesting raw logs, correlating events, and suggesting root causes in real time.

- Collaborate with platform engineers to design scalable, resilient telemetry ingestion, aggregation, and analytics pipelines supporting real-time operational intelligence.

- Contribute to defining platform standards, technology selections, and engineering practices aligned with the long-term product vision.

- Partner with SRE, engineering, and product teams to translate enterprise observability challenges into scalable, AI-powered platform solutions.

- Participate in design reviews and help optimize platform reliability, performance, security, automation, and scalability.

- Mentor and guide junior data scientists and cross-functional teams in applied ML and observability techniques.

- Document model methodologies and assumptions and create dashboards for model performance and stakeholder visibility.

- Stay abreast of emerging observability technologies, AI workflows, and telemetry-driven operational intelligence trends.

- Support the development of backend services or APIs in languages such as Python, Java, or .NET to integrate ML components with platform services.

Required Skills & Qualifications :

- Bachelor's degree in Data Science, Computer Science, Statistics, or a related field, or equivalent experience.

- 3+ years of professional experience in data science, machine learning engineering, or observability analytics.

- Strong programming skills in Python (preferred), with optional experience in Java, Go, or .NET for production code integration.

- Experience developing ML models using frameworks such as TensorFlow, PyTorch, or scikit-learn.

- Expertise in NLP techniques for log analysis and time-series anomaly detection.

- Familiarity with distributed computing environments such as Spark, Flink, or Kafka.

- Knowledge of knowledge graph technologies and graph neural networks, such as Neo4j, PyG, or DGL.

- Experience with telemetry and observability platforms such as ELK, Datadog, Splunk, Prometheus, or similar tools.

- Understanding of observability concepts, including metrics, logs, traces, dashboards, alerting, SLOs/SLIs, incident management, and operational analytics.

- Exposure to AI-powered operational intelligence, including agentic AI workflows, LLM-based assistants, graph-based dependency mapping, automated incident detection and triage, root cause analysis, and remediation recommendations.

- Strong problem-solving skills and the ability to communicate complex models to technical and non-technical stakeholders.

Preferred Qualifications :

- Master's degree in Machine Learning, Computer Science, or a related discipline.

- Experience in AIOps, Site Reliability Engineering, or IT Operations domains.

- Published research or contributions to open-source ML or observability projects.

- Knowledge of causal inference techniques for root cause analysis.

- Familiarity with containerization technologies, including Docker and Kubernetes, and CI/CD pipelines.

- Experience with incident management systems and on-call tooling.

- Background working with microservices architectures or cloud platforms; Azure experience is preferred.

- Awareness of SRE practices, self-healing automation, capacity prediction, and operational decision intelligence.

Skills

AIOps, Machine Learning, Python, Observability Services, MLOps, Data Science, AI Observability, Site Reliability, NLP, PyTorch

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
1,433,388 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

DevOps
Similar stack
Same company
Hyderabad
$60k per year • In office • Full-Time • Graz
DevOps
VMWare
Windows Server
Proxmox VE
Linux
Windows
Management
ITIL
Apply
$42k per year • In office • Full-Time • Amstetten
DevOps
VMWare
Azure
Windows Server
Docker
Kubernetes
Ubuntu
Hyper-V
Linux
Windows
DNS
Cybersecurity
Active Directory
PKI
Apply
$50k per year • In office • Full-Time • Salzburg
DevOps
VPN
Apply
$44k per year • In office • Full-Time • Wels
DevOps
Linux
Windows
VPN
Cybersecurity
Active Directory
Apply
$64k per year • In office • Full-Time • Vienna
Python
Bash
DevOps
Rest API
Splunk
Ansible
Red Hat
OpenShift
Loki
Git
Kubernetes
Ubuntu
Grafana
Configuration Management
Bitbucket
Linux
Apply
≈ $19k – $54k per year (Estimated) • Remote (likely EAEU) • 2+ years exp • Tbilisi
Java
Java
Spring Boot
DevOps
Rest API
GCP
Azure
HAProxy
AWS
Docker
Kubernetes
Cloudflare
Nginx
kubectl
Fastly
Akamai
Incident Management
SLI/SLO/SLA
Amazon S3
Linux
Cybersecurity
Snyk
SonarQube
Trivy
Management
Confluence
Jira
QA
Chrome DevTools
Postman
Apply
≈ $6k – $17k per year (Estimated) • Remote (likely EAEU) • 2+ years exp • Minsk
Java
Java
Spring Boot
DevOps
Rest API
GCP
Azure
HAProxy
AWS
Docker
Kubernetes
Cloudflare
Nginx
kubectl
Fastly
Akamai
Incident Management
SLI/SLO/SLA
Amazon S3
Linux
Cybersecurity
Snyk
SonarQube
Trivy
Management
Confluence
Jira
QA
Chrome DevTools
Postman
Apply
≈ $26k – $49k per year (Estimated) • Hybrid • Internship • Hong Kong
Python
Apply
≈ $60k – $147k per year (Estimated) • Hybrid • Full-Time • PhD • Paris
Python
C++
AI/ML
Machine Learning
DevOps
CI/CD
Robotics
SLAM
Localization
Path Planning
Apply
Hybrid • Full-Time • Bachelor's Degree • Bengaluru
Python
SQL
Databases
PostgreSQL
Oracle
AI/ML
AI Agents
Edge AI
DevOps
Ansible
Chef
Azure
AWS
Configuration Management
Incident Management
Management
Agile
ITIL
Apply
Senior Data Engineer 17 days ago
$36k – $42k per year (gross) • In office • 6+ years exp • Bengaluru
Python
SQL
Scala
Python
pySpark
Databases
Databricks
Delta Lake
AI/ML
Hadoop
Spark
DevOps
Azure
Apply
$31k – $36k per year (gross) • In office • 8+ years exp • Bengaluru
Management
Agile
Scrum
Waterfall
Apply
≈ $28k – $49k per year (Estimated) • In office • 8+ years exp • Noida • Gurgaon
Apex
Apex
Visualforce
AI/ML
AI Agents
Agentforce
Apply
Salesforce Developer 2 months ago
≈ $15k – $34k per year (Estimated) • Hybrid • 7+ years exp • Noida • Gurgaon
JavaScript
Apex
Apex
Lightning Web Components
Visualforce
DevOps
Rest API
Incident Management
SOAP
Analytics
Power BI
Apply
≈ $17k – $42k per year (Estimated) • In office • Full-Time • 10+ years exp • Bachelor's Degree • Hyderabad
Java
DevOps
Ansible
Dynatrace
CI/CD
Git
Platform Engineering
Configuration Management
Incident Management
Management
ServiceNow
Agile
Scrum
Apply
≈ $12k – $28k per year (Estimated) • Remote (India) • Full-Time • 3+ years exp • Hyderabad
Apply
≈ $10k – $20k per year (Estimated) • Remote (India) • Full-Time • 1+ year exp • Bachelor's Degree • Hyderabad
DevOps
SLI/SLO/SLA
Analytics
Microsoft Excel
Apply
≈ $11k – $22k per year (Estimated) • Remote (India) • Full-Time • Hyderabad
Management
Google Sheets
Apply
Remote (India) • Full-Time • 6+ years exp • Bachelor's Degree • Hyderabad
DevOps
SLI/SLO/SLA
Analytics
Microsoft Excel
Apply
See all jobs
This is one of many
1,433,388 more open roles from verified company boards, updated every day.