{"id":1285269,"url":"https://alion.io/job/tavant-site-reliability-engineer","title":"Site Reliability Engineer","company":{"id":1938580,"name":"Tavant","domain":"tavant.com","url":"https://alion.io/company/tavant","size_band":"201-500","is_staffing_agency":false,"employer_type":"direct","is_intermediary":false,"listed_via":null,"ats_vendor":null,"truth_index":null},"role":"DevOps","role_family":"DevOps","seniority":"staff","employment_type":null,"work_mode":"hybrid","remote_scope":null,"remote_scope_basis":null,"remote_working_hours":null,"hiring_geo_confidence":"structured","locations":["Bengaluru, India","India"],"countries":["IN"],"hiring_countries":[],"hiring_countries_total":0,"salary":null,"salary_estimate":{"min_usd":25000,"max_usd":59000,"period":"year","method":"global_role_cell_scaled_by_country","sample_n":603},"experience_years_min":15,"visa_sponsorship":false,"relocation_package":false,"has_equity":false,"technologies":[{"name":"AIOps","optional":false},{"name":"Anomaly Detection","optional":false},{"name":"AWS","optional":false},{"name":"Azure","optional":false},{"name":"Datadog","optional":false},{"name":"Dynatrace","optional":false},{"name":"Grafana","optional":false},{"name":"Incident Management","optional":false},{"name":"ITIL","optional":false},{"name":"ITSM","optional":false},{"name":"Kubernetes","optional":false},{"name":"New Relic","optional":false},{"name":"OpenSearch","optional":false},{"name":"PagerDuty","optional":false},{"name":"Platform Engineering","optional":false},{"name":"Prometheus","optional":false},{"name":"ServiceNow","optional":false},{"name":"SLI/SLO/SLA","optional":false},{"name":"Splunk","optional":false}],"status":"live","first_seen_at":"2026-08-17T08:24:02Z","employer_posted_date":null,"last_verified_at":"2026-08-17T08:24:02Z","board_verified":false,"closed_at":null,"days_open":42,"trust":{"level":"not_scored","repost_count":null,"flags":[],"days_open":42},"description":"Job Details : \n\n- Experience : 10 to 15 years\n\n- Location : BLR\n\n- Shift : 2 to 10.30pm\n\n- Notice : Immediate joiner\n\n- Work Mode : Hybrid 3 Days from office\n\nRole Overview : \n\nWe are seeking a highly experienced Enterprise Observability & AIOps Architect with 15+ years of experience in designing and modernizing enterprise-scale observability ecosystems across applications, infrastructure, cloud platforms, databases, integrations, and operational workflows. \n\nThe ideal candidate should possess strong expertise in AIOps & Event Correlation, ITSM Integration, Telemetry Governance, SRE & Operational Excellence, Enterprise Monitoring Rationalization, and AI-driven Operational Transformation. This role requires both strategic architecture leadership and strong hands-on expertise across modern observability and AIOps platforms in large enterprise environments.\n\nKey Responsibilities : \n\n1. Enterprise Observability Architecture : \n\n- Lead enterprise-wide observability assessments across applications, infrastructure, cloud, databases, and operational workflows.\n\n- Define current-state and target-state observability architecture.\n\n- Develop monitoring rationalization and consolidation strategies across enterprise toolsets.\n\n- Establish standards for telemetry, tagging, service identity, alerting, dashboards, and governance.\n\n- Define scalable operating models aligned to SRE, ITSM, and platform engineering practices.\n\n2. Application Observability : \n\n- Architect observability solutions across APM, Distributed tracing, Logs & metrics, and RUM & synthetics.\n\n- Define SLI/SLO-driven monitoring and alerting strategies.\n\n- Improve service dependency visibility, transaction tracing, and telemetry quality.\n\n- Design monitoring patterns for microservices, APIs, Kubernetes, Azure-native, and legacy applications.\n\n3. Infrastructure & Platform Observability : \n\n- Design observability solutions for cloud infrastructure, middleware, databases, platform services, and batch ecosystems.\n\n- Assess alert quality, duplication, routing inefficiencies, and monitoring overlaps.\n\n- Define event correlation, severity models, enrichment standards, and operational ownership structures.\n\n4. AIOps & Intelligent Operations : \n\n- Design AIOps capabilities including event correlation, noise reduction, intelligent alert prioritization, anomaly detection, predictive insights, and root-cause contextualization.\n\n- Define AI-assisted operational workflows for incident reduction, MTTR optimization, and automated remediation.\n\n5. ITSM & Operational Integration : \n\n- Integrate observability platforms with ServiceNow, incident workflows, CMDB, and collaboration tools.\n\n- Define monitoring-to-incident operational workflows and governance standards.\n\n- Establish KPI-driven operational maturity frameworks.\n\n6. Governance & Blueprinting : \n\n- Develop enterprise standards, onboarding blueprints, engineering playbooks, and reusable observability patterns.\n\n- Create reference architectures, dashboard standards, and operational governance frameworks.\n\n- Define Day-1 Observability onboarding models for new services.\n\nRequired Experience : \n\n- 15+ years of experience in observability, infrastructure, SRE, production operations, platform engineering, or AIOps architecture.\n\n- Strong experience in enterprise-scale hybrid cloud and distributed environments.\n\n- Proven experience leading observability transformation and monitoring rationalization initiatives.\n\n- Experience working with executive leadership, enterprise architects, platform teams, and operations organizations.\n\n- Strong understanding of enterprise operational workflows, incident management, and reliability engineering.\n\nRequired Technical Expertise : \n\n1. Observability Platforms : \n\n- Strong hands-on expertise in : Dynatrace, Azure Monitor, Azure Application Insights, Azure Log Analytics, LogicMonitor, ManageEngine.\n\n- Preferred : Splunk, ELK/OpenSearch, Prometheus/Grafana, Datadog, New Relic, BigPanda, PagerDuty.\n\n2. Core Skills : \n\n- Event correlation & alert engineering, Distributed tracing & topology mapping, AIOps & intelligent operations, Cloud monitoring & telemetry, Kubernetes & microservices observability, ITIL / ITSM integration, SRE principles & operational governance.\n\n3. Cloud & Platform Experience : \n\n- Azure, AWS, Kubernetes, APIs & integrations, Middleware & distributed systems.\n\nPreferred Qualifications : \n\n- Experience defining enterprise observability standards and governance models.\n\n- Experience with operational transformation initiatives involving AI/AIOps.\n\n- Strong workshop facilitation, stakeholder management, and executive presentation skills.\n\n- Certifications in Cloud, Observability, ITIL, SRE, or AIOps preferred.\nSkills\nSite Reliability, DevOps, Cloud, Cloud Services, AIOps, DynaTrace, Monitoring Tools, Kubernetes, Production Support, Observability Services","description_format":"text","description_chars":4876,"description_truncated":false,"requirements":{"experience_years_min":15,"management_years_min":null,"team_size_min":null,"manages_managers":false,"education":null,"security_clearance":false,"languages":[]},"benefits":[],"hiring_locations":[],"hiring_excludes":[],"relocation_offered":false,"industries":["Artificial Intelligence","Farming & Agriculture","Real Estate"],"lifecycle":[{"event":"open","at":"2026-09-26T04:00:00Z"}],"liveness":{"score":39,"band":"fade","label":"Fading","p_open":0.6,"p_active":0.86,"p_room":0.75,"age_days":41,"expected_fill_days":53,"reasons":["seen:41","win:late"],"computed_at":"2026-09-28T05:45:00Z"},"pay":null,"html_url":"https://alion.io/job/tavant-site-reliability-engineer","json_url":"https://alion.io/job/tavant-site-reliability-engineer.json","meta":{"generated_at":"2026-09-29T00:46:57Z","cache_seconds":300,"methodology":"https://alion.io/methodology","terms":"https://alion.io/terms","contact":"https://alion.io/contact","api":"https://alion.io/developers","usage":{"tier":"crawler","counted_by":"address","units_charged":1,"used_today":687,"day_limit":5000,"remaining_today":4313,"minute_limit":60,"resets_at":"2026-09-30T00:00:00Z"}}}