{"id":1216872,"url":"https://alion.io/job/bce-global-tech-lead-aiops-support-engineer","title":"Lead AIOps Support Engineer","company":{"id":6489,"name":"BCE Global Tech","domain":"bceglobaltech.com","url":"https://alion.io/company/bce-global-tech","size_band":null,"is_staffing_agency":false,"employer_type":"direct","is_intermediary":false,"listed_via":null,"ats_vendor":null,"truth_index":null},"role":"Support","role_family":"Support","seniority":"lead","employment_type":null,"work_mode":"on_site","remote_scope":null,"remote_scope_basis":null,"remote_working_hours":null,"hiring_geo_confidence":"structured","locations":["Bengaluru, India"],"countries":["IN"],"hiring_countries":[],"hiring_countries_total":0,"salary":null,"salary_estimate":null,"experience_years_min":7,"visa_sponsorship":false,"relocation_package":false,"has_equity":false,"technologies":[{"name":"AI Agents","optional":false},{"name":"AIOps","optional":false},{"name":"Amazon CloudWatch","optional":false},{"name":"AWS","optional":false},{"name":"Azure","optional":false},{"name":"DNS","optional":false},{"name":"Dynatrace","optional":false},{"name":"GCP","optional":false},{"name":"Incident Management","optional":false},{"name":"ITIL","optional":false},{"name":"Jira","optional":false},{"name":"Kubernetes","optional":false},{"name":"New Relic","optional":false},{"name":"OpenShift","optional":false},{"name":"OpenTelemetry","optional":false},{"name":"Postman","optional":false},{"name":"Python","optional":false},{"name":"ServiceNow","optional":false},{"name":"SLI/SLO/SLA","optional":false},{"name":"SOAP","optional":false},{"name":"SQL","optional":false},{"name":"TCP/IP","optional":false},{"name":"Windows","optional":false},{"name":"Knowledge Graph","optional":true},{"name":"Prompt Engineering","optional":true},{"name":"RAG","optional":true}],"status":"live","first_seen_at":"2026-09-25T10:00:01Z","employer_posted_date":null,"last_verified_at":"2026-09-25T10:00:01Z","board_verified":false,"closed_at":null,"days_open":2,"trust":{"level":"not_scored","repost_count":null,"flags":[],"days_open":2},"description":"We are starting with Tier 1 (L1) triage support for 50 applications and scaling to roughly 320 applications within two years. The AIOps Support Lead will manage a team of 14 AIOps Support Engineers, own the quality of day-to-day manual triage, and drive the transition toward automated, proactive triage as the OpenTelemetry backbone matures. This role owns the people, process, and data-quality discipline that makes reliable triage possible; it does not own the agentic automation layer itself but prepares the team and the data for it.\nResponsibilities:\nManage the Tier 1 team: directly manage a team of 14 AIOps Support Engineers performing manual triage of alarms and alerts across a diverse, 50-to-320-application portfolio, including hiring, coaching, scheduling, and performance management.\nOwn triage quality and speed: set and monitor standards for how quickly and accurately the team detects, classifies, and routes incidents, and drive continuous improvement in mean time-to-triage.\nDrive data stewardship: partner with application teams to standardize alarm and alert data across heterogeneous log aggregation tools (Dynatrace, New Relic, ManageEngine, Glassbox, and others) into a clean, consistent telemetry backbone built on OpenTelemetry.\nManage the reactive-to-proactive shift: reduce reliance on reactive, manual triage over time by improving alert quality, correlation, and early-warning signals, laying the groundwork for future automated and agentic triage.\nNavigate a diverse, moving application landscape: support applications spanning different technology stacks and different architecture dispositions (Invest, Tolerate, Retire, Migrate), reprioritizing team focus as the portfolio shifts.\nCoordinate onboarding of new apps: run a repeatable process for bringing new applications into Tier 1 coverage as the program scales from 50 to 320 applications, including support group and application owner mapping.\nManage stakeholders: act as the primary point of contact for support groups, application owners, and AIOps program leadership on Tier 1 status, incidents, and data-quality issues.\nManage shift/roster coverage: ensure the team of 14 provides consistent triage coverage across required hours as the application count grows. Report on outcomes: track and report team KPIs (triage time, alert-to-incident accuracy, false-positive rates, and coverage growth to program leadership).\nRequirements:\nExperience: 7+ years in application/production support (L1/L1.5/L2) or site reliability, with 2+ years directly managing a technical support team.\nHybrid environment expertise: proven experience supporting applications across both on-premises and cloud environments, with exposure to modern microservices architectures.\nObservability tooling: hands-on experience with monitoring and observability platforms such as Dynatrace, New Relic, AWS CloudWatch, ManageEngine, or Glassbox; working knowledge of OpenTelemetry and distributed tracing concepts.\nITIL discipline: strong grounding in incident, problem, and change management practices, with ServiceNow or Jira ticket management experience.\nTechnical range: comfortable with Linux and Windows troubleshooting, basic networking (TCP/IP, DNS, HTTP/HTTPS, SSL, load balancers), SQL/database query analysis, and API/integration troubleshooting.\nPeople management: demonstrated ability to hire, coach, and retain a team of 10+ technical support staff through a period of significant scale-up (5x application coverage growth).\nAnalytical mindset: able to turn noisy, inconsistent alert data into clear, actionable insight and to build repeatable frameworks rather than one-off fixes.\nComfort with ambiguity: willing to support a moving target, a portfolio spanning Invest, Tolerate, Retire, and Migrate applications, and to adapt priorities as the program evolves.\nMust Have:\nIncident Management Lifecycle and working knowledge of Problem, Change Request, and Service Request concepts (ITIL).\nCMDB concepts and their use in incident and asset traceability.\nHands-on experience with log aggregation technologies (Dynatrace, New Relic, ManageEngine, Glassbox, or similar).\nWorking knowledge of JSON and XML and basic file/task automation.\nUnderstanding of IT infrastructure and basic networking: VMs, firewalls, load balancers, containers, OpenShift (OCP), and Kubernetes.\nUnix Shell scripting; Windows batch file creation.\nBasic cloud concepts: compute, storage, and security fundamentals.\nSecurity fundamentals: TLS, SSL, tokens, and secret management.\nFamiliarity with API gateways and API testing toolkits (Postman, SOAP UI, or similar).\nOutage response management and experience leading cross-functional coordination during major incidents.\nAbility to drive Root Cause Analyses (RCAs) and build reusable knowledge articles/runbooks.\nSLA/SLO management and reporting, including availability calculation.\nWorking knowledge of data concepts: data latency, data fragmentation, data lineage, and data marts.\nNice-to-Have Skills:\nFamiliarity with AI concepts such as prompt engineering, knowledge graphs, and Retrieval-Augmented Generation (RAG).\nExperience with cloud-native observability on AWS, Azure, or GCP.\nExposure to Agentic AI or automation-driven triage tooling\nSuccess Looks Like:\nA 14-person Tier 1 team that reliably triages alerts across 50+ applications with clear, standardized data.\nA measurable, ongoing reduction in reactive manual triage as proactive detection improves\nA clean, well-governed OpenTelemetry-based data backbone that the future agentic automation layer can build on.\nA repeatable onboarding process ready to scale coverage from 50 to 320 applications.","description_format":"text","description_chars":5652,"description_truncated":false,"requirements":{"experience_years_min":7,"management_years_min":null,"team_size_min":10,"manages_managers":false,"education":null,"security_clearance":false,"languages":[]},"benefits":[],"hiring_locations":[],"hiring_excludes":[],"relocation_offered":false,"industries":["Mobile Networks","Custom Software Development"],"lifecycle":[{"event":"open","at":"2026-09-25T10:00:01Z"}],"liveness":{"score":90,"band":"hot","label":"Hiring now","p_open":1,"p_active":0.903,"p_room":1,"age_days":1,"expected_fill_days":24,"reasons":["seen:1","velocity","win:early","comp:attention"],"computed_at":"2026-09-27T05:45:00Z"},"pay":null,"html_url":"https://alion.io/job/bce-global-tech-lead-aiops-support-engineer","json_url":"https://alion.io/job/bce-global-tech-lead-aiops-support-engineer.json","meta":{"generated_at":"2026-09-28T01:14:15Z","cache_seconds":300,"methodology":"https://alion.io/methodology","terms":"https://alion.io/terms","contact":"https://alion.io/contact","api":"https://alion.io/developers","usage":{"tier":"crawler","counted_by":"address","units_charged":1,"used_today":814,"day_limit":5000,"remaining_today":4186,"minute_limit":60,"resets_at":"2026-09-29T00:00:00Z"}}}