{"id":1128835,"url":"https://alion.io/job/servicenow-senior-manager-data-storage-reliability-engineering","title":"Senior Manager, Data & Storage Reliability Engineering","company":{"id":97,"name":"ServiceNow","domain":"servicenow.com","url":"https://alion.io/company/servicenow","size_band":"1001-5000","is_staffing_agency":false,"is_intermediary":false,"ats_vendor":"SmartRecruiters","truth_index":{"grade":"B","score":80,"open_postings":25,"ghost_share":0,"stale_share":1,"repost_share":0,"time_to_fill_p50_days":15,"computed_at":"2026-09-23T05:45:00Z"}},"role":"DevOps","role_family":"DevOps","seniority":"lead","employment_type":"full_time","work_mode":"on_site","remote_scope":null,"hiring_geo_confidence":"structured","locations":["Dublin, Ireland"],"countries":["IE"],"hiring_countries":[],"hiring_countries_total":0,"salary":null,"salary_estimate":{"min_usd":93000,"max_usd":221000,"period":"year","method":"global_role_cell_scaled_by_country","sample_n":475},"experience_years_min":10,"visa_sponsorship":false,"relocation_package":false,"has_equity":false,"technologies":[{"name":"Linux","optional":false},{"name":"Platform Engineering","optional":false},{"name":"ServiceNow","optional":false},{"name":"Anomaly Detection","optional":true},{"name":"MariaDB","optional":true},{"name":"MySQL","optional":true},{"name":"Oracle","optional":true},{"name":"PostgreSQL","optional":true},{"name":"SQL","optional":true}],"status":"live","first_seen_at":"2026-09-22T22:21:33Z","employer_posted_date":"2026-09-22","last_verified_at":"2026-09-23T11:04:03Z","board_verified":true,"closed_at":null,"days_open":0,"trust":{"level":"ok","repost_count":null,"flags":[],"days_open":0},"description":"It all started when engineer Fred Luddy wrote code that automated a tedious task for his coworker, Phyllis. She cried tears of joy. That moment inspired Fred to build a company that could do that for everyone-freeing people from busywork so they could focus on meaningful work. Today, ServiceNow is the AI control tower for business reinvention. Our ServiceNow AI platform brings together any AI, any data, and any workflow- helping 85% of the Fortune 500® work smarter, faster, and better. We're building an AI-native culture where technology and talent are unstoppable together. And we're just getting started.\nJoin us to put AI to work for people.\nServiceNow is seeking an experienced Sr Manager for our Data & Storage Reliability Engineering team.\nThis leader will drive engineering excellence across prevention engineering, reliability engineering, observability, incident learnings, diagnostics, automation, capacity planning, and platform risk reduction. The role requires deep technical expertise in database and distributed systems architectures, large-scale SaaS production environments, and customer-facing operations, combined with strong people leadership and execution rigor.\nThe ideal candidate has experience leading high-performing engineering teams responsible for identifying recurring production patterns, converting incident and escalation learnings into durable engineering improvements, strengthening observability, and building sustainable solutions that improve platform resilience at scale.\nYou should have experience with large-scale web applications, database platforms, distributed systems, Linux-based production environments, and a strong problem-solving mindset for reliability, automation, diagnostics, observability, and prevention. Qualified candidates will be responsible for leading a team of highly skilled engineers that push the limits of scalability, resiliency, and operational excellence.\nDo you\nHave experience leading teams of engineers and developing people? \nEnjoy problem solving and using an analytical mindset to understand why systems fail and how to prevent repeat issues? \nHave a technical background in roles including database engineering, reliability engineering, systems/cloud engineering, SRE, DevOps, or production engineering? \nKnow Linux operating systems, databases, observability, diagnostics, and production troubleshooting well enough to guide engineers through complex investigations? \nHave an attitude of continuous improvement and a passion for removing inefficient, repetitive, or reactive processes through automation and engineering prevention? \nAnswer 'yes' to these questions and we want to hear from you. Hit the Apply button and let's have a chat about the role and your skills and experiences.\nLet’s start with the role\nAs a Sr Manager of the Data & Storage Reliability Engineering team your responsibilities will be:\nDefine and execute the team-level strategy for prevention engineering, reliability, observability, resilience, and operational risk reduction across large-scale production environments. \nLead initiatives that turn production signals, incident learnings, customer escalations, migration outcomes, and platform telemetry into durable engineering improvements. \nPartner closely with SWAT and Customer & Production Engineering to establish a continuous feedback loop between production operations and platform improvement. \nDrive improvements in observability, diagnostics, automation, reliability reviews, resiliency validation, migration readiness, and engineering guardrails. \nIdentify recurring failure patterns, reliability risks, observability gaps, operational inefficiencies, scalability constraints, and performance bottlenecks, and drive action to reduce future customer impact. \nEstablish reliability, resilience, observability, automation, and prevention goals for critical database and storage services. \nChampion proactive monitoring, production analytics, and automation to improve operational health and reduce repetitive manual work. \nLead deep root cause analysis and ensure sustainable corrective actions are implemented for recurring issues and customer-impacting events. \nPartner with engineering leaders to influence database, storage, reliability, observability, and platform architecture priorities based on production evidence. \nBuild and develop a world-class team of reliability, prevention, observability, and platform engineers. \nOwn team management, recruitment, career development, objective setting, project prioritization, onboarding, and performance reviews. \nManage an engineering team that supports production-facing work, including on-call or escalation participation where required. \nDrive a culture of intolerance for repetitive manual activities by promoting automation, self-service diagnostics, guardrails, and scalable engineering practices. \nDrive initiatives with partner teams to improve the reliability, resilience, scalability, and operational efficiency of the ServiceNow application and platform. \nAct as part of the escalation and crisis management ecosystem by helping convert immediate recovery learnings into sustainable engineering prevention. \nAnalyze and evaluate existing processes to drive continuous improvement, operational efficiency, and prevention-oriented engineering practices. \nProvide training, documentation, dashboards, playbooks, and support to partner teams that interface with the Data & Storage Reliability Engineering team. \nOnboard new hires, new technologies, new systems, and new automations into the team to enable successful execution and scale. \n Experience in leveraging or critically thinking about how to integrate AI into work processes, decision-making, or problem-solving - using AI-powered tools, automating workflows, analyzing AI-driven insights, or exploring AI's potential impact on the function.\n\n10+ years of experience in database engineering, reliability engineering, distributed systems, platform engineering, infrastructure engineering, production engineering, or large-scale SaaS platform operations.\n\n4+ years of engineering leadership experience, including leading engineers and cross-functional or distributed teams.\n\nExperience leading Reliability Engineering, Database Engineering, Platform Engineering, Infrastructure Engineering, Production Engineering, Performance Engineering, or related technical teams.\n\nStrong expertise in database technologies, operating system performance, distributed systems, cloud-native architectures, and large-scale production environments.\n\nSolid understanding of reliability engineering, observability, diagnostics, root cause analysis, capacity planning, scalability engineering, resiliency, automation, and operational excellence.\n\nExperience translating production insights, customer escalations, incident learnings, platform telemetry, and recurring operational challenges into prioritized engineering work.\n\nExperience designing and improving observability, diagnostics, reliability reviews, migration readiness checks, resiliency validation, automation, and engineering guardrails.\n\nStrong understanding of tuning and troubleshooting across database, operating system, storage, network, and application layers.\n\nProven experience identifying and resolving complex reliability, scalability, performance, efficiency, and operational bottlenecks in large-scale distributed environments.\n\nExperience leading customer-critical investigations involving reliability, capacity, performance, scalability, resilience, or operational risk challenges.\n\nExperience leveraging observability and telemetry platforms to analyze system behavior and drive platform improvements.\n\nExperience partnering with software engineering, infrastructure, production operations, and escalation organizations to improve platform reliability, scalability, efficiency, and production readiness.\n\nExperience driving engineering initiatives through data, metrics, incident learnings, telemetry, benchmarking, and measurable outcomes.\n\nStrong communication, stakeholder management, and leadership skills.\n\nBachelor's degree in Computer Science, Engineering, or a related technical field, or equivalent practical experience.\n\nDesired Skills\nExperience operating large-scale enterprise database and storage platforms supporting mission-critical workloads.\n\nExperience growing Reliability Engineering, Database Engineering, Platform Engineering, Performance Engineering, Scalability Engineering, or Production Engineering teams.\n\nExperience with observability platforms, telemetry systems, diagnostics frameworks, production analytics, reliability scorecards, engineering metrics, and impact reporting.\n\nExperience with reliability reviews, resiliency validation, migration readiness, workload simulation, capacity forecasting, prevention programs, and operational risk reduction frameworks.\n\nExperience leveraging AI technologies to improve anomaly detection, forecasting, incident analysis, prioritization, operational efficiency, and engineering productivity.\n\nUnderstanding of distributed systems architecture, cloud platform operations, Linux-based production environments, and hyperscale environments.\n\nExperience contributing to platform architecture, database strategy, reliability investments, scalability roadmaps, and long-term engineering improvements.\n\nExperience with performance testing, benchmarking, workload simulation, and capacity modeling as part of broader reliability and prevention engineering programs.\n\nExperience supporting enterprise database technologies such as MySQL, MariaDB, PostgreSQL, Oracle, SQL Server, or cloud-native database platforms.\n\nFamiliarity with ServiceNow platform architecture and large-scale SaaS operations.\n\n Work Personas\nWe approach our distributed world of work with flexibility and trust. Work personas (flexible, remote, or required in office) are categories that are assigned to ServiceNow employees depending on the nature of their work and their assigned work location. Learn more here. To determine eligibility for a work persona, ServiceNow may confirm the distance between your primary residence and the closest ServiceNow office using a third-party service.\nEqual Opportunity Employer\nServiceNow is an equal opportunity employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, sexual orientation, national origin, age, disability, gender identity, veteran status, or any other category protected by law. In addition, all qualified applicants with arrest or conviction records will be considered for employment in accordance with legal requirements.\nAccommodations\nWe strive to create an accessible and inclusive experience for all candidates. If you require a reasonable accommodation to complete any part of the application process, or are unable to use this online application and need an alternative method to apply, please contact  for assistance.\nExport Control Regulations\nFor positions requiring access to controlled technology subject to export control regulations, including the U.S. Export Administration Regulations (EAR), ServiceNow may be required to obtain export control approval from government authorities for certain individuals. All employment is contingent upon ServiceNow obtaining any export license or other approval that may be required by relevant export control authorities.\nFrom Fortune. ©2026 Fortune Media IP Limited. All rights reserved. Used under license.","description_format":"text","description_chars":11543,"description_truncated":false,"requirements":{"experience_years_min":10,"management_years_min":null,"team_size_min":null,"manages_managers":false,"education":{"level":"bachelor","optional":false},"security_clearance":false,"languages":[{"language":"English","level":"All levels","optional":false}]},"benefits":[],"hiring_locations":[{"name":"United States","iso":"US","kind":"country"}],"hiring_excludes":[],"relocation_offered":false,"industries":["Business Process Automation (BPA)","IT Management","AI Agents"],"lifecycle":[{"event":"open","at":"2026-09-22T23:44:52Z"}],"liveness":{"score":63,"band":"ok","label":"Likely open","p_open":1,"p_active":0.632,"p_room":1,"age_days":0,"expected_fill_days":15,"reasons":["conf:6","stale_co","velocity","win:early","comp:brand"],"computed_at":"2026-09-23T05:45:00Z"},"pay":null,"html_url":"https://alion.io/job/servicenow-senior-manager-data-storage-reliability-engineering","json_url":"https://alion.io/job/servicenow-senior-manager-data-storage-reliability-engineering.json","meta":{"generated_at":"2026-09-23T16:14:35Z","cache_seconds":300,"methodology":"https://alion.io/methodology","terms":"https://alion.io/terms","contact":"https://alion.io/contact","api":"https://alion.io/developers"}}