{"id":1584058,"url":"https://alion.io/job/mastercard-lead-site-reliability-engineering","title":"Lead, Site Reliability Engineering","company":{"id":252,"name":"Mastercard","domain":"mastercard.com","url":"https://alion.io/company/mastercard","size_band":"5000+","is_staffing_agency":false,"employer_type":"direct","is_intermediary":false,"listed_via":null,"ats_vendor":"Workday","truth_index":{"grade":"A","score":91,"open_postings":152,"ghost_share":0.013,"stale_share":0.48,"repost_share":0.053,"time_to_fill_p50_days":21,"computed_at":"2026-10-01T05:45:00Z"}},"role":"DevOps","role_family":"DevOps","seniority":"lead","employment_type":"full_time","work_mode":"on_site","remote_scope":null,"remote_scope_basis":null,"remote_working_hours":null,"hiring_geo_confidence":"structured","locations":["Dublin, Ireland"],"countries":["IE"],"hiring_countries":[],"hiring_countries_total":0,"salary":null,"salary_estimate":{"min_usd":87000,"max_usd":213000,"period":"year","method":"global_role_cell_scaled_by_country","sample_n":602},"experience_years_min":null,"visa_sponsorship":false,"relocation_package":false,"has_equity":false,"technologies":[{"name":"Linux","optional":false},{"name":"Machine Learning","optional":false},{"name":"Self-Healing","optional":false}],"status":"live","first_seen_at":"2026-10-01T13:50:24Z","employer_posted_date":"2026-10-01","last_verified_at":"2026-10-01T13:50:24Z","board_verified":true,"closed_at":null,"days_open":0,"trust":{"level":"ok","repost_count":null,"flags":[],"days_open":0},"description":"Our Purpose\nMastercard powers economies and empowers people in 200+ countries and territories worldwide. Together with our customers, we’re helping build a sustainable economy where everyone can prosper. We support a wide range of digital payments choices, making transactions secure, simple, smart and accessible. Our technology and innovation, partnerships and networks combine to deliver a unique set of products and services that help people, businesses and governments realize their greatest potential.\nTitle and Summary\nLead, Site Reliability EngineeringSite Reliability Engineer (SRE) - GeneralistRole Summary\nThe Site Reliability Engineer (SRE) - Generalist is a senior level engineer and cross stack reliability expert who proactively ensures system stability, performance, and operational resilience by deeply understanding application behavior and how it manifests across infrastructure.\nThis role emphasizes anticipation over reaction. While the SRE Generalist participates in incident response, their primary value is in converting operational signals, incidents, and patterns into preventative actions-improving observability, reducing risk, and eliminating classes of failure before they impact customers. They partner closely with application, platform, and infrastructure teams to continuously reduce mean time to detect (MTTD), mean time to resolve (MTTR), and overall incident frequency through data driven insight, automation, and engineering rigor.\nKey Responsibilities\nProactive Reliability Engineering\nAnticipate reliability risks by analyzing application behavior, system signals, and historical incidents to identify failure patterns and systemic weaknesses before they result in outages.\nTranslate deep application knowledge into reliability requirements, architectural guidance, and infrastructure improvements that prevent incidents rather than simply respond to them.\nContinuously assess system health, resiliency gaps, and operational debt, driving improvements that increase service robustness over time.\nIncident Response as an Input to Prevention\nParticipate in and lead troubleshooting efforts during high severity and cross domain incidents, applying structured, data driven investigation techniques.\nUse incidents as learning opportunities-performing root cause analysis that focuses on why systems allowed failure, not just what broke.\nEnsure incident outcomes result in concrete, measurable improvements such as better instrumentation, safer defaults, automation, or architectural changes.\nObservability, Monitoring & Signal Quality\nProactively design and evolve observability strategies by onboarding new data sources and improving signal quality across logs, metrics, traces, and events.\nBuild dashboards, alerts, and monitors that surface early indicators of degradation, not just failure states.\nApply analytical techniques to detect emerging trends, weak signals, and anomalous behavior before customers are impacted.\nCommunicate insights through clear data storytelling that enables engineering teams and leaders to act decisively and early.\nAutomation & Continuous Improvement\nLead automation efforts that reduce manual intervention, shorten feedback loops, and eliminate repetitive operational work.\nConvert operational learnings into reusable tools, standards, documentation, and patterns that raise the reliability baseline across teams.\nActively reduce operational toil and risk by improving system defaults, guardrails, and self healing capabilities.\nCollaboration, Influence & Mentorship\nPartner across application, infrastructure, and platform teams to drive shared ownership of reliability outcomes and proactive operational thinking.\nInfluence design and delivery decisions by representing the reliability perspective early in the development lifecycle.\nMentor engineers by modeling proactive troubleshooting, systems thinking, and data driven decision making.\nKnowledge, Skills & Abilities\nStrong ability to reason about systems end to end, connecting application behavior to infrastructure performance and failure modes.\nExpertise in observability, monitoring, and troubleshooting tools, with a focus on signal quality and actionable insight.\nProficiency in scripting and automation to operationalize reliability improvements and accelerate learning.\nBroad infrastructure knowledge (networking, Linux, databases, containers, storage), with depth in at least one domain.\nStrong data analysis and storytelling skills, enabling proactive identification of risks and clear communication of technical insights.\nWorking knowledge of machine learning concepts and their application to predictive and proactive operational problem solving.\nCuriosity, ownership, and a mindset oriented toward preventing tomorrow’s incidents, not just fixing today’s.\nWhat Defines Success in This Role\nA successful SRE Generalist:\nSees incidents as signals, not endpoints.\nUses observability and data to shift reliability work left and upstream.\nReduces incident frequency and impact over time-not just MTTR.\n• Acts as a connective force across teams, turning complexity into clarity and prevention.\nCorporate Security Responsibility\nAll activities involving access to Mastercard assets, information, and networks comes with an inherent risk to the organization and, therefore, it is expected that every person working for, or on behalf of, Mastercard is responsible for information security and must:\nAbide by Mastercard’s security policies and practices;\n\nEnsure the confidentiality and integrity of the information being accessed;\n\nReport any suspected information security violation or breach, and\n\nComplete all periodic mandatory security trainings in accordance with Mastercard’s guidelines.","description_format":"text","description_chars":5727,"description_truncated":false,"requirements":{"experience_years_min":null,"management_years_min":null,"team_size_min":null,"manages_managers":false,"education":null,"security_clearance":false,"languages":[]},"benefits":[],"hiring_locations":[],"hiring_excludes":[],"relocation_offered":false,"industries":["Cards & Card Issuing","Payment Processing & Gateways"],"lifecycle":[{"event":"open","at":"2026-10-01T13:50:24Z"}],"liveness":{"score":86,"band":"hot","label":"Hiring now","p_open":1,"p_active":0.86,"p_room":1,"age_days":0,"expected_fill_days":21,"reasons":["conf:12","win:early","comp:brand"],"computed_at":"2026-10-02T02:36:48Z"},"pay":null,"html_url":"https://alion.io/job/mastercard-lead-site-reliability-engineering","json_url":"https://alion.io/job/mastercard-lead-site-reliability-engineering.json","meta":{"generated_at":"2026-10-02T02:36:48Z","cache_seconds":300,"methodology":"https://alion.io/methodology","terms":"https://alion.io/terms","contact":"https://alion.io/contact","api":"https://alion.io/developers","usage":{"tier":"crawler","counted_by":"address","units_charged":1,"used_today":3531,"day_limit":5000,"remaining_today":1469,"minute_limit":60,"resets_at":"2026-10-03T00:00:00Z"}}}