{"id":1492667,"url":"https://alion.io/job/salesforce-senior-site-reliability-engineer-2","title":"Senior Site Reliability Engineer","company":{"id":70,"name":"Salesforce","domain":"salesforce.com","url":"https://alion.io/company/salesforce","size_band":"5000+","is_staffing_agency":false,"employer_type":"direct","is_intermediary":false,"listed_via":null,"ats_vendor":"Workday","truth_index":{"grade":"A","score":89,"open_postings":313,"ghost_share":0.006,"stale_share":0.61,"repost_share":0.032,"time_to_fill_p50_days":21,"computed_at":"2026-10-01T05:45:00Z"}},"role":"DevOps","role_family":"DevOps","seniority":"senior","employment_type":"full_time","work_mode":"on_site","remote_scope":null,"remote_scope_basis":null,"remote_working_hours":null,"hiring_geo_confidence":"structured","locations":["Dublin, Ireland"],"countries":["IE"],"hiring_countries":[],"hiring_countries_total":0,"salary":null,"salary_estimate":{"min_usd":84000,"max_usd":168000,"period":"year","method":"role_seniority_country_remote_cell","sample_n":10},"experience_years_min":5,"visa_sponsorship":false,"relocation_package":false,"has_equity":false,"technologies":[{"name":"Agentforce","optional":false},{"name":"AI Agents","optional":false},{"name":"Anomaly Detection","optional":false},{"name":"Argo Workflows","optional":false},{"name":"AWS","optional":false},{"name":"Chaos Engineering","optional":false},{"name":"CI/CD","optional":false},{"name":"Claude Code","optional":false},{"name":"Commander.js","optional":false},{"name":"Copilot","optional":false},{"name":"Cursor","optional":false},{"name":"Datadog","optional":false},{"name":"DNS","optional":false},{"name":"Docker","optional":false},{"name":"GCP","optional":false},{"name":"Go","optional":false},{"name":"Grafana","optional":false},{"name":"Incident Management","optional":false},{"name":"Kubernetes","optional":false},{"name":"Linux","optional":false},{"name":"LLM","optional":false},{"name":"Model Context Protocol","optional":false},{"name":"OpenAI Codex","optional":false},{"name":"Prometheus","optional":false},{"name":"Prompt Engineering","optional":false},{"name":"Python","optional":false},{"name":"Self-Healing","optional":false},{"name":"Splunk","optional":false},{"name":"Unix","optional":false},{"name":"JavaScript","optional":true},{"name":"Node JS","optional":true}],"status":"live","first_seen_at":"2026-09-29T00:00:00Z","employer_posted_date":"2026-09-29","last_verified_at":"2026-10-01T06:28:26Z","board_verified":true,"closed_at":null,"days_open":2,"trust":{"level":"ok","repost_count":null,"flags":[],"days_open":2},"description":"To get the best candidate experience, please consider applying for a maximum of 3 roles within 12 months to ensure you are not duplicating efforts.\nJob Category\nSoftware EngineeringJob Details\nAbout Salesforce\nSalesforce is the #1 AI CRM, where humans with agents drive customer success together. Here, ambition meets action. Tech meets trust. And innovation isn’t a buzzword - it’s a way of life. The world of work as we know it is changing and we're looking for Trailblazers who are passionate about bettering business and the world through AI, driving innovation, and keeping Salesforce's core values at the heart of it all.\nReady to level-up your career at the company leading workforce transformation in the agentic era? You’re in the right place! Agentforce is the future of AI, and you are the future of Salesforce.\nSalesforce is seeking a senior engineering candidate to join the Site Reliability organization in Dublin. Working closely with counterparts in the Infrastructure and R&D organizations, this organization provides a global team of engineers monitoring cloud service availability and ready to swiftly repair any service-impacting issues. Five days a week, 24 hours a day, in a follow-the-sun model with weekend oncall, the Site Reliability team keeps the Salesforce cloud and our customers protected.\nThe Experience\nAs an SRE, you will be a technical leader of the team driving Salesforce’s operational resilience by engineering solutions that blend automation, observability, and AI-powered platforms. You will not only respond to incidents but proactively design systems that prevent them, applying software engineering principles to operations to reduce toil and improve reliability at scale. By leveraging cutting-edge software engineering practices within SRE function and AI-driven insights, you will help transform how services are built, monitored, and operated - ensuring that Salesforce delivers always-on, high-performance experiences to customers worldwide.\nBuild and run reliable, scalable, and efficient systems by applying software engineering principles to operations. Our mission is to ensure services are highly available, performant, and resilient - while continuously improving the balance between operational work and engineering innovation.\nReliability as the Priority: Ensure that systems meet defined Service Level Indicators (SLIs) and Service Level Objectives (SLOs), using error budgets to guide engineering and release decisions.\nEngineering for Operations: Apply software engineering practices - automation, monitoring, self-healing systems - to eliminate toil and improve operational efficiency.\nIncident Management: Lead the coordinated response to incidents as an Incident Commander, drive fast recovery (low TTR), and ensure lasting improvements through blameless postmortems.\nContinuous Improvement: Identify and remove sources of toil, enhance observability, and optimize systems to reduce Time to Detect (TTD) and Time to Restore (TTR).\nCollaboration with Development: Partner with product and engineering teams early in the lifecycle to design, build, and operate systems that are reliable by default.\nLong-Term Focus: Leverage AI-driven automation to eliminate manual workflows, enabling the team to focus on complex problem-solving and strategic innovation while reducing operational overhead to less than 20% of capacity.\nWhat You'll Actually Be Doing:\nLead incident detection, response, and resolution-driving root cause analysis, postmortems, and proactive measures to ensure high uptime, rapid recovery, and prevention of future issues.\nLead post-incident reviews, drive systemic fixes through corrective actions, and ensure customer-facing services maintain peak performance and reliability.\nUnderstanding of AI/ML concepts applied to operations (e.g., anomaly detection, predictive analysis).\nIndependently drive the design and implementation of complex automation platforms, self-healing systems, and AI-powered operational tooling using durable workflow engines (Temporal, Airflow, Argo Workflows).\nArchitect and build production-grade observability solutions - monitoring, logging, alerting, and tracing systems - that enable proactive detection and autonomous remediation.\nDesign and implement AI/ML-powered operations tools including anomaly detection systems, predictive analysis pipelines, intelligent runbook automation, and prompt-engineered operational agents (MCP-based).\nDrive optimization of system performance, reliability, and cost-effectiveness through proactive monitoring and tuning.\nEnsuring that work carried out by the Site Reliability team is executed in such a way as to comply with the company’s internal compliance policy and directives.\nIdentifying opportunities and driving the creation of comprehensive technical epics that include well-defined problem statements, detailed project and implementation documentation, and clearly measurable business outcomes aligned with team objectives.\nProvide technical coaching to junior team members through pair programming, design reviews, and code reviews - helping grow their skills and knowledge.\nCollaborate with engineering and product teams to define and uphold SLAs/SLOs, driving improvements in service reliability and customer experience.\nBuild and ship high-quality, production-grade software using modern engineering practices, with AI as a core part of your development workflow by pushing the boundaries of AI development tools to deliver secure, optimized, and high-quality code.\nDesign and orchestrate complex systems where AI agents integrate seamlessly into human workflows, driving efficiency and innovation at scale.\nCritically evaluate code (Human or AI-generated) for correctness, quality, security, and performance\nContribute to building and maintaining the shared system context, an explicit repository of system designs, constraints, and standards that enables AI to operate accurately and reliably.\nYou're Our Person If You Have:\n5+ years of experience in systems engineering and software engineering for large-scale, internet-facing services.\nHands-on expertise with containerized architectures (Docker, Kubernetes) and orchestration platforms.\nStrong knowledge of distributed systems and Linux/Unix internals, with experience tuning performance and troubleshooting at scale.\nFamiliarity with large-scale internet service architectures (DNS, HTTP, Load Balancing, caching, etc.).\nProven proficiency in Python and Go (GoLang) with strong software engineering practices (testing, code review, CI/CD).\nProduction experience building and operating observability platforms (Grafana, Prometheus, ELK, Splunk, Datadog, or similar)\nSolid background in incident management, including on-call participation, root cause analysis, and postmortem practices.\nStrong understanding of SRE principles: SLIs/SLOs, error budgets, toil reduction, blameless culture, and capacity planning.\nHands-on experience with workflow/orchestration engines (Temporal, Airflow, Argo Workflows, or similar) for building durable automation pipelines.\nExperience applying AI/ML to operations - including anomaly detection, predictive analysis, LLM-based automation, and prompt engineering to build intelligent operational agents and workflows.\nExcellent communication skills with demonstrated ability to lead during high-pressure incidents, present technical designs to leadership, and mentor junior engineers.\nTrack record of mentoring and technically coaching other engineers.\nAbility to work in a 24/7 global operations model, managing multiple priorities under time-sensitive conditions.\nGrowth mindset with curiosity to explore new technologies and drive continuous improvement.\nA demonstrated, genuine AI-first approach to engineering. Using AI to move faster, build fluency across the stack, and contribute well beyond your core specialty.\nExperience using AI tools (e.g., Claude Code, GitHub Copilot, Codex, Cursor, etc.) in development workflows\nAdvanced prompt engineering skills and the ability to write precise, structured prompts and cultivate the system context that makes AI outputs reliable, secure, and production-ready.\nA related technical degree required.\nEven Better If You Have:\nExperience with AI agent frameworks, MCP (Model Context Protocol), or building LLM-powered operational tools.\nContributions to open-source reliability/observability tooling.\nAWS/GCP professional-level certifications.\nPrior experience in SRE organizations supporting multi-cloud or hyperscale environments.\nPython and Go proficiency for systems-level tooling.\nExperience with chaos engineering and game day exercises.\nUnleash Your Potential\nWhen you join Salesforce, you’ll be limitless in all areas of your life. Our benefits and resources support you to find balance andbe your best, and our AI agents accelerate your impact so you cando your best. Together, we’ll bring the power of Agentforce to organizations of all sizes and deliver amazing experiences that customers love. Apply today to not only shape the future - but to redefine what’s possible - for yourself, for AI, and the world.\nAccommodations\nIf you need a reasonable accommodation during the application or the recruiting process, please submit a request via this Accommodations Request Form.\nPlease note that Salesforce uses artificial intelligence (AI) tools to help our recruiters assess and evaluate candidates’ resumes and qualifications throughout the recruiting process. Humans will always make any candidate selection and hiring decisions. Please see our Candidate Privacy Statement for more information about how we use your personal data and your rights, including with regard to use of AI tools and opt out options.\nPosting Statement\nSalesforce is an equal opportunity employer and maintains a policy of non-discrimination with all employees and applicants for employment. What does that mean exactly? It means that at Salesforce, we believe in equality for all. And we believe we can lead the path to equality in part by creating a workplace that’s inclusive, and free from discrimination. Know your rights: workplace discrimination is illegal. Any employee or potential employee will be assessed on the basis of merit, competence and qualifications - without regard to race, religion, color, national origin, sex, sexual orientation, gender expression or identity, transgender status, age, disability, veteran or marital status, political viewpoint, or other classifications protected by law. This policy applies to current and prospective employees, no matter where they are in their Salesforce employment journey. It also applies to recruiting, hiring, job assignment, compensation, promotion, benefits, training, assessment of job performance, discipline, termination, and everything in between. Recruiting, hiring, and promotion decisions at Salesforce are fair and based on merit. The same goes for compensation, benefits, promotions, transfers, reduction in workforce, recall, training, and education.","description_format":"text","description_chars":11008,"description_truncated":false,"requirements":{"experience_years_min":5,"management_years_min":null,"team_size_min":null,"manages_managers":false,"education":{"level":"phd","optional":false},"security_clearance":false,"languages":[]},"benefits":[],"hiring_locations":[],"hiring_excludes":[],"relocation_offered":false,"industries":["Marketing Automation","AI Agents","Team Communication & Collaboration","Customer Support & Help Desk Software"],"lifecycle":[{"event":"open","at":"2026-09-30T01:07:31Z"}],"liveness":{"score":63,"band":"ok","label":"Likely open","p_open":1,"p_active":0.632,"p_room":1,"age_days":2,"expected_fill_days":21,"reasons":["conf:7","stale_co","velocity","win:early","comp:brand"],"computed_at":"2026-10-01T05:45:00Z"},"pay":null,"html_url":"https://alion.io/job/salesforce-senior-site-reliability-engineer-2","json_url":"https://alion.io/job/salesforce-senior-site-reliability-engineer-2.json","meta":{"generated_at":"2026-10-01T12:27:31Z","cache_seconds":300,"methodology":"https://alion.io/methodology","terms":"https://alion.io/terms","contact":"https://alion.io/contact","api":"https://alion.io/developers","usage":{"tier":"crawler","counted_by":"address","units_charged":1,"used_today":3866,"day_limit":5000,"remaining_today":1134,"minute_limit":60,"resets_at":"2026-10-02T00:00:00Z"}}}