{"id":1038509,"url":"https://alion.io/job/bet365-site-reliability-engineer-4","title":"Site Reliability Engineer","company":{"id":3490,"name":"bet365","domain":"bet365.com","url":"https://alion.io/company/bet365","size_band":"1001-5000","is_staffing_agency":false,"employer_type":"direct","is_intermediary":false,"listed_via":null,"ats_vendor":"SmartRecruiters","truth_index":{"grade":"B","score":75,"open_postings":3,"ghost_share":0,"stale_share":1,"repost_share":0,"time_to_fill_p50_days":22,"computed_at":"2026-10-03T05:45:00Z"}},"role":"DevOps","role_family":"DevOps","seniority":null,"employment_type":"full_time","work_mode":"hybrid","remote_scope":null,"remote_scope_basis":null,"remote_working_hours":null,"hiring_geo_confidence":"structured","locations":["Manchester, United Kingdom"],"countries":["GB"],"hiring_countries":[],"hiring_countries_total":0,"salary":null,"salary_estimate":{"min_usd":50000,"max_usd":123000,"period":"year","method":"role_country_seniority_unknown","sample_n":122},"experience_years_min":null,"visa_sponsorship":false,"relocation_package":false,"has_equity":false,"technologies":[{"name":"Ansible","optional":false},{"name":"Cloudflare","optional":false},{"name":"DNS","optional":false},{"name":"Go","optional":false},{"name":"Grafana","optional":false},{"name":"Incident Management","optional":false},{"name":"JavaScript","optional":false},{"name":"LLM","optional":false},{"name":"New Relic","optional":false},{"name":"OpenTelemetry","optional":false},{"name":"PagerDuty","optional":false},{"name":"Python","optional":false},{"name":"Splunk","optional":false},{"name":"Terraform","optional":false}],"status":"live","first_seen_at":"2026-09-16T11:01:31Z","employer_posted_date":"2026-09-16","last_verified_at":"2026-10-04T01:13:13Z","board_verified":true,"closed_at":null,"days_open":17,"trust":{"level":"ok","repost_count":null,"flags":[],"days_open":17},"description":"We’re one of the world’s leading online gambling companies, revolutionising the industry since 2000. Founded by Denise Coates CBE, we now employ over 10,000 people and serve over 120 million customers in 26 languages.\nWe empower our employees to push boundaries and explore new ideas, cultivating a culture that celebrates and rewards creativity. This offers employees a wealth of growth opportunities, giving them the opportunity to make a real impact in the world of online gambling. As a forward-thinking company, we’re breaking new ground in software innovation too, redefining what’s possible for our global worldwide.\nOur focus on In-Play betting has solidified our market-leading position, featuring more than 1.38 million In-Play sporting events a year. With over 750 concurrent sporting fixtures at peak and more live sports streamed than anyone else in Europe (750,000), we handle over 6 million HTTP requests daily and process more than 1.5 million bets per hour at peak.\n As a Site Reliability Engineer, you will shape the stability of the systems behind every click, query and live change.\nOur Site Reliability team protects and improves the availability, performance and resilience of the systems that support our global product. This role combines software engineering, automation and incident response to reduce toil, sharpen observability and strengthen service health across a complex technical estate.\nYou will work with Open Telemetry, logging, telemetry and automation to surface issues faster and improve operational control. The role also includes using AI tools, LLM platforms and coding assistants to boost productivity, support autonomous operations and improve system insight.\nWorking across SRE, development and IT Operations, you will help embed reliability throughout the software development lifecycle, lead technical work and share knowledge that lifts standards across the wider engineering community.\nThis role is eligible for inclusion in the company’s hybrid work from home policy.\n Software engineering background with Python, Golang, JavaScript or similar language.\nKnowledge of modern development practices, including testing, source control and delivery lifecycles.\nAn understanding of SRE principles, including SLIs, SLOs, reliability measurement and incident management.\nHands-on experience with observability tools such as OpenTelemetry, Splunk, New Relic, Grafana or PagerDuty.\nProficiency in shell scripting for automation and system management.\nExperience with Infrastructure as Code, including Terraform and Ansible.\nKnowledge of Cloudflare or a comparable edge platform, including DNS, CDN, WAF, DDoS protection and traffic management.\nAbility to troubleshoot distributed systems across edge, network, platform, application, dependency and origin layers.\nExperience working in a large-scale, 24/7 enterprise where uptime, performance and stability are critical.\nPractical experience using LLM platforms and coding assistants safely to improve productivity, quality and root-cause analysis.\n Develop and maintain resilient tools, operational APIs and automation for effective system management.\nUse orchestration and scripting to remove manual activity, reduce toil and improve operational consistency.\nWrite and contribute to code, telemetry and instrumentation that improve service reliability and observability.\nBuild dashboards and operational views using telemetry from Grafana, Splunk, New Relic and related platforms.\nConfigure and manage Cloudflare edge services using Infrastructure as Code and integrate edge telemetry with observability platforms.\nDiagnose incidents end to end, trace issues from the edge through to origin systems and coordinate effective remediation.\nParticipate in live incident response, post-mortems and root-cause analysis to prevent recurrence.\nMaintain and administer monitoring, alerting, APM and analytics toolsets, including PagerDuty workflows.\nDrive initiatives that improve reliability, observability, performance and continuous improvement across teams.\nMentor colleagues, share knowledge and work with IT Operations to deliver tooling that increases business value.\nBy applying to us you are agreeing to share your Personal Data in accordance with our Recruitment Privacy Notice - https://www.bet365careers.com/privacy-policy\nAt bet365, we're committed to creating an environment where everyone feels welcome, respected and valued. Where all individuals can grow and develop, regardless of their background. We're Never Ordinary, and we're always striving to be better. If you need any adjustments or accommodations to the recruitment process, at either application or interview, please don’t hesitate to reach out.","description_format":"text","description_chars":4706,"description_truncated":false,"requirements":{"experience_years_min":null,"management_years_min":null,"team_size_min":null,"manages_managers":false,"education":null,"security_clearance":false,"languages":[{"language":"English","level":"All levels","optional":false}]},"benefits":["Growth opportunities","Hybrid work"],"hiring_locations":[{"name":"United Kingdom","iso":"GB","kind":"country"}],"hiring_excludes":[],"relocation_offered":false,"industries":["Gaming","Sports","Sports Betting","iGaming & Gambling"],"lifecycle":[{"event":"open","at":"2026-09-18T17:34:00Z"}],"visa":[],"liveness":{"score":52,"band":"ok","label":"Likely open","p_open":1,"p_active":0.688,"p_room":0.75,"age_days":16,"expected_fill_days":22,"reasons":["conf:26","velocity","win:late"],"computed_at":"2026-10-03T05:45:00Z"},"pay":null,"html_url":"https://alion.io/job/bet365-site-reliability-engineer-4","json_url":"https://alion.io/job/bet365-site-reliability-engineer-4.json","meta":{"generated_at":"2026-10-04T01:46:14Z","cache_seconds":300,"methodology":"https://alion.io/methodology","terms":"https://alion.io/terms","contact":"https://alion.io/contact","api":"https://alion.io/developers","usage":{"tier":"crawler","counted_by":"address","units_charged":1,"used_today":2412,"day_limit":5000,"remaining_today":2588,"minute_limit":60,"resets_at":"2026-10-05T00:00:00Z"}}}