{"id":1434159,"url":"https://alion.io/job/img-site-reliability-engineer-studios","title":"Site Reliability Engineer, Studios","company":{"id":1869062,"name":"IMG","domain":"img.com","url":"https://alion.io/company/img-com","size_band":"201-500","is_staffing_agency":false,"employer_type":"direct","is_intermediary":false,"listed_via":null,"ats_vendor":"Workday","truth_index":{"grade":"B","score":75,"open_postings":4,"ghost_share":0,"stale_share":1,"repost_share":0,"time_to_fill_p50_days":null,"computed_at":"2026-10-04T05:45:00Z"}},"role":"DevOps","role_family":"DevOps","seniority":null,"employment_type":"full_time","work_mode":"hybrid","remote_scope":null,"remote_scope_basis":null,"remote_working_hours":null,"hiring_geo_confidence":"structured","locations":["London, United Kingdom"],"countries":["GB"],"hiring_countries":[],"hiring_countries_total":0,"salary":null,"salary_estimate":{"min_usd":70000,"max_usd":171000,"period":"year","method":"role_country_seniority_unknown","sample_n":122},"experience_years_min":null,"visa_sponsorship":false,"relocation_package":false,"has_equity":false,"technologies":[{"name":"AWS","optional":false},{"name":"Azure","optional":false},{"name":"Bash","optional":false},{"name":"CI/CD","optional":false},{"name":"CloudFormation","optional":false},{"name":"Docker","optional":false},{"name":"GCP","optional":false},{"name":"Kubernetes","optional":false},{"name":"Linux","optional":false},{"name":"Python","optional":false},{"name":"Terraform","optional":false},{"name":"Incident Management","optional":true}],"status":"live","first_seen_at":"2026-09-09T00:00:00Z","employer_posted_date":"2026-09-09","last_verified_at":"2026-10-05T00:45:42Z","board_verified":true,"closed_at":null,"days_open":26,"trust":{"level":"ok","repost_count":null,"flags":[],"days_open":26},"description":"Who We Are:\nIMG is a leading global sports marketing agency, specializing in media rights management and sales, multi-channel content production and distribution, brand partnerships, strategic consulting, digital services, and event management. It powers growth of revenues, fanbases and IP for more than 250 federations, associations, events, and teams, including the National Football League, English Premier League, International Olympic Committee, National Hockey League, Major League Soccer, ATP and WTA Tours, the AELTC (Wimbledon), Euroleague Basketball, CONMEBOL, World Rugby, DP World Tour, and The R&A, as well as UFC, WWE, and PBR. IMG is a subsidiary of TKO Group Holdings, Inc. (NYSE: TKO), a premium sports and entertainment company.\n\nTKO Group Holdings, Inc. (NYSE: TKO) is a premium sports and entertainment company. TKO owns iconic properties including UFC, the world’s premier mixed martial arts organization; WWE, the global leader in sports entertainment; and PBR, the world’s premier bull riding organization. Together, these properties reach 1 billion households across 210 countries and territories and organize more than 500 live events year-round, attracting more than three million fans. TKO also services and partners with major sports rights holders through IMG, an industry-leading global sports marketing agency; and On Location, a global leader in premium experiential hospitality.\n\nWorking Conditions\nPermanent Position, Mon-Fri, 9am-5pm\n\nThis role is based at our facilities in Stockley Park, Uxbridge, with hybrid working options where applicable.\n\nYou may be required to work unsociable hours, including occasional weekends or on-call rotations, to support live operations and critical systems.\n\nOccasional travel may be required depending on project and client needs.\n\nIMG is looking for a Site Reliability Engineer to help design, build, operate, and continuously improve resilient, secure, and highly available platforms that underpin our digital, cloud, and broadcast-adjacent services. This role is suited to someone who combines strong infrastructure and software engineering capability with an operational mindset, and who can help embed reliability engineering practices across systems that support live, business-critical environments.\nThe successful candidate will play a key role in improving service reliability, observability, incident response, automation, and disaster recovery readiness across IMG platforms, while working closely with engineering, operations, and project stakeholders.\nKey Responsibilities and Accountabilities\nDesign, build, and maintain reliable, scalable infrastructure and platform services across on-premises and cloud environments.\n\nImprove service availability, latency, performance, and operational efficiency through engineering-led reliability practices.\n\nBuild and enhance observability across services and infrastructure, including monitoring, logging, alerting, dashboards, and service health indicators.\n\nDefine and maintain SLIs, SLOs, alerting standards, and operational runbooks for critical services.\n\nAutomate infrastructure provisioning, configuration, deployment, and recovery processes using Infrastructure as Code and scripting.\n\nPartner with software, platform, broadcast engineering, and operational teams to improve release quality, resilience, and supportability.\n\nAct as an escalation point for production incidents, leading or supporting rapid diagnosis, mitigation, communication, and post-incident follow-up.\n\nDrive root cause analysis and corrective actions following incidents, with a focus on prevention and continuous improvement.\n\nSupport the design, testing, and documentation of high availability, backup, failover, and disaster recovery arrangements.\n\nHelp enforce security, access control, patching, and operational best practices across infrastructure and services.\n\nOptimise system capacity, cost, and performance across environments.\n\nProduce and maintain clear technical documentation, operational procedures, and support handover materials.\n\nSupport live event and critical operational workflows where reliability, rapid response, and stakeholder communication are essential.\n\nContribute to technical planning for new services, migrations, and platform enhancements, ensuring resilience is designed in from the start.\n\nImprove reliability, stability, and recoverability of IMG’s platform services.\n\nReduced mean time to detect and resolve incidents through better observability and response processes.\n\nHigher levels of automation across provisioning, deployment, remediation, and operational support.\n\nClearer operational ownership, documentation, and service standards across critical environments.\n\nStronger resilience for live and client-facing workflows through tested failover and recovery approaches.\n\nKnowledge and Experience\nMandatory\nProven experience in a Site Reliability Engineer, DevOps Engineer, Platform Engineer, or similar role.\n\nStrong knowledge of Linux and operating system fundamentals.\n\nStrong hands-on experience with cloud platforms such as AWS, Azure, or Google Cloud.\n\nExperience with containerisation and orchestration technologies such as Docker and Kubernetes.\n\nStrong experience with CI/CD tooling and modern software delivery practices.\n\nHands-on experience with Infrastructure as Code tools such as Terraform or CloudFormation.\n\nExperience with monitoring, logging, and alerting tooling, and with designing actionable observability solutions.\n\nSolid understanding of networking, security, system architecture, and distributed systems principles.\n\nStrong scripting or programming capability in Python, Bash, or similar languages.\n\nExperience working in high-availability, live production, or other business-critical operational environments.\n\nStrong troubleshooting skills, calm decision-making under pressure, and a continuous improvement mindset.\n\nExcellent communication and collaboration skills, including the ability to work effectively with technical and non-technical stakeholders.\n\nDesirable\nExperience supporting media, broadcast, streaming, or live event platforms.\n\nFamiliarity with incident management, postmortem practice, and error-budget based operational models.\n\nExperience with resilience engineering, multi-site failover, and disaster recovery testing.\n\nExposure to event-driven or low-latency systems, media transport, or hybrid on-prem/cloud architectures.\n\nUnderstanding of compliance, operational risk management, and support processes in client-facing environments.\n\nPersonal Attributes\nProactive and ownership driven.\n\nMethodical, analytical, and detail oriented.\n\nComfortable operating in fast-moving, high-pressure environments.\n\nPragmatic in balancing engineering excellence with operational needs.\n\nCollaborative, service oriented, and committed to raising reliability standards across teams.\n\nIn addition, success at IMG is driven by four core competencies that apply to all employees:\nBusiness Acumen - Understanding financial drivers, interpreting business data, aligning decisions to strategic outcomes\nOperational Excellence - Driving efficiency, governance, and continuous improvement in delivery & operations\nInnovation Mindset - Cultivating curiosity, experimentation, and forward-looking capability development\nLeadership & Collaboration - Inspiring others, building trust, and enabling collaboration across teams\nTKO EEO Statement\nTKO is an Equal Opportunity Employer and complies with all applicable federal, state, and local laws regarding non-discrimination in employment. TKO makes employment decisions based on merit and qualifications, without considering an employee’s or applicant’s race, color, religion, sex, sexual orientation, gender identity or expression, national origin, age, disability, marital status, veteran status, or any other basis prohibited under federal or local laws governing non-discrimination in employment in every location in which the Company has facilities. TKO also provides reasonable accommodations for qualified individuals with disabilities in accordance with the Americans with Disabilities Act (ADA) and applicable state or local laws. For information about Privacy and Information Security for TKO employment candidates, please review our Privacy Policy. For information regarding Terms of Use for this and other TKO websites, please review our Terms of Use.\n\n TKO EEO Statement:\nTKO is an Equal Opportunity Employer and complies with all applicable federal, state, and local laws regarding non-discrimination in employment. TKO makes employment decisions based on merit and qualifications, without considering an employee’s or applicant’s race, color, religion, sex, sexual orientation, gender identity or expression, national origin, age, disability, marital status, veteran status, or any other basis prohibited under federal, state or local laws governing non-discrimination in employment in every location in which the Company has facilities. TKO also provides reasonable accommodations for qualified individuals with disabilities in accordance with the Americans with Disabilities Act (ADA) and applicable state or local laws. For information about Privacy and Information Security for TKO employment candidates, please review our Privacy Policy. For information regarding Terms of Use for this and other TKO websites, please review our Terms of Use.","description_format":"text","description_chars":9383,"description_truncated":false,"requirements":{"experience_years_min":null,"management_years_min":null,"team_size_min":null,"manages_managers":false,"education":null,"security_clearance":false,"languages":[]},"benefits":["Hybrid work"],"hiring_locations":[{"name":"United Kingdom","iso":"GB","kind":"country"}],"hiring_excludes":[],"relocation_offered":false,"industries":["Events & Ticketing"],"lifecycle":[{"event":"open","at":"2026-09-29T02:27:38Z"}],"visa":[],"liveness":{"score":50,"band":"ok","label":"Likely open","p_open":1,"p_active":0.673,"p_room":0.75,"age_days":25,"expected_fill_days":29,"reasons":["conf:1","velocity","win:late"],"computed_at":"2026-10-04T05:45:00Z"},"pay":null,"html_url":"https://alion.io/job/img-site-reliability-engineer-studios","json_url":"https://alion.io/job/img-site-reliability-engineer-studios.json","meta":{"generated_at":"2026-10-05T01:38:27Z","cache_seconds":300,"methodology":"https://alion.io/methodology","terms":"https://alion.io/terms","contact":"https://alion.io/contact","api":"https://alion.io/developers","usage":{"tier":"crawler","counted_by":"address","units_charged":1,"used_today":2192,"day_limit":5000,"remaining_today":2808,"minute_limit":60,"resets_at":"2026-10-06T00:00:00Z"}}}