368,611open jobs
9,439companies
50,719added this week
Browse all
Salary
$106k – $252k per year (Estimated)
Location
In office (Singapore)
Seniority
Principal · 5+ years exp
Overview
Company
Impact
Profile match
Riot Games is a global video game developer, publisher, and esports tournament organizer headquartered in Los Angeles, California. Founded in 2006 by Brandon Beck and Marc Merrill and acquired by Tencent in 2011, the company is best known for creating League of Legends, one of the most played PC games and dominant esports titles in the world. Riot's product portfolio includes the tactical shooter VALORANT, Teamfight Tactics, League of Legends: Wild Rift, and the fighting game 2XKO, alongside its multimedia production wing responsible for the Emmy Award-winning animated series Arcane.

Riot Games was established in 2006 by entrepreneurial gamers who believe that player-focused game development can result in great games. In 2009, Riot released its debut title League of Legends to critical and player acclaim. As the most played PC game in the world, over 100 million play every month. Players form the foundation of our community and it’s for them that we continue to evolve and improve the League of Legends experience. 

We’re looking for humble but ambitious, razor-sharp professionals who can teach us a thing or two. We promise to return the favor. Like us, you take play seriously; you’re passionate about games. We embrace those who see things differently, aren’t afraid to experiment, and who have a healthy disregard for constraints.

That's where you come in.

The AI Efficiency team at Riot Games builds the platforms, tools, and technical foundations that help Rioters safely and effectively use AI to accelerate how we work. As these systems become increasingly important to creative, product, and development workflows across Riot, we need dedicated engineering leadership to ensure these systems remain stable, scalable, secure, and dependable in production.

As a Principal DevOps / Site Reliability Enginee r on the AI Efficiency team, you will own and evolve the operational foundations that allow the AI Efficiency team’s tech platform and the tools deployed within it to run reliably at growing scale. You will establish the systems, standards, automation, and support practices required to move quickly without compromising availability, deployment safety, maintainability, or user trust.

You will partner closely with software engineers, ML platform engineers, technical artists, data scientists, and Riot’s infrastructure and security teams to improve developer experience, production readiness, observability, incident response, capacity planning, and service resilience. You will also help evaluate and operationalize AI-native engineering workflows such as agent-assisted code review, automated bug triage, AI-driven performance and security analysis, and browser-based UI validation. This role ensures the broader platform and its services are safely operated, supported, and continuously improved in production.

You’re right for this role if you enjoy making complex systems reliable, reducing operational toil, improving how engineers build and ship software, and anticipating how systems will fail before those failures affect users. You are comfortable taking ownership of production health, leading through incidents, building sustainable operational practices, and creating paved roads that help teams move quickly and safely. You are also energized by the opportunity to responsibly bring new AI-native automation patterns into real engineering workflows, thoughtfully applying emerging capabilities to reduce friction, improve reliability, and enhance how engineers interact with production systems without compromising safety or control.

Responsibilities:

  • Own and continuously improve the reliability, availability, scalability, performance, and operational health of the Efficiency team’s (web) platform and the tools deployed within it
  • Design, build, and maintain the infrastructure, deployment systems, and operational foundations required to support a growing portfolio of production AI services and internal tools
  • Improve CI/CD pipelines, release engineering practices, environment management, and deployment automation so software can be shipped safely, quickly, and consistently
  • Establish production-readiness standards and ensure new utilities have appropriate monitoring, alerting, ownership, documentation, rollback strategies, and support plans before launch
  • Define and operationalize service health indicators, SLIs, SLOs, error budgets, and reliability metrics that guide engineering priorities and tradeoffs between reliability, velocity, cost, and complexity
  • Build comprehensive observability across applications, infrastructure, service dependencies, and user workflows using metrics, logs, traces, dashboards, synthetic monitoring, and actionable alerts
  • Establish sustainable incident-management and on-call practices, including escalation paths, runbooks, severity definitions, communication protocols, and clear service ownership
  • Lead or contribute to the diagnosis and resolution of production incidents, coordinating across teams and driving blameless post-incident reviews and durable corrective actions
  • Build automation that reduces operational toil, improves mean time to detect and recover, and eliminates recurring sources of failure or manual intervention
  • Implement safe deployment patterns such as automated validation, progressive delivery, canary releases, feature flags, health checks, rollback mechanisms, and controlled environment promotion
  • Perform capacity planning, load testing, performance analysis, and resource forecasting to ensure the Toolkit can support increasing adoption and usage across Riot
  • Design and validate resilience, backup, recovery, failover, and disaster-recovery strategies for critical services, data, configurations, and infrastructure
  • Identify single points of failure and systemic risks across applications, cloud infrastructure, networking, databases, queues, caches, third-party dependencies, and operational workflows
  • Improve developer experience by building self-service workflows, reusable infrastructure components, local development environments, test environments, deployment tooling, and clear operational documentation
  • Establish and maintain infrastructure-as-code, configuration-management, secrets-management, and environment-governance practices that make infrastructure changes safe, repeatable, and auditable
  • Partner with engineers throughout the software development lifecycle to embed reliability, operability, security, and maintainability into system design rather than addressing them only after launch
  • Troubleshoot complex production issues across web applications, APIs, distributed services, containerized workloads, cloud infrastructure, network boundaries, authentication systems, and external service dependencies
  • Partner with ML Platform Engineers to ensure model-serving and inference systems integrate cleanly with the team’s broader observability, deployment, incident-management, and reliability standards
  • Collaborate with Riot infrastructure, information security, IT, developer-platform, and compliance teams to ensure the team follows appropriate operational and security requirements
  • Evaluate and implement AI-assisted operational workflows such as automated anomaly investigation, log analysis, remediation recommendations, regression detection, and runbook automation
  • Define guardrails, approval requirements, auditability, and escalation paths for agentic or automated operational systems that can interact with production environments
  • Champion operational excellence through technical leadership, mentoring, documentation, standards, architecture reviews, and tooling that raise the reliability bar across the team

Required Qualifications:

  • Bachelor’s degree in Computer Science or a related field, or equivalent professional experience
  • 5+ years of experience in Site Reliability Engineering, DevOps, Infrastructure Engineering, Platform Engineering, Production Engineering, Developer Experience, or a similar role supporting production systems
  • Strong programming and automation skills in one or more languages such as Python, Go, JavaScript, or TypeScript
  • Experience designing, operating, and improving cloud-based production systems in AWS, GCP, Azure, or comparable environments
  • Experience building and maintaining CI/CD pipelines, release systems, deployment automation, and environment-management workflows
  • Strong understanding of observability practices, including metrics, logging, distributed tracing, dashboards, synthetic monitoring, and alert design
  • Experience participating in or leading incident response, on-call support, root-cause analysis, and post-incident improvement work
  • Experience improving the reliability, availability, scalability, and performance of distributed systems, service-oriented architectures, APIs, or web platforms
  • Strong understanding of containerized environments and orchestration technologies such as ECS, Docker, Kubernetes, or comparable systems
  • Experience with infrastructure-as-code and configuration-management tools such as Terraform, Pulumi, CloudFormation, or similar technologies
  • Working knowledge of Linux systems, networking, DNS, load balancing, service discovery, authentication, secrets management, and cloud security fundamentals
  • Ability to identify systemic operational risks and drive durable improvements across systems owned by multiple engineers or teams
  • Ability to collaborate across organizational boundaries, influence technical direction, and communicate clearly during both planned work and high-pressure incidents
  • Experience providing technical leadership, mentoring engineers, and establishing engineering standards across a team or organization

Desired Qualifications: 

  • Experience supporting AI/ML platforms, inference services, model-serving systems, GPU-backed workloads, data pipelines, or other compute-intensive services
  • Experience defining and using SLOs, error budgets, and reliability metrics to guide prioritization and engineering decisions
  • Experience designing or improving internal developer platforms, self-service infrastructure, paved roads, golden paths, or shared engineering services
  • Experience building sustainable on-call rotations and operational support models for services used by multiple teams
  • Experience with progressive delivery, canary deployments, blue-green deployments, feature-flag systems, and automated rollback strategies
  • Experience with performance testing, capacity modeling, chaos engineering, fault injection, resilience testing, or failure-mode analysis
  • Experience designing backup, disaster-recovery, business-continuity, and regional failover strategies
  • Experience operating databases, caches, message queues, object storage, service meshes, API gateways, and other common distributed-system components
  • Experience improving security posture through access controls, secrets management, dependency management, vulnerability remediation, network segmentation, and infrastructure hardening
  • Experience balancing availability, latency, engineering velocity, infrastructure efficiency, and cost in systems operating at scale
  • Familiarity with browser automation and end-to-end testing frameworks such as Playwright for validating critical user workflows and detecting production regressions
  • Experience evaluating or integrating AI-assisted tools for incident investigation, anomaly detection, operational diagnostics, code review, test generation, or automated remediation
  • Familiarity with the risks and operational controls required when AI agents interact with source control, CI/CD pipelines, cloud infrastructure, or production systems
  • Experience establishing governance, approval workflows, audit trails, and quality controls for automated operational systems
  • Experience working in environments where experimental tools must be transitioned into reliable, supported, and maintainable production services

For this role, you'll find success through craft expertise, a collaborative spirit, and decision-making that prioritizes your fellow Rioters, who are the customers of your work. Being a dedicated fan of games is not necessary for this position!

Our Perks:

  • Full relocation support
  • Comprehensive health insurance for you, your spouse, and children
  • Open paid time off
  • Retirement benefits with company matching
  • Life insurance, parental leave, plus short-term and long-term disability
  • Play Fund so you can deepen your knowledge of our players and community through games
  • We’ll double down on your donations of time and money to non-profits
Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
368,611 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
Singapore
$191k – $267k per year • Equity • Remote • Full-Time • 4+ years exp • Master's Degree
Python
SQL
AI/ML
Anomaly Detection
Apply
$105k – $252k per year • Remote • Full-Time • 18+ years exp • Bachelor's Degree
Python
Java
Java
Gradle
DevOps
Ansible
AWS
CI/CD
CloudFormation
Configuration Management
Docker
GitHub Actions
GitLab CI
Helm
Jenkins
Kubernetes
Platform Engineering
Terraform
GitHub
GitLab
Cybersecurity
Sonatype Nexus IQ
Management
Confluence
Jira
Apply
$54k – $175k per year (Estimated) • Remote • Full-Time
JavaScript
Node JS
TypeScript
Frontend
React.js
Sass
DevOps
AWS
CI/CD
Docker
Kubernetes
Rest API
Apply
$35k – $113k per year (Estimated) • Remote • Full-Time
JavaScript
Node JS
TypeScript
Frontend
React.js
Sass
DevOps
AWS
CI/CD
Docker
Kubernetes
Rest API
Apply
$35k – $116k per year (Estimated) • Remote • Full-Time
JavaScript
Node JS
TypeScript
Frontend
React.js
Sass
DevOps
AWS
CI/CD
Docker
Kubernetes
Rest API
Apply
$129k – $235k per year (Estimated) • In office • 5+ years exp • Los Angeles
Game Dev
Unreal Engine
Design
Adobe Photoshop
Figma
Apply
$188k – $352k per year (Estimated) • In office • 6+ years exp • Bachelor's Degree • Los Angeles
C#
C++
Java
DevOps
CI/CD
Game Dev
Unreal Engine
Management
Slack
Apply
In office • Guangzhou
Game Dev
Unity
Design
Axure RP
Figma
Apply
$29k – $110k per year (Estimated) • Remote/Hybrid • Contractor • 4+ years exp • Singapore
Python
DevOps
Git
Game Dev
GLSL
HLSL
Marmoset Toolbag
RenderDoc
Unity
Design
Adobe Photoshop
Maya
Apply
In office • Contractor • 5+ years exp • Shanghai
Design
Adobe After Effects
Adobe Photoshop
Blender
Cinema 4D
Figma
Management
Miro
Apply
$117k – $251k per year (Estimated) • Remote/Hybrid • Full-Time • Singapore
Apply
$74k – $126k per year (Estimated) • In office • Full-Time • Singapore
Python
Apply
$88k – $191k per year (Estimated) • Remote/Hybrid • Full-Time • 5+ years exp • Singapore
C++
Java
Kotlin
Python
Mobile
JUnit
DevOps
Git
gRPC
Jenkins
JFrog Artifactory
Shift-Left
Cybersecurity
Shift-Left Security
QA
Pytest
Robot Framework
TestNG
Apply
Senior AI Architect 6 hours ago
$138k – $304k per year (Estimated) • In office • Full-Time • 6+ years exp • Bachelor's Degree • Singapore
Python
SQL
Databases
Databricks
AI/ML
AI Agents
LangGraph
OpenAI
RAG
Spark
LangChain
DevOps
Azure
Apply
$64k – $189k per year (Estimated) • Remote/Hybrid • Full-Time • 1+ year exp • Bachelor's Degree • Singapore
Python
SQL
Databases
Apache Kafka
AI/ML
Amazon SageMaker
Kubeflow
MLFlow
Spark
Vertex AI
DevOps
AWS
Azure
Azure DevOps
CI/CD
Docker
GCP
GitLab
GitLab CI
Jenkins
Kubernetes
Apply
See all jobs
This is one of many
368,611 more open roles from verified company boards, updated every day.