657,963open jobs
38,241companies
93,447added this week
Browse all
Salary
$96k – $208k per year (Estimated)
Location
Remote (United States)
Seniority
Senior · 5+ years exp
Employment
Full-Time
Overview
Company
Impact
Profile match
Play chess online for free on Chess.com with over 250 million members from around the world. Have fun playing with friends or challenging the computer!

About Us

Chess.com is one of the largest gaming sites in the world and the #1 platform for playing, learning, and enjoying chess.

We are a team of 600+ fully remote people in 60+ countries working hard to serve the global chess community. We are here to support 250M+ chess players worldwide with the best possible product, content, and tools to serve the community!

We are a tech company. A gaming company. A content company. And we do it all with passion and commitment to the game. Above all we prize our mission-driven, flat, life-celebrating, no-corporate culture, and we look forward to meeting you and learning more about what you can bring to the team.

About the role

The Site Reliability Engineer will play a critical role in ensuring the stability, performance, and scalability of our global gaming platform infrastructure. This position exists to bridge the gap between development and operations, maintaining high availability for millions of concurrent users while supporting rapid feature development and deployment. The SRE will be instrumental in building resilient systems that can handle massive scale across multiple regions, directly impacting user experience and platform reliability.

As our platform continues to grow and serve a global community, this role will drive the technical infrastructure decisions that enable seamless gaming experiences. The position requires both deep technical expertise and collaborative leadership to work across engineering teams, ensuring our systems can scale efficiently while maintaining the performance standards our users expect.

What you'll do

  • Design and implement multi-regional resilient infrastructure capable of handling millions of concurrent sessions and transactions daily across global data centers
  • Lead the hybrid cloud migration strategy, integrating bare-metal datacenter resources with cloud services for optimal performance and cost efficiency
  • Own the on-call rotation and incident response procedures, ensuring rapid resolution of critical system issues and maintaining high availability SLAs
  • Architect monitoring and alerting systems using industry-standard tools to proactively identify and resolve performance bottlenecks before they impact users
  • Collaborate with development teams to implement infrastructure-as-code practices and establish deployment pipelines that support continuous integration and delivery
  • Optimize system performance through capacity planning, load testing, and resource allocation across distributed computing environments
  • Establish and maintain security protocols and risk assessment procedures for infrastructure components and data protection
  • Partner with engineering teams to design scalable solutions for high-traffic applications and real-time processing requirements
  • Drive automation initiatives to reduce manual operational overhead and improve system reliability through scripting and configuration management
  • Mentor team members on SRE best practices and contribute to the development of infrastructure standards and documentation

Preferred Skills

  • Bachelor's degree in Computer Science, Engineering, or related technical field, or equivalent practical experience
  • 5+ years of experience in site reliability engineering, DevOps, or infrastructure engineering roles
  • Experience managing bare-metal server infrastructure and datacenter operations
  • Strong proficiency with UNIX/Linux operating systems and command-line administration
  • Experience with cloud platforms (GCP, AWS, or Azure) and infrastructure-as-code tools (Terraform, CloudFormation, or similar)
  • Hands-on experience with configuration management systems (Ansible, Chef, Puppet, or similar)
  • Solid understanding of networking fundamentals, protocols (TCP/IP, HTTP/HTTPS, DNS), and network troubleshooting
  • Experience with containerization and orchestration technologies (Docker, Kubernetes, or similar)
  • Proficiency with monitoring and observability tools (Datadog, Prometheus, Grafana, ELK stack, or similar)
  • Experience with relational and NoSQL databases, including performance optimization and scaling strategies
  • Strong collaboration and communication skills for working effectively in a distributed team environment
  • Demonstrated sense of ownership and accountability for system reliability and performance

Nice to have

  • Advanced knowledge of content delivery networks (CDNs) and edge computing
  • Experience with server-side automation and scripting languages (Python, Go, Bash, or similar)
  • Background in high-availability architectures and disaster recovery planning
  • Familiarity with security frameworks and compliance requirements
  • Experience with game server infrastructure or real-time application hosting
  • Knowledge of database administration and optimization for high-concurrency applications
  • Understanding of CI/CD pipelines and deployment automation
  • Experience with capacity planning and performance testing tools
  • Previous experience in a fully remote, distributed work environment
  • Continuous learning mindset with interest in emerging infrastructure technologies

About the Opportunity

  • This is a full-time opportunity
  • We are 100% remote (work from anywhere!)

---

You can learn more about us here:

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
657,963 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
In your city
$160k – $210k per year • Remote/Hybrid • Full-Time • 5+ years exp • Bachelor's Degree • San Francisco
JavaScript
TypeScript
SQL
Node JS
Databases
PostgreSQL
DynamoDB
AI/ML
Edge AI
Frontend
GraphQL
DevOps
GCP
Azure
CI/CD
AWS
Docker
Kubernetes
SRE
Cybersecurity
HIPAA
Auth0
Apply
$29k – $63k per year (Estimated) • Remote • Full-Time • 6+ years exp • Bachelor's Degree
Python
Rust
Bash
DevOps
Terraform
Ansible
Prometheus
CI/CD
Jenkins
Git
Kubernetes
Grafana
SaltStack
Configuration Management
Apply
$87k – $175k per year (Estimated) • Remote • Full-Time • 5+ years exp • Bachelor's Degree
Python
C++
C++
CMake
DevOps
GitHub Actions
CI/CD
Jenkins
Docker
Buildkite
Bazel
Management
Agile
Apply
$85k – $171k per year (Estimated) • Equity • Remote • Full-Time • 7+ years exp • Bachelor's Degree
Python
TypeScript
Python
Flask
FastAPI
Django
Databases
Snowflake
Databricks
Google BigQuery
Amazon Redshift
BigQuery
AI/ML
Cursor
AI Agents
LLM
RAG
Human-in-the-Loop
LLM Guardrails
Agentic Workflows
DevOps
CI/CD
AWS
Docker
Kubernetes
Apply
$200k – $300k per year • Remote/Hybrid • Full-Time • 8+ years exp • San Francisco • New York
Python
Databases
Redis
Amazon DocumentDB
DevOps
Terraform
GitHub Actions
Terragrunt
CI/CD
AWS
Amazon S3
IAM
Amazon ECS
Amazon CloudWatch
Cybersecurity
HIPAA
Apply
$104k – $210k per year (Estimated) • Remote • Full-Time • 5+ years exp
AI/ML
Claude
AI Agents
Apply
$51k – $128k per year (Estimated) • Remote • Full-Time • 2+ years exp
AI/ML
Cursor
Claude
AI Agents
Management
Slack
Apply
Ad Sales Director 12 days ago
$145k – $262k per year (Estimated) • Remote • Full-Time • 8+ years exp • Bachelor's Degree
Apply
Remote • Contractor
AI/ML
AI Agents
Apply
$113k – $245k per year (Estimated) • Remote • Full-Time
JavaScript
TypeScript
AI/ML
Cursor
Claude
Claude Code
Prompt Engineering
AI Agents
LLM
Tool Use
Frontend
Vue.js
React.js
Design
Figma
Apply
See all jobs
This is one of many
657,963 more open roles from verified company boards, updated every day.