657,709open jobs
38,284companies
94,121added this week
Browse all
Salary
$93k – $186k per year (Estimated)
Location
Remote (United States)
Seniority
Senior
Employment
Full-Time
Overview
Company
Impact
Profile match
Play chess online for free on Chess.com with over 250 million members from around the world. Have fun playing with friends or challenging the computer!

About Us

Chess.com is one of the largest gaming sites in the world and the #1 platform for playing, learning, and enjoying chess.

We are a team of 600+ fully remote people in 60+ countries working hard to serve the global chess community. We are here to support 250M+ chess players worldwide with the best possible product, content, and tools to serve the community!

We are a tech company. A gaming company. A content company. And we do it all with passion and commitment to the game. Above all we prize our mission-driven, flat, life-celebrating, no-corporate culture, and we look forward to meeting you and learning more about what you can bring to the team.

About You

Chess.com is the world's largest chess platform with 235M+ members and ~20 million daily games. Our Elasticsearch and OpenSearch infrastructure underpins search, user activity, analytics, logging, and operational intelligence at massive scale, hundreds of terabytes across a dozen production clusters running on bare-metal Kubernetes.

We're looking for a Senior Elasticsearch Engineer who can own the full lifecycle of our search and analytics data platform: capacity planning, cluster architecture, performance tuning, incident response, migration strategy, and operational excellence. You'll be the single point of deep expertise across all Elasticsearch and OpenSearch clusters at Chess.com.

This is not a monitoring-from-dashboards role. You'll be hands-on with cluster internals, write ILM/ISM policies, push infrastructure changes through GitOps, and make real-time decisions about replica allocation when a cluster goes red.

What you'll do

Incident Response & Reliability

  • Shard allocation strategy for write-heavy data streams at high throughput (millions of documents per minute)
  • Disk watermark management, retention policy tuning, and rollover orchestration for high-volume indices
  • Performance optimization and I/O tuning on bare-metal nodes
  • Write queue analysis, thread pool diagnostics, and shard rebalancing under load
  • Capacity planning and growth forecasting across clusters

Incident Response & Reliability

  • On-call ownership for Elasticsearch-related incidents: cluster health degradation, node loss, disk pressure, shard imbalance, and write rejection cascades
  • Real-time cluster triage and cross-team coordination during production incidents
  • Post-mortem authoring and systemic reliability improvements
  • Snapshot and disaster recovery management across clusters

Migration & Strategy

  • Elasticsearch-to-OpenSearch migration analysis and execution, including compatibility evaluation across ILM/ISM, security models, and plugin ecosystems
  • Version upgrade planning and rolling restart orchestration with zero-downtime requirements
  • End-to-end new cluster provisioning and onboarding

Cross-Team Enablement

  • Advise engineering teams on index design, mapping strategy, retention policies, and query optimization
  • Manage Kibana and OpenSearch Dashboards access and configuration for internal consumers
  • Define and maintain workload priority tiers across clusters

Preferred Skills

  • 7+ years operating Elasticsearch at scale (multi-TB clusters, dozens of nodes, high write throughput)
  • Deep understanding of Elasticsearch internals: segment merging, translog, shard allocation, and cluster state management
  • Production experience with ECK (Elastic Cloud on Kubernetes) or equivalent operator-based deployments
  • Proficiency with Kubernetes operations for stateful workloads (StatefulSets, persistent storage, resource management)
  • Hands-on Linux systems administration with a focus on storage and I/O performance
  • Experience managing both Elasticsearch and OpenSearch in production, including an informed opinion on their respective trade-offs
  • Incident command experience: ability to diagnose and mitigate cluster emergencies under pressure while communicating clearly to stakeholders
  • Git-based infrastructure management (GitOps): Helm charts, ArgoCD/Flux, infrastructure-as-code for cluster configuration
  • Fluency with the Elastic stack APIs: cluster administration, index templates, data streams, ILM policies, snapshot/restore

Bonus Experience

  • OpenSearch ISM policies and security plugin (fine-grained access control)
  • GCS or S3 snapshot repository configuration and cross-cluster replication
  • Grafana + Prometheus monitoring for Elasticsearch metrics
  • Kibana Discover, Dev Tools, and data view management at scale
  • Java internals relevant to Elasticsearch JVM tuning (heap sizing, GC tuning, circuit breakers)
  • Vault integration for secrets management in Kubernetes-deployed search clusters
  • Fluentd/Fluent Bit log pipeline configuration feeding OpenSearch
  • Hardware selection experience for search-optimized server configurations
  • Python or scripting for operational analysis and automation

What Makes This Role Special

  • Full autonomy. You are the Elasticsearch authority. You make the architecture calls, set the priorities, and own the outcomes.
  • Real scale. Hundreds of terabytes of data, billions of documents, millions of daily queries. The problems here don't exist at smaller companies.
  • Bare metal. No managed Elastic Cloud. You're operating directly on the hardware. This is hands-on engineering.
  • Strategic impact. Your decisions on ES vs. OpenSearch migration, cluster topology, and capacity planning directly affect product capabilities and infrastructure costs.
  • Small team, big trust.Chess.com runs lean. You won't be buried in process or approvals. Ship changes, fix problems, improve systems.

About the Opportunity

  • This is a full-time opportunity
  • We are 100% remote (work from anywhere!)

---

You can learn more about us here:

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
657,709 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
In your city
$214k – $356k per year • Equity • Remote/Hybrid • Full-Time • 10+ years exp • Bachelor's Degree • San Jose
Python
Go
Java
AI/ML
AI Agents
DevOps
Splunk
Terraform
GCP
GitHub Actions
OpenTelemetry
CloudFormation
Prometheus
GitLab CI
Azure
CI/CD
GitOps
ArgoCD
Jenkins
AWS
Kubernetes
Grafana
Platform Engineering
Backstage
Bazel
Apply
$112k – $207k per year (Estimated) • In office • Full-Time • 8+ years exp • Charlotte • Dallas
Java
Databases
Apache Kafka
DevOps
AWS
Kubernetes
Amazon EKS
Apply
$214k – $356k per year • Equity • In office • Full-Time • 12+ years exp • Bachelor's Degree • Milpitas
Python
Databases
MySQL
Redis
Apache Kafka
AI/ML
AI Agents
DevOps
Splunk
GCP
Datadog
Prometheus
Azure
CI/CD
AWS
Kubernetes
Amazon EKS
Amazon ECS
Cybersecurity
SOC 2
GDPR
FedRAMP
Zero Trust
Apply
$168k – $270k per year • In office • Full-Time • 8+ years exp • Bachelor's Degree • Santa Clara
Python
AI/ML
vLLM
CUDA Toolkit
SGLang
TensorRT
TensorRT-LLM
PyTorch
CUDA
NCCL
DevOps
Splunk
Terraform
Puppet
Ansible
GCP
OpenTelemetry
Chef
Prometheus
Azure
AWS
Kubernetes
Grafana
Argo Workflows
KubeVirt
Incident Management
AWS Step Functions
Apply
$272k – $431k per year • Remote • Full-Time • 5+ years exp • Bachelor's Degree • Westford • Austin • Durham
AI/ML
LLM
DevOps
SLURM
Kubernetes
Shift-Left
Self-Healing
Cybersecurity
Shift-Left Security
Apply
$104k – $210k per year (Estimated) • Remote • Full-Time • 5+ years exp
AI/ML
Claude
AI Agents
Apply
$51k – $128k per year (Estimated) • Remote • Full-Time • 2+ years exp
AI/ML
Cursor
Claude
AI Agents
Management
Slack
Apply
Ad Sales Director 12 days ago
$145k – $262k per year (Estimated) • Remote • Full-Time • 8+ years exp • Bachelor's Degree
Apply
Remote • Contractor
AI/ML
AI Agents
Apply
$113k – $245k per year (Estimated) • Remote • Full-Time
JavaScript
TypeScript
AI/ML
Cursor
Claude
Claude Code
Prompt Engineering
AI Agents
LLM
Tool Use
Frontend
Vue.js
React.js
Design
Figma
Apply
See all jobs
This is one of many
657,709 more open roles from verified company boards, updated every day.