368,657open jobs
9,442companies
50,883added this week
Browse all
Location
Remote (Azerbaijan)
Seniority
Middle · 4+ years exp
Employment
Full-Time
Overview
Company
Impact
Profile match
Xsolla is a global video game commerce company that provides specialized financial and operational tools tailored for the gaming industry. The firm helps game developers and publishers fund, launch, market, and monetize their titles across PC, mobile, web, and cloud platforms. By operating as a merchant of record with support for over 1,000 local payment methods, it enables direct-to-consumer sales and seamless cross-border transactions for gaming studios worldwide.

ABOUT YOU

We are looking for an Operations Engineer who is technically curious, detail-oriented, a strong communicator, and proactive to join our Global Technical Operations (GTO) team. The best candidate will be someone who thrives in a fast-paced, highly collaborative, and exceptionally dynamic setting and is excited to monitor and investigate production issues across a global platform, help improve how we detect and respond to incidents, analyze trends and patterns in production data, and contribute to better communication with partners and stakeholders during incidents.

Strong troubleshooting skills, observability platform experience, and scripting ability are essential, along with experience in SRE, DevOps, production operations, or NOC environments supporting high-availability platforms (payments, e-commerce, SaaS, or gaming). The ability to communicate clearly and effectively in English - both written and verbal - when writing incident updates, shift handoffs, and status page communications will be key to your success in this role.

If you're passionate about keeping critical systems running and continuously improving operational processes and love being the first to spot issues and the one who drives them to resolution for game developers and players worldwide, we would love to hear from you!

ABOUT US

Xsolla is a global commerce company with robust tools and services to help developers solve the inherent challenges of the video game industry. From indie to AAA, companies partner with Xsolla to help them fund, distribute, market, and monetize their games. Grounded in the belief in the future of video games, Xsolla is resolute in the mission to bring opportunities together, and continually make new resources available to creators.

Headquartered and incorporated in Los Angeles, California, Xsolla operates as the merchant of record and has helped over 1,500+ game developers to reach more players and grow their businesses around the world. With more paths to profits and ways to win, developers have all the things needed to enjoy the game.

For more information, visit xsolla.com.

Responsibilities

    • Serve as the primary dashboard monitor during your shift - continuously watch the GTO Operational Dashboard in Datadog, detect anomalies by correlating signals across APM, logs, metrics, synthetic tests, and Real User Monitoring, and determine whether alerts warrant an incident ticket or can be resolved through immediate investigation.

    • Triage and investigate production incidents - create incident tickets in JIRA Service Management, perform initial technical investigation using Datadog (traces, logs, infrastructure and application metrics), determine blast radius and likely root cause domain, and route to the correct team (Product SRE, Infrastructure SRE, or Engineering) using the smart routing model.

    • Own lower-severity incidents end-to-end from detection through resolution - diagnose, execute runbook procedures, and resolve without escalation where possible. Escalate promptly when an incident is unresolved within defined thresholds or requires a code-level fix.

    • Support the TSO Lead during major incidents as the technical right hand in the war room - surface real-time data (error rates, impact scope, deployment history, related alerts), maintain the incident ticket with live timeline entries and linked evidence, and execute mitigation actions as directed.

    • Draft incident communications under TSO Lead direction, including internal Slack updates, stakeholder notifications, and customer-facing status page updates (status.xsolla.com). Support clear, timely communication throughout the incident lifecycle.

    • During non-incident periods, analyze incident trends, recurring issues, and production bugs - compile data from Datadog, JIRA, and Slack, identify patterns, and contribute findings to regular reports for product and engineering teams.

    • Compile incident timelines and draft initial PIR documents for Post-Incident Review preparation. Track PIR action items post-session and flag overdue items to the TSO Lead.

    • Build and maintain operational automation (alert enrichment scripts, incident templates, Slack workflows, dashboard widgets) and contribute to runbook development - documenting new resolution procedures so they can be repeated by any Operations Engineer on any shift.

    • Conduct structured shift handoffs covering active incidents, at-risk services, upcoming deployments, and follow-up items. Participate in knowledge transfer sessions with SREs to continuously expand independent resolution capability.

    • Cover for the TSO Lead during vacations, absences, or emergencies - including severity classification, escalation decisions, stakeholder communications, and basic Incident Commander functions.

    • Publish health reports of critical apps periodically.

Qualifications and Skills

    • 4+ years of experience in SRE, DevOps, production operations, NOC, or technical operations in a high-availability environment. Experience with platforms that handle payments, e-commerce, SaaS, or gaming workloads is preferred.

    • Strong troubleshooting and investigation skills - ability to take an alert or user-reported symptom and methodically trace it through the stack: application logs, APM traces, infrastructure metrics, database queries, and network paths.

    • Hands-on experience with Datadog (or equivalent observability platform: Grafana, Splunk, New Relic, Elastic) - navigating APM, building log queries, reading infrastructure dashboards, interpreting SLO burn rates, and configuring monitors and alerts.

    • Proficiency in at least one scripting language: Python, Go, or Bash. You will write automation scripts, build operational tooling, and work with APIs.

    • Clear written and verbal communication skills in English - ability to write incident tickets, investigation notes, Slack updates, shift handoff reports, status page communications, and PIR drafts that are clear, concise, and useful to both technical and non-technical audiences.

    • Working knowledge of Kubernetes and cloud infrastructure (GCP preferred, AWS/Azure acceptable) - understanding of pods, deployments, services, ingress, node health, and how to investigate Kubernetes-related production issues.

    • Understanding of SLOs, error budgets, and burn-rate alerting - knowing what a multi-window burn-rate alert means, how error budgets deplete, and how SLO breaches translate into incident severity.

    • Experience with incident management tooling: JIRA or JIRA Service Management, PagerDuty or OpsGenie, Slack, and Confluence.

    • Experience with or strong interest in AI/ML-assisted operations: anomaly detection, alert correlation, predictive monitoring, or automated remediation.

    • Comfort with 24x7 shift-based operations as part of a follow-the-sun model with handoff overlaps. Weekend on-call (rotating) is required.

Nice to Have

    • Experience in the gaming, payments, or fintech industry - particularly environments where transaction processing, checkout flows, or player-facing services must meet strict uptime requirements.

    • Familiarity with Datadog Service Catalog, synthetic monitoring, and RUM; exposure to database operations (MySQL, PostgreSQL, Redis, Kafka); and experience with CI/CD pipelines and deployment tooling (GitLab CI, ArgoCD, Helm).

    • JIRA Service Management administration experience (workflows, automation rules, SLA timers) or ITIL Foundation certification - practical experience matters more than credentials.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
368,657 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
Baku
$169k – $321k per year (Estimated) • In office • 8+ years exp • Bachelor's Degree • Phoenix
AI/ML
AI Agents
Anomaly Detection
LLM Guardrails
DevOps
AWS
Kong
Amazon S3
API Gateway
Cybersecurity
Zero Trust
Apply
$105k – $252k per year • Remote • Full-Time • 18+ years exp • Bachelor's Degree
Python
Java
Java
Gradle
DevOps
Ansible
AWS
CI/CD
CloudFormation
Configuration Management
Docker
GitHub Actions
GitLab CI
Helm
Jenkins
Kubernetes
Platform Engineering
Terraform
GitHub
GitLab
Cybersecurity
Sonatype Nexus IQ
Management
Confluence
Jira
Apply
$185k – $260k per year • Remote • Full-Time • 8+ years exp • Bachelor's Degree
DevOps
AWS
CI/CD
GCP
Kubernetes
GitHub
Cybersecurity
Clair
Dependabot
OWASP Top 10
OWASP ZAP
Snyk
Trivy
Apply
$110k – $131k per year • Remote • Full-Time • 10+ years exp • Bachelor's Degree
DevOps
AWS
Incident Management
VMWare
Apply
$158k – $189k per year • In office • Full-Time • 8+ years exp • Bachelor's Degree • Westminster
Python
DevOps
Bitbucket
Git
GitLab
Management
Confluence
SpaceTech
GMAT
Apply
$14k – $26k per year (Estimated) • In office • Contractor • 4+ years exp • Bachelor's Degree • Vladivostok
Node JS
TypeScript
JavaScript
Frontend
esbuild
GraphQL
React.js
styled-components
Webpack
Zod
DevOps
CI/CD
GitLab CI
GitLab
QA
Cypress
Jest
Playwright
Vitest
Apply
$15k – $37k per year (Estimated) • In office • 5+ years exp • Bachelor's Degree • Perm
Bash
Python
SQL
Databases
MySQL
PostgreSQL
DevOps
AWS
Datadog
GCP
Grafana
Prometheus
Puppet
Terraform
Zabbix
Apply
$14k – $35k per year (Estimated) • In office • Perm
Go
PHP
SQL
DevOps
CI/CD
Apply
$12k – $27k per year (Estimated) • In office • 3+ years exp • Perm
Go
PHP
Python
DevOps
CI/CD
Datadog
GCP
GitHub Actions
GitLab CI
Google GKE
Grafana
Helm
Kubernetes
OpenTelemetry
Prometheus
SLI/SLO/SLA
Terraform
Terragrunt
GitHub
GitLab
IAM
Apply
IT Support Engineer 5 days ago
$31k – $67k per year (Estimated) • In office • Full-Time • 3+ years exp • Bachelor's Degree • Berlin
Cybersecurity
Okta
Management
Google Workspace
Apply
Remote • 2+ years exp • Baku
Go
SQL
TypeScript
JavaScript
Frontend
React.js
Apply
$24k – $36k per year (gross) • Remote • Full-Time • Baku
TypeScript
Databases
PostgreSQL
Supabase
Apply
Remote • 7+ years exp • Baku
DevOps
AWS
AWS Fargate
CI/CD
CircleCI
CloudFormation
Datadog
GitHub Actions
Grafana
Kubernetes
New Relic
OpenTelemetry
Platform Engineering
Terraform
Amazon ECS
GitHub
IAM
Cybersecurity
Snyk
SonarQube
Apply
Remote • Internship • Bachelor's Degree • Baku
PHP
Python
SQL
Databases
Apache Kafka
RabbitMQ
DevOps
CI/CD
Git
Helm
Jenkins
Kubernetes
Terraform
WebSockets
Web3
DeFi
Smart Contracts
Apply
Tech Lead 5 days ago
In office • Full-Time • 5+ years exp • Baku
Go
JavaScript
PHP
Frontend
Next.js
React.js
DevOps
CI/CD
Git
Incident Management
Apply
See all jobs
This is one of many
368,657 more open roles from verified company boards, updated every day.