1,008,666open jobs
59,992companies
167,270added this week
Browse all
Salary
≈ $20k – $49k per year (Estimated)
Location
Hybrid (Kuala Lumpur, Malaysia)
Seniority
Principal · 8+ years exp

Confirmed on the employer's own hiring board on Oct 1, 2026. First seen by Alion on Aug 7, 2026. Bybit scores A on the Alion truth index.

Overview
Company
Impact
Profile match
Bybit is a centralized cryptocurrency exchange offering spot and derivatives trading, copy trading, earn products, payments, custody and a Web3 wallet to retail and institutional users, headquartered in Dubai. Founded in 2018 by Ben Zhou, it reports more than 80 million users in over 200 countries and regions, ranks among the largest exchanges by trading volume, and suffered a theft of about 1.5 billion dollars in digital assets in a February 2025 hack. Its openings are mostly for product managers, compliance and legal counsel, VIP relationship and user operations managers, marketing and social media specialists and principal-level backend, frontend and site reliability engineers.

About Us

Established in 2018, Bybit is one of the world’s leading cryptocurrency exchanges and digital financial platforms, serving over 80 million users across more than 200 countries and regions. Powered by world-class technology and a user-first mindset, Bybit delivers a seamless ecosystem across trading, payments, wealth management, custody, institutional services, and Web3 - connecting users to the future of digital finance.

Our core values define how we build. We listen, care and improve to create products and experiences that put users first. Backed by a global team of ambitious builders, problem-solvers, and innovators, we foster a high-performance and fast-moving environment where talent is empowered to drive real impact at the global scale. Supported by 24/7 multilingual customer service and a strong commitment to innovation, we are shaping the future of finance through technology, collaboration, and bold execution.

Today, Bybit is recognized as one of the most trusted and transparent platforms in the digital asset industry, continuing to expand its global presence while building the infrastructure for the next generation of financial services.

Core Responsibilities

Chaos Engineering Platform Architecture & Development (50%)

  • Design and build an enterprise-grade chaos engineering platform supporting multi-cluster (K8s + EC2 hybrid), multi-region, and multi-environment (testnet/mainnet) deployments
  • Core capability development:
  • Fault Injection Engine: Pod-level / Node-level / AZ-level fault simulation, network latency / packet loss / partition, dependency timeout / error injection
  • Production Safety Assurance: Blast radius control, one-click Kill Switch, automatic rollback, real-time impact monitoring
  • Traffic Isolation: Experiment traffic tagging and isolation to ensure fault injection does not impact real users
  • Fault Isolation: Precise impact scoping at service / cluster / AZ granularity
  • Design experiment orchestration capabilities supporting complex fault scenario composition (e.g., simultaneous network latency + downstream timeout + cache invalidation)
  • Deep integration with existing monitoring, alerting, and SLO systems to achieve an automated closed loop: inject fault → observe impact → determine pass/fail

Production Resilience Validation Framework (30%)

  • Define safety standards and approval workflows for mainnet fault injection
  • Design and drive routine chaos experiments:
  • Daily patrol-level experiments: Low-risk experiments executed automatically on a daily/weekly basis
  • Periodic validation experiments: Monthly/quarterly resilience verification of critical paths
  • Large-scale drills: Cross-AZ / cross-region disaster recovery failover validation
  • Establish a resilience scoring system to quantify system health based on experiment results
  • Deliver improvement recommendations and drive business teams to remediate identified weaknesses

3. Technology Selection & Team Enablement (20%)

  • Evaluate and select the technology foundation (Chaos Mesh / Litmus / custom components - hybrid strategy)
  • Develop chaos engineering best practices and playbooks to enable SRE teams and application developers
  • Mentor and grow the team (2-3 engineers) in chaos engineering capabilities
  • Stay current with industry developments and introduce cutting-edge practices (e.g., AI-driven fault scenario discovery)

──────

Requirements

Must-Have:

  • 8+ years of backend / infrastructure engineering experience, with 3+ years dedicated to chaos engineering or stability engineering
  • Hands-on experience with large-scale fault injection in production environments (not just test environments), with deep understanding of production safety constraints
  • Expert-level proficiency in Kubernetes fault injection (Chaos Mesh / Litmus / custom solutions), familiar with CRD / Operator development
  • Proficient in at least one backend language (Go preferred), with platform-level system architecture design capability
  • Deep understanding of distributed system failure modes (network partitions, split-brain, cascading failures, data inconsistency, etc.)
  • Familiarity with observability tech stack (Prometheus / Grafana / Thanos / OpenTelemetry)
  • Excellent technical documentation and solution design skills

Nice-to-Have:

  • Experience in financial / trading system stability (understanding of transaction consistency and fund safety constraints)
  • Experience building SLO / Error Budget frameworks
  • Experience building automated fault recovery (self-healing) systems
  • Familiarity with AWS infrastructure (EC2 / EKS / Multi-AZ / Multi-Region)
  • Knowledge of Netflix Chaos Engineering / AWS Fault Injection Simulator / Gremlin
  • Open-source community contributions (Chaos Mesh / Litmus or similar projects)

Soft Skills:

  • Ability to balance "safety" and "validation depth" - not afraid of production injection, while maintaining strict risk control
  • Strong cross-team collaboration and influence - chaos engineering requires buy-in from business teams; this role demands persuasion skills
  • Self-driven, capable of independently planning and executing in ambiguous situations

Why Join Us

At Bybit, we are committed to fostering a supportive and enriching work environment. 

Our benefits include:

- Study Growth Fund: We support your professional development and continuous learning.

- Internal Events: Participate in regular team-building activities, workshops, and events designed to promote collaboration and innovation.

- Global Collaboration: Be part of a diverse, international team, working alongside colleagues from around the world.

- Career Advancement: Access opportunities for growth and advancement within a rapidly expanding global company.

- Internal Mobility: Grow with us- Your long-term development is important to us. We offer internal job opportunities to help build your career path.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
1,008,666 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

DevOps
Similar stack
Same company
Kuala Lumpur
Production Engineer 13 days ago
In office • 3+ years exp • Bachelor's Degree • Kulai
AI/ML
Hadoop
DevOps
HPC
Apply
≈ $17k – $43k per year (Estimated) • In office • 5+ years exp • Bachelor's Degree • Kulai
AI/ML
Hadoop
DevOps
HPC
Linux
Apply
≈ $17k – $43k per year (Estimated) • In office • 5+ years exp • Bachelor's Degree • Kulai
AI/ML
Hadoop
DevOps
HPC
Linux
Apply
≈ $16k – $40k per year (Estimated) • Remote (Malaysia) • Contractor
Python
Bash
DevOps
Terraform
Ansible
Red Hat
VMWare
Azure
CI/CD
GitOps
AWS
Kubernetes
Ubuntu
KVM
Linux
DNS
DHCP
Apply
Remote (Malaysia) • Contractor
Python
Bash
AI/ML
CUDA Toolkit
CUDA
NCCL
DevOps
Terraform
Ansible
Red Hat
OpenShift
OpenTelemetry
VMWare
Prometheus
SLURM
Azure
CI/CD
GitOps
AWS
Kubernetes
Ubuntu
Grafana
KVM
HPC
Linux
DNS
DHCP
Apply
≈ $86k – $156k per year (Estimated) • Equity • Hybrid • 7+ years exp • Munich
Go
Java
Rust
Databases
MySQL
PostgreSQL
AI/ML
Cursor
Claude Code
AI Agents
OpenAI Codex
DevOps
Terraform
Datadog
Kustomize
Azure
ArgoCD
AWS
Docker
Kubernetes
GitHub
Robotics
Digital Twin
Apply
Hybrid • Full-Time • Diegem
DevOps
Cilium
Kubernetes
Service Mesh
eBPF
Apply
Remote (South Korea, China, Vietnam) • Full-Time • 4+ years exp • Bachelor's Degree • Busan • Seongnam • Seoul • Daejeon
Python
C++
AI/ML
Multimodal AI
Physical AI
DevOps
RTOS
AWS
AWS Lambda
Amazon EC2
Amazon S3
Cybersecurity
GDPR
Robotics
Sensor Fusion
Apply
Remote (China, Vietnam) • Full-Time • 4+ years exp • Bachelor's Degree • Shenzhen • Shanghai • Beijing • Guangzhou
Python
C++
AI/ML
Multimodal AI
Physical AI
DevOps
RTOS
AWS
AWS Lambda
Amazon EC2
Amazon S3
Cybersecurity
GDPR
Robotics
Sensor Fusion
Apply
≈ $30k – $75k per year (Estimated) • Hybrid • 5+ years exp • Moscow
Python
Databases
PostgreSQL
Redis
OpenSearch
AI/ML
Model Context Protocol
AI Agents
LLM
RAG
DevOps
Docker
Kubernetes
Apply
≈ $20k – $50k per year (Estimated) • In office • 8+ years exp • PhD • Kuala Lumpur
Python
Go
Rust
DevOps
Terraform
Pulumi
GitOps
Kubernetes
Platform Engineering
SLI/SLO/SLA
Cybersecurity
SBOM
SLSA
Apply
≈ $16k – $42k per year (Estimated) • In office • 5+ years exp • Kuala Lumpur
Java
C++
C++
TensorFlow C++
PyTorch C++
Databases
Milvus
FAISS
Apache Kafka
AI/ML
Quantization
Flink
TensorFlow
PyTorch
Feature Store
Recommender Systems
DevOps
Jaeger
Prometheus
Platform Engineering
SLI/SLO/SLA
Apply
≈ $20k – $48k per year (Estimated) • In office • 3+ years exp • Bachelor's Degree • Kuala Lumpur
Python
Databases
Apache Kafka
DevOps
ZooKeeper
etcd
AWS
Docker
Kubernetes
Linux
Apply
≈ $16k – $41k per year (Estimated) • In office • 8+ years exp • Kuala Lumpur
Python
Go
Java
TypeScript
AI/ML
LangGraph
AutoGen
LangChain
Model Context Protocol
Prompt Engineering
Function Calling
AI Agents
CrewAI
RAG
A2A
Context Engineering
Multi-Agent Systems
Tool Use
DevOps
CI/CD
Git
Platform Engineering
Apply
≈ $65k – $158k per year (Estimated) • Hybrid • 8+ years exp • Hong Kong
Go
DevOps
OpenTelemetry
Prometheus
AWS
Kubernetes
Grafana
Chaos Engineering
Self-Healing
Thanos
Amazon EKS
Amazon EC2
Error Budget
SLI/SLO/SLA
Apply
In office • Internship • Kuala Lumpur
Apply
In office • Full-Time • 7+ years exp • Kuala Lumpur
Analytics
Microsoft Excel
Apply
≈ $20k – $45k per year (Estimated) • In office • Full-Time • 15+ years exp • Bachelor's Degree • Kuala Lumpur
DevOps
SLI/SLO/SLA
Apply
≈ $8.5k – $24k per year (Estimated) • In office • Full-Time • 8+ years exp • Kuala Lumpur
Apply
≈ $16k – $29k per year (Estimated) • In office • Full-Time • Kuala Lumpur
Analytics
Microsoft Excel
Apply
See all jobs
This is one of many
1,008,666 more open roles from verified company boards, updated every day.