Salary
≈ $37k – $91k per year (Estimated)
Location
In office (Hong Kong)
Seniority
Senior
Overview
Company
Impact
Profile match
kodypay is on a mission to make in-person payment acceptance easy. today, paying in person presents common problems for businesses, such as high costs, long queues, and limited choice of payment methods. kodypay fully integrates the payment ecosys...
Job Summary
Kody is seeking a Senior Site Reliability Engineer (8+ years of experience) to drive the reliability, availability, scalability, and operational excellence of our global payment platform. Based in Hong Kong or Shenzhen, you will take end-to-end ownership of production observability, incident response, service-level management, and cloud infrastructure reliability across mission-critical payment processing systems operating across Europe, Asia, and North America.
Key Responsibilities
- Incident Management & On-Call: Participate in a follow-the-sun production on-call rotation as a senior incident responder. Lead incident management during SEV1/SEV2 events to optimize MTTR and operational effectiveness.
- Production Operations: Diagnose, triage, mitigate, and coordinate the resolution of complex production incidents across payment services, Kubernetes platforms, databases, messaging systems, and cloud infrastructure.
- SLO & Reliability Engineering: Define, implement, and maintain SLOs, SLIs, error budgets, alerting standards, and operational readiness processes across distributed services.
- Continuous Optimization: Drive systemic reliability improvements through infrastructure automation, observability enhancement, capacity planning, performance tuning, and post-incident root-cause analysis (RCA).
- Security & Compliance: Partner with global engineering teams to strengthen architectural resilience, security posture, and operational maturity in PCI-DSS-regulated payment environments.
- Technical Leadership: Mentor junior engineers, eliminate operational toil through automation, and influence engineering teams to adopt resilience-by-design practices.
Requirements
Qualifications & Requirements
- Experience: 8+ years of hands-on experience in Site Reliability Engineering, Platform Engineering, DevOps, or Cloud Infrastructure roles supporting high-availability, mission-critical production systems.
- Core Technical Stack: Strong expertise in AWS, Kubernetes (EKS), Terraform, PostgreSQL, Redis, Kafka, Linux, networking, and modern observability platforms (e.g., Datadog, Prometheus, Grafana).
- Distributed Systems Mastery: Deep understanding of distributed systems architecture, high availability, disaster recovery, capacity planning, and microservices orchestration.
- Domain Expertise: Proven track record operating in payment, banking, fintech, or other highly regulated environments with strict PCI-DSS, security, and uptime standards.
- SRE Methodology: Deep knowledge of core SRE principles, including SLO/SLI design, error budget management, alert governance, and toil reduction.
- Location & Communication: Based in Hong Kong or Shenzhen. Excellent command of English (written and spoken) to lead cross-functional incident responses and collaborate seamlessly with global teams.
Leadership & Operational Excellence
- Ownership: Demonstrates strong end-to-end accountability for service reliability and customer impact under high pressure.
- Structured Problem Solving: Applies a systematic and data-driven approach to troubleshooting, telemetry analysis, and incident resolution in complex distributed environments.
- Crisis Management: Proven ability to command cross-functional incident response efforts, align stakeholders, and maintain clear communication during critical outages.
- Engineering Culture: Champions a blameless post-incident culture, operational readiness, continuous learning, and technical mentorship.
Benefits
- Competitive Package
- A dynamic and innovative team
- Collaborative, inclusive working environment
Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
539,450 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Free forever. No card. Under a minute.
Your match
How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.
Recommended for you based on this role
Similar stack
Same company
Hong Kong
DevOps Engineer
1 hour ago
≈ $16k – $45k per year (Estimated) • In office • 3+ years exp • Bengaluru
Python
Bash
DevOps
Terraform
GCP
Helm
Azure DevOps
Azure
CI/CD
GitOps
AWS
Docker
Kubernetes
SRE
Platform Engineering
IAM
Cybersecurity
ISO 27001
SOC 2
Apply
AI Engineer
1 min ago
≈ $22k – $59k per year (Estimated) • In office • 3+ years exp • Bengaluru
Python
Python
FastAPI
Databases
PostgreSQL
Weaviate
Chroma
Milvus
pgvector
Pinecone
AI/ML
LangGraph
AutoGen
LangChain
Claude
LlamaIndex
vLLM
Fine-tuning
Prompt Engineering
Multimodal AI
AI Agents
TensorRT
TensorRT-LLM
TGI
Llama
Mistral
TensorFlow
PyTorch
CrewAI
Gemini
LLM
RAG
Hallucination
Agentic Workflows
Multi-Agent Systems
DevOps
Rest API
GCP
Azure
CI/CD
Git
AWS
Docker
Kubernetes
Vector
Apply
Application Support Engineer
10 min ago
≈ $17k – $40k per year (Estimated) • In office • Full-Time • 5+ years exp • Hyderabad
DevOps
Incident Management
Apply
MLOps Engineer
21 min ago
In office • Bengaluru
Python
Rust
Bash
AI/ML
CUDA Toolkit
Multimodal AI
Computer Vision
PyTorch
MediaPipe
Ray
CUDA
cuDNN
DevOps
Terraform
GitHub Actions
CI/CD
AWS
Docker
Kubernetes
Amazon EKS
AWS Lambda
Amazon S3
IAM
Amazon ECS
Amazon CloudWatch
AWS Step Functions
Robotics
Localization
Visual-Inertial Odometry
Apply
Senior Site Reliability Engineer
3 days ago
≈ $21k – $52k per year (Estimated) • In office • Shenzhen
Databases
PostgreSQL
Redis
Apache Kafka
DevOps
Terraform
Datadog
Prometheus
AWS
Kubernetes
Grafana
Platform Engineering
Amazon EKS
Incident Management
Error Budget
SLI/SLO/SLA
Cybersecurity
PCI DSS
Apply
HR Specialist- UK (Chinese Speaking)
17 days ago
≈ $46k – $116k per year (Estimated) • In office • London
Apply
Apply
Apply
VP of Commercial & Partnership
22 days ago
≈ $107k – $257k per year (Estimated) • Equity • Remote • Singapore
Apply
Leader, Solutions Engineer (AI & Agentic Systems)
10 hours ago
≈ $63k – $150k per year (Estimated) • In office • Full-Time • 2+ years exp • Bachelor's Degree • Shanghai
AI/ML
Model Context Protocol
AI Agents
DevOps
Splunk
Cilium
Apply
Account Executive - Architecture
10 hours ago
$50k – $750k per year • In office • Full-Time • 7+ years exp • Bachelor's Degree • Hong Kong
Apply
≈ $52k – $116k per year (Estimated) • In office • Full-Time • 15+ years exp • Master's Degree • Hong Kong
Apply
Senior Business Developer
1 day ago
≈ $37k – $95k per year (Estimated) • Remote/Hybrid • Full-Time • Hong Kong
Marketing
HubSpot
LinkedIn
Apply
≈ $43k – $90k per year (Estimated) • In office • Full-Time • 5+ years exp • Sydney
Apply
This is one of many
539,450 more open roles from verified company boards, updated every day.

