379,681open jobs
9,923companies
50,127added this week
Browse all
Location
In office (Singapore)
Overview
Company
Impact
Profile match
Headquartered in Jersey City, New Jersey, AvePoint is a global provider of enterprise cloud data management and governance solutions. The company specializes in data migration, backup protection, security compliance, and lifecycle management for platforms such as Microsoft 365, Salesforce, and Google Workspace. By automating administrative workflows and risk mitigation, it helps organizations secure their digital collaboration environments and optimize SaaS infrastructure.

We are seeking a skilled and passionate Engineer to join our team to build and operate a Whole-of-Government (WoG) runtime platform.

As a Site Reliability Engineer, you will be responsible for designing and operating GitLab, AWS and Kubernetes-based infrastructure and solutions that power our platform, to ensure the stability, scalability, and performance of our runtime platform.

Responsibilities:

As a Site Reliability Engineer, you will be responsible for:

Toil Reduction & Automation

  • Identify repetitive tasks and develop automation via CI/CD pipelines, ensuring integration with cross-functional teams to reduce manual intervention and improve operational efficiency.

Observability & System Health

  • Implement comprehensive observability solutions (logs, metrics, traces, alerts) around the four Golden Signals (latency, traffic, errors, saturation), and build automation for proactive system health assessments and self-remediation.

Production Support & Incident Management

  • Participate in on-call rotations, promptly respond to incidents to minimize MTTR, and conduct thorough post-incident reviews to implement preventive measures and improve system resilience.

Security & Compliance

  • Design and implement solutions that are secure and compliant by collaborating with dedicated security teams, conducting regular audits, and integrating advanced vulnerability scanning tools.

Maintenance, Optimisation & Performance

  • Identify and resolve performance bottlenecks and operational issues, define and track KPIs (e.g., MTTR, system uptime, cost efficiency), and drive ongoing optimisation efforts.

Strategic Customer Engagement

  • Act as a technical advisor for tenants, guiding them on containerization, and best practices for cloud-native deployments, and participating in strategic initiatives to enhance platform scalability and performance.

Knowledge Sharing & Documentation

  • Develop and maintain detailed playbooks, runbooks, and documentation to facilitate team-wide knowledge sharing, streamline incident response, and ensure that critical processes are well understood across the team.

Continuous Learning & Innovation

  • Stay current with the latest AWS, Kubernetes, and industry developments, and proactively recommend improvements and innovative solutions to maintain a competitive and reliable platform.

Requirements:

  • Bachelor's degree or Diploma in Computer Science, Engineering, or a related field (or equivalent experience).
  • Proven experience as a Site Reliability Engineer or similar role, with a strong background in containerization, orchestration, and cloud-native technologies.
  • Proven ability to troubleshoot and resolve complex technical issues in containerized applications.
  • Demonstrated experience with incident management, including post-incident reviews and continuous improvement.
  • Strong documentation skills and experience in knowledge sharing across teams.
  • Deep understanding of AWS, Kubernetes (including AWS EKS), and operational best practices, with familiarity in multi-cloud or hybrid environments.
  • Solid grasp of networking, security, and storage in both AWS and Kubernetes contexts.
  • Experience integrating Kubernetes with AWS cloud technologies (e.g., Secrets Manager, Load Balancers) and using infrastructure-as-code (Terraform or similar).
  • Hands-on experience with containerization tools (Kubernetes, Kustomize, Helm) and automation scripting (Go, Python, Bash, or equivalent).
  • Ability to write and maintain automated tests or conduct thorough manual testing for automation scripts, ensuring the reliability and effectiveness of automated solutions.
  • Familiarity with CI/CD tools (GitLab CI/CD, ArgoCD) and version control systems (Git).
  • Experience with observability/monitoring tools (Prometheus, Grafana, ELK Stack) and defining SLOs and Error Budgets.
  • Certifications such as Certified Kubernetes Administrator (CKA) or Certified Kubernetes Application Developer (CKAD) are a plus.
  • Experience with developing Kubernetes operators using Go, service mesh technologies, and Chaos Engineering is a plus.

Soft skills:

  • Proactive in identifying problems and recommending strategic solutions.
  • Excellent problem-solving skills with a robust analytical mindset.
  • Clear, concise, and effective communication skills; adept at collaborating across crossfunctional teams, including development, security, and customer-facing groups.
  • Ability to remain calm and effective under pressure, especially during incident response.
  • Adaptability to rapid change with a continuous learning mindset, sharing knowledge to foster team growth.
  • Customer-focused with the ability to translate technical insights into understandable, actionable guidance.
  • Leadership and mentoring capabilities, contributing to the development of a resilient and collaborative team environment are a plus.

Any personal data you share with us during the application process will be processed strictly in compliance with applicable data protection laws and our Privacy Notice.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
379,681 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
Singapore
$74k – $209k per year (Estimated) • Remote • 5+ years exp • Bachelor's Degree • Singapore
JavaScript
TypeScript
C#
Node JS
C#
ASP.NET Core
Node JS
Nest.JS
Databases
Microsoft Fabric
AI/ML
LLM
Prompt Engineering
Frontend
Angular
React.js
Vue.js
DevOps
AWS
Azure
GCP
Git
Analytics
Power BI
Apply
$120k – $140k per year • In office • 3+ years exp
Java
Python
SQL
Python
FastAPI
Databases
PostgreSQL
DevOps
AWS
CI/CD
Rest API
QA
Playwright
Postman
Pytest
Rest-Assured
Apply
$15k – $38k per year (Estimated) • In office • Full-Time • 4+ years exp • Tashkent
Databases
Apache Kafka
PostgreSQL
DevOps
Ansible
ArgoCD
CI/CD
Docker
GitLab
GitLab CI
GitOps
Grafana
Helm
Kubernetes
Prometheus
Terraform
Ubuntu
Apply
$15k – $38k per year (Estimated) • In office • 4+ years exp • Tashkent
Databases
Apache Kafka
PostgreSQL
DevOps
Ansible
ArgoCD
CI/CD
Docker
GitLab
GitLab CI
GitOps
Grafana
Helm
Kubernetes
Prometheus
Terraform
Ubuntu
Apply
Backend Developer 1 day ago
$20k – $84k per year (Estimated) • Remote • Contractor
Go
Python
SQL
Python
Django
FastAPI
Databases
Apache Kafka
Databricks
DevOps
Azure
Azure AKS
Bicep
CI/CD
Docker
GCP
Git
Kubernetes
Terraform
Apply
In office • Bachelor's Degree • Johor Bahru
JavaScript
SQL
Databases
MySQL
PostgreSQL
Frontend
React.js
DevOps
CI/CD
Apply
In office • Bachelor's Degree • Kuala Lumpur
JavaScript
SQL
Databases
MySQL
PostgreSQL
Frontend
React.js
DevOps
CI/CD
Apply
$190k – $220k per year • Equity • Remote/Hybrid • 5+ years exp • Bachelor's Degree • Jersey City
AI/ML
LLM
Marketing
Salesforce
Apply
$8k – $21k per year (Estimated) • Remote/Hybrid • 1+ year exp • Bachelor's Degree • Manila
SQL
DevOps
Azure
Windows Server
Apply
$180k – $230k per year • Equity • In office • 5+ years exp • Arlington
C#
Python
TypeScript
Databases
Chroma
Milvus
Pinecone
Weaviate
AI/ML
AI Agents
Anthropic
AWS Bedrock
EU AI Act
Function Calling
ISO 42001
LangChain
LLM
Model Context Protocol
NIST AI RMF
OpenAI
RAG
Semantic Kernel
DevOps
AWS
Azure
GCP
Marketing
Salesforce
Apply
In office • Full-Time • Master's Degree • Singapore
C++
Python
AI/ML
Multimodal AI
Reinforcement Learning
Robotics
MoveIt
Perception
Reinforcement Learning
ROS
Apply
In office • Full-Time • Bachelor's Degree • Singapore
AI/ML
Automatic1111
ComfyUI
Gradio
LocalAI
Stable Diffusion
Apply
$55k – $121k per year (Estimated) • In office • Full-Time • 4+ years exp • Bachelor's Degree • Singapore
Apply
$75k – $129k per year (Estimated) • In office • Contractor • Singapore
AI/ML
ChatGPT
Copilot
Apply
$60k – $153k per year (Estimated) • In office • Full-Time • 1+ year exp • Singapore
Apply
See all jobs
This is one of many
379,681 more open roles from verified company boards, updated every day.