588,947open jobs
26,464companies
82,741added this week
Browse all
Salary
$76k – $186k per year (Estimated)
Location
In office (Singapore)
Seniority
Senior · 2+ years exp
Employment
Full-Time
Overview
Company
Impact
Profile match
Bitdeer Technologies Group is a global technology company specializing in Bitcoin mining, proprietary ASIC hardware manufacturing, and high-performance computing infrastructure. Headquartered in Singapore, the firm operates data center facilities across North America, Europe, and Asia to provide self-mining, cloud hash rate sharing, and colocation hosting services. By expanding into GPU-accelerated cloud platform capabilities, it delivers scalable computing solutions for both cryptocurrency networks and artificial intelligence workloads.

About Bitdeer:

Bitdeer is a world-leading technology company for Bitcoin mining and AI cloud.

Bitdeer is committed to providing comprehensive Bitcoin mining solutions for its customers. Apart from designing industry-leading ASIC chips and manufacturing mining rigs, the Group handles complex processes involved in computing across the value chain. This includes equipment procurement, transport logistics, datacenter design and construction, equipment management, and network and facility operations. Bitdeer also offers advanced cloud capabilities to customers with a high demand for artificial intelligence.

Headquartered in Singapore, Bitdeer operates globally with a diversified 3 GW energy portfolio, and deploys Bitcoin mining and HPC datacenters in the United States, Bhutan, Norway, Canada, Malaysia, and Ethiopia.

What you will be responsible for:

What you'll own

  • End-to-end reliability of the customer-facing GPU cloud service - availability, job completion, provisioning latency, and tenant experience.
  • Production Kubernetes clusters optimized for GPU workloads at scale (100-10,000 GPUs) as the runtime substrate for customer workloads.
  • Nvidia GPU operator, device plugin, MIG configuration, GPU time-slicing, and multi-tenant GPU allocation policies.
  • Topology-aware scheduling: GPU locality, NVLink domain awareness, network rail affinity - placing customer jobs on the right hardware.
  • Customer & tenant lifecycle: onboarding, quota management, isolation enforcement (namespaces, network policies, RBAC, resource quotas, pod security), and offboarding/reclamation.
  • Bare-Metal-as-a-Service (BMaaS): automated provisioning, tenant handoff, lifecycle, and reclamation.
  • SLIs/SLOs/SLAs for the customer cloud service: cluster availability, job completion rates, provisioning latency, API availability.
  • Incident management with customer communication: runbook automation, escalation, customer-facing status updates, and post-incident reviews.
  • Monitoring & observability stack: Prometheus, Grafana, Alertmanager, PagerDuty - tenant-aware dashboards and alerting.
  • GPU node failure handling: automated detection, drain/cordon/taint, and workload rescheduling - minimizing customer-visible impact.
  • Infrastructure-as-code: Terraform providers/modules, Helm, and GitOps (ArgoCD/Flux) across GPU clusters.
  • Customer-facing operational readiness: service documentation, tenant runbooks, capacity planning, and support tiering.

Customer-facing ownership

  • You are accountable for the customer's reliability experience - when a tenant's job fails or a node drops, you own the detection, remediation, and communication loop.
  • Define and publish customer-facing SLAs/SLOs and drive error-budget-based prioritization between feature work and reliability.
  • Partner with customer success / support to close the feedback loop between customer-reported issues and systemic improvements.
  • Build self-service observability that lets customers answer their own questions - status, quota, job health - reducing support load.

Feed the AIOps substrate

  • The remediation-actuator and workflow engine land here - you make the control plane safe for automated action.
  • Your CRDs and runbooks are the schema the platform's predictors and remediators write against.
  • Every human intervention you do this quarter becomes an autonomous workflow next quarter - turning customer-impacting incidents into self-healing events.

What success looks like in year 1

  • Customer-facing GPU cloud service SLAs published and met - availability, job completion, provisioning latency.
  • Automated drain/reschedule around predicted GPU faults, at scale, without customer-visible impact.
  • BMaaS live for external tenants with self-service onboarding.
  • MTTD and MTTR for customer-impacting incidents reduced through automation.
  • Tenant self-service observability live - customers can see their own job health, quota, and status.

    How you will stand out:

    • 5+ years in SRE / cloud operations, with at least 2 years operating GPU workloads at scale.
    • Deep understanding of Kubernetes operations and GPU workload management (Nvidia GPU operator, device plugin, MIG, time-slicing, GPU scheduling).
    • Experience with topology-aware scheduling and GPU-specific resource management.
    • Hands-on experience building multi-tenant cloud platforms with strong isolation guarantees.
    • Customer-facing cloud service experience - defining and operating against customer SLAs/SLOs, handling tenant incidents and communications.
    • Experience with bare-metal server provisioning and lifecycle automation (Ironic, MAAS, or custom).
    • Proficiency in Terraform, Helm, and GitOps workflows (ArgoCD/Flux).
    • Strong SRE background: SLI/SLO/SLA frameworks, error budgets, incident management, capacity planning.
    • Experience with Prometheus, Grafana, and alerting at scale.
    • Strong programming skills in Go or Python for automation / operator development.
    • AIOps aptitude - you view the control plane as an execution surface for automated remediation, not just a scheduler.
    • Runbook-as-code mindset - every SRE playbook you write should be executable by the platform.

    What you will experience working with us:

    • A culture that values authenticity and diversity of thoughts and backgrounds;
    • An inclusive and respectable environment with open workspaces and exciting start-up spirit;
    • Fast-growing company with the chance to network with industrial pioneers and enthusiasts;
    • Ability to contribute directly and make an impact on the future of the digital asset industry;
    • Involvement in new projects, developing processes/systems;
    • Personal accountability, autonomy, fast growth, and learning opportunities;
    • Attractive welfare benefits and developmental opportunities such as training and mentoring.

    --------------------------------------------------------------------

    Bitdeer is committed to providing equal employment opportunities in accordance with country, state, and local laws. Bitdeer does not discriminate against employees or applicants based on conditions such as race, colour, gender identity and/or expression, sexual orientation, marital and/or parental status, religion, political opinion, nationality, ethnic background or social origin, social status, disability, age, indigenous status, and union.

    Free account
    Stop reading job ads. Get the ones that fit.
    One free account turns this page into a shortlist built around your stack, your level and your pay.
    Match on every job. Stack, seniority, pay and location, scored against your profile.
    588,947 open roles. Read straight off company career pages, refreshed every day.
    Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
    3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
    Create a free account Continue with Google
    Free forever. No card. Under a minute.

    Your match

    How well do you fit this role?
    Two answers are enough for a real match. No account needed.
    Check my fit
    Answers stay in this browser until you create an account.

    Recommended for you based on this role

    Similar stack
    Same company
    Singapore
    Cloud Engineer 1 day ago
    $55k – $98k per year (Estimated) • Remote/Hybrid • Bachelor's Degree • Wrocław
    Python
    PHP
    Bash
    PHP
    WooCommerce
    DevOps
    Terraform
    AWS
    Immutable Infrastructure
    Configuration Management
    Management
    Agile
    Apply
    $21k – $52k per year (Estimated) • Remote/Hybrid • Full-Time • 5+ years exp • Bengaluru
    Python
    AI/ML
    Model Context Protocol
    DevOps
    Terraform
    GCP
    Istio
    OpenTelemetry
    Datadog
    FluxCD
    Azure
    CI/CD
    GitOps
    ArgoCD
    AWS
    Kubernetes
    Platform Engineering
    Service Mesh
    AIOps
    Incident Management
    Apply
    In office • Full-Time • Kwidzyn
    Python
    C#
    Apply
    $23k – $49k per year (Estimated) • In office • 8+ years exp • Bengaluru
    JavaScript
    TypeScript
    Databases
    Apache Kafka
    Frontend
    Angular
    DevOps
    Terraform
    CI/CD
    Docker
    Kubernetes
    Management
    Agile
    QA
    Selenium
    Playwright
    Postman
    Rest-Assured
    SoapUI
    Apply
    $24k – $69k per year (Estimated) • In office • 6+ years exp • Bachelor's Degree • Bengaluru
    Python
    Go
    JavaScript
    Java
    TypeScript
    Node JS
    Databases
    MySQL
    PostgreSQL
    Frontend
    Vue.js
    GraphQL
    Angular
    React.js
    DevOps
    Rest API
    GCP
    Azure
    CI/CD
    AWS
    Apply
    $78k – $192k per year (Estimated) • In office • Full-Time • 5+ years exp • Singapore
    AI/ML
    NCCL
    InfiniBand
    DevOps
    AIOps
    HPC
    Web3
    Bitcoin
    Apply
    $75k – $186k per year (Estimated) • In office • Full-Time • 2+ years exp • Singapore
    Python
    AI/ML
    Kubeflow
    Ray
    NVLink
    DevOps
    Terraform
    Helm
    PagerDuty
    Prometheus
    SLURM
    GitOps
    ArgoCD
    Kubernetes
    Grafana
    Alertmanager
    AIOps
    Incident Management
    SLI/SLO/SLA
    HPC
    Web3
    Bitcoin
    Apply
    $75k – $184k per year (Estimated) • In office • Full-Time • 5+ years exp
    Python
    DevOps
    Terraform
    Ansible
    AIOps
    SLI/SLO/SLA
    HPC
    Web3
    Bitcoin
    Apply
    $76k – $187k per year (Estimated) • In office • Full-Time • 2+ years exp • Singapore
    AI/ML
    KV Cache
    DevOps
    AIOps
    HPC
    Cybersecurity
    Shuffle
    Web3
    Bitcoin
    Apply
    $171k – $317k per year (Estimated) • In office • Full-Time • Austin
    Web3
    Bitcoin
    Apply
    $68k – $154k per year (Estimated) • In office • Full-Time • 2+ years exp • Bachelor's Degree • Singapore
    Apply
    GTM Engineer 8 hours ago
    $100k – $150k per year • In office • Full-Time • 2+ years exp • Singapore
    AI/ML
    Reinforcement Learning
    Post-training
    Apply
    $150k – $250k per year • In office • Full-Time • 2+ years exp • Singapore
    Python
    AI/ML
    Reinforcement Learning
    DevOps
    Docker
    Apply
    $150k – $250k per year • In office • Full-Time • 2+ years exp • Singapore
    Python
    AI/ML
    Reinforcement Learning
    AI Agents
    DevOps
    Docker
    Apply
    $150k – $250k per year • In office • Full-Time • 2+ years exp • Singapore
    Python
    AI/ML
    Reinforcement Learning
    LLM
    Post-training
    DevOps
    Docker
    Apply
    See all jobs
    This is one of many
    588,947 more open roles from verified company boards, updated every day.