582,867open jobs
25,570companies
81,102added this week
Browse all
Salary
$122k – $216k per year (Estimated)
Location
Remote (United States)
Seniority
Senior · 2+ years exp
Employment
Full-Time
Overview
Company
Impact
Profile match
Bitdeer Technologies Group is a global technology company specializing in Bitcoin mining, proprietary ASIC hardware manufacturing, and high-performance computing infrastructure. Headquartered in Singapore, the firm operates data center facilities across North America, Europe, and Asia to provide self-mining, cloud hash rate sharing, and colocation hosting services. By expanding into GPU-accelerated cloud platform capabilities, it delivers scalable computing solutions for both cryptocurrency networks and artificial intelligence workloads.

Bitdeer is a world-leading technology company for AI and Bitcoin mining infrastructure.

Bitdeer is committed to providing comprehensive Bitcoin mining solutions for its customers and building AI computational infrastructure to support the AI revolution. Bitdeer handles complex processes involved in computing such as equipment procurement, transport logistics, data center design and construction, equipment management, and daily operations. Bitdeer also offers advanced cloud capabilities to customers with high demand for artificial intelligence.

Headquartered in Singapore, Bitdeer has deployed data centers across multiple countries, including the United States, Norway, Bhutan, and Ethiopia.

To learn more, visit https://ir.bitdeer.com/

Position Overview

You run the control plane where AIOps meets tenants - where topology-aware scheduling, self-healing, and agent-driven remediation actually execute.

NeoCloud is building an AI-operated GPU cloud. Kubernetes is where all of that lands on real customer workloads: the topology-aware scheduler places jobs on the right NVLink domain, the operator drains and reschedules around predicted faults, and the tenant boundary is enforced against a Bare-Metal-as-a-Service backend. In this role you design, deploy, and operate that control plane - and you make sure the AIOps substrate can reach in and remediate without a human on the pager.

What you'll own

  • Production Kubernetes clusters optimized for GPU workloads at scale (100-10,000 GPUs).
  • Nvidia GPU operator, device plugin, MIG configuration, and GPU time-slicing policies.
  • Topology-aware scheduling: GPU locality, NVLink domain awareness, network rail affinity.
  • Custom Resource Definitions (CRDs) for GPU workload lifecycle management.
  • AI framework integrations: Slurm on K8S, Ray on K8S, Kubeflow.
  • Multi-tenant isolation: namespaces, network policies, resource quotas, RBAC, pod security standards.
  • Bare-Metal as a Service (BMaaS): automated provisioning, tenant onboarding, lifecycle, reclamation.
  • Terraform providers and modules for infrastructure-as-code across GPU clusters.
  • SLIs/SLOs for cluster availability, job completion rates, and provisioning latency.
  • Incident management: runbook automation, escalation, post-incident reviews.
  • Monitoring stack: Prometheus, Grafana, Alertmanager, PagerDuty.
  • GPU node failure handling: automated detection, drain/cordon/taint, workload rescheduling.

Feed the AIOps substrate

  • The remediation-actuator and workflow engine land here - you make the control plane safe for automated action.
  • Your CRDs are the schema the platform's predictors and remediators write against.
  • Every human intervention you do this quarter becomes an autonomous workflow next quarter.

What success looks like in year 1

  • Automated drain/reschedule around predicted GPU faults, at scale, without customer impact.
  • BMaaS live for external tenants with self-service onboarding.
  • Cluster availability and job-completion SLOs published and met.

Job Requirement:

  • 5+ years in Kubernetes operations, with at least 2 years managing GPU workloads on K8S
  • Deep understanding of Nvidia GPU operator, device plugin, and GPU scheduling in K8S
  • Experience with topology-aware scheduling and GPU-specific resource management
  • Hands-on experience building multi-tenant K8S platforms with strong isolation guarantees
  • Experience with bare-metal server provisioning and lifecycle automation (Ironic, MAAS, or custom)
  • Proficiency in Terraform, Helm, and GitOps workflows (ArgoCD/Flux)
  • Strong SRE background: SLI/SLO frameworks, incident management, capacity planning
  • Experience with Prometheus, Grafana, and alerting at scale
  • Strong programming skills in Go or Python for operator/CRD development
  • AIOps aptitude - you think of the K8S control plane as an execution surface for automated remediation, not just a scheduler. You've either wired an autoscaler/remediator loop into K8S or you can design one.
  • Runbook-as-code mindset - every SRE playbook you write should be executable by the platform.

--------------------------------------------------------------------

Bitdeer is committed to providing equal employment opportunities in accordance with country, state, and local laws. Bitdeer does not discriminate against employees or applicants based on conditions such as race, color, gender identity and/or expression, sexual orientation, marital and/or parental status, religion, political opinion, nationality, ethnic background or social origin, social status, disability, age, indigenous status, and union.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
582,867 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
San Jose
$44k – $89k per year (Estimated) • Remote/Hybrid • Full-Time • 2+ years exp • Bachelor's Degree • Kraków
Python
Go
Bash
DevOps
CI/CD
Git
Kubernetes
Platform Engineering
Argo Workflows
Apply
$115k – $158k per year • In office • Full-Time • 6+ years exp • Bachelor's Degree • Louisville • Columbia • New York
Python
SQL
Analytics
Tableau
Power BI
Alteryx
Apply
$145k – $261k per year • Equity • In office • Full-Time • 3+ years exp • Bachelor's Degree • San Jose • McLean • San Francisco • Lehi
Python
AI/ML
Prompt Engineering
Cybersecurity
OWASP Top 10
Apply
$122k – $148k per year • Remote/Hybrid • Full-Time • PhD • Leverkusen
Python
AI/ML
AI Agents
LLM
Anomaly Detection
Red Teaming
DevOps
CI/CD
AWS
Cybersecurity
Threat Modeling
Apply
In office • Moscow
Python
C
C++
C
FFmpeg
GCC
C++
Conan
CMake
VCPKG
Databases
PostgreSQL
SQLite
AI/ML
OpenCV
YOLO
ONNX
DevOps
Git
Docker
Robotics
GStreamer
Chips/EDA
PoC Library
Apply
$78k – $160k per year (Estimated) • Remote • Full-Time • San Jose
Python
Java
Rust
AI/ML
Ray
Anomaly Detection
InfiniBand
DevOps
GitHub Actions
Loki
OpenTelemetry
Prometheus
GitLab CI
SLURM
CI/CD
GitOps
Git
Kubernetes
Grafana
AIOps
SLI/SLO/SLA
Web3
Bitcoin
Apply
$122k – $216k per year (Estimated) • Remote • Full-Time • 2+ years exp • San Jose
AI/ML
KV Cache
DevOps
AIOps
HPC
Cybersecurity
Shuffle
Web3
Bitcoin
Apply
$120k – $212k per year (Estimated) • Remote • Full-Time • 5+ years exp • San Jose
Python
DevOps
Terraform
Ansible
AIOps
SLI/SLO/SLA
Web3
Bitcoin
Apply
$78k – $192k per year (Estimated) • In office • Full-Time • 5+ years exp • Singapore
AI/ML
NCCL
InfiniBand
DevOps
AIOps
HPC
Web3
Bitcoin
Apply
$75k – $186k per year (Estimated) • In office • Full-Time • 2+ years exp • Singapore
Python
AI/ML
Kubeflow
Ray
NVLink
DevOps
Terraform
Helm
PagerDuty
Prometheus
SLURM
GitOps
ArgoCD
Kubernetes
Grafana
Alertmanager
AIOps
Incident Management
SLI/SLO/SLA
HPC
Web3
Bitcoin
Apply
DMTS 3 hours ago
$89k – $249k per year (Estimated) • In office • Boise • San Jose
Apply
$112k – $215k per year • Equity • In office • Full-Time • 8+ years exp • San Francisco • San Jose
Apply
$135k – $234k per year • Equity • In office • Full-Time • 5+ years exp • San Jose • San Francisco • Seattle • Los Angeles • Lehi
Design
Figma
Canva
Apply
$70k – $90k per year • In office • 1+ year exp • Bachelor's Degree • San Jose
Analytics
Microsoft Excel
Apply
$144k – $205k per year • Remote/Hybrid • Full-Time • 8+ years exp • San Jose
Cybersecurity
Zscaler
Zero Trust
Management
Agile
Apply
See all jobs
This is one of many
582,867 more open roles from verified company boards, updated every day.