412,170open jobs
14,281companies
71,440added this week
Browse all
Salary
$95k – $234k per year (Estimated)
Location
In office (Singapore)
Seniority
Senior
Employment
Full-Time
Overview
Company
Impact
Profile match
Bitdeer Technologies Group is a global technology company specializing in Bitcoin mining, proprietary ASIC hardware manufacturing, and high-performance computing infrastructure. Headquartered in Singapore, the firm operates data center facilities across North America, Europe, and Asia to provide self-mining, cloud hash rate sharing, and colocation hosting services. By expanding into GPU-accelerated cloud platform capabilities, it delivers scalable computing solutions for both cryptocurrency networks and artificial intelligence workloads.

About Bitdeer:

Bitdeer is a world-leading technology company for Bitcoin mining and AI cloud.

Bitdeer is committed to providing comprehensive Bitcoin mining solutions for its customers. Apart from designing industry-leading ASIC chips and manufacturing mining rigs, the Group handles complex processes involved in computing across the value chain. This includes equipment procurement, transport logistics, datacenter design and construction, equipment management, and network and facility operations. Bitdeer also offers advanced cloud capabilities to customers with a high demand for artificial intelligence.

Headquartered in Singapore, Bitdeer operates globally with a diversified 3 GW energy portfolio, and deploys Bitcoin mining and HPC datacenters in the United States, Bhutan, Norway, Canada, Malaysia, and Ethiopia.

About the team:

We are seeking a Senior Inference Runtime Engineer to own the performance-critical serving layer of the MaaS platform. This role focuses on making self-hosted LLMs faster, cheaper, and more stable by optimizing the runtime stack behind OpenAI- and Anthropic-compatible APIs.

What you will be responsible for:

  • Optimize prefill/decode scheduling, continuous batching, KV cache behavior, speculative decoding, long-context serving, and streaming smoothness.
  • Tune and operate vLLM/Dynamo/SGLang/TensorRT-LLM-style runtimes for model-specific latency, throughput, GPU utilization, and cost efficiency.
  • Profile bottlenecks across GPU memory, HBM bandwidth, NCCL/network, tokenizer, frontend/proxy, and model worker paths.
  • Lead high-value model onboarding, including runtime selection, tensor/pipeline parallelism, quantization, context length, and rollback strategy.
  • Define runtime playbooks and safe defaults for reasoning, tool calling, multimodal, prompt cache, and provider-specific parameters.
  • Partner with SRE and performance/evaluation engineers to turn benchmark findings into production runtime improvements.

How you will stand out:

  • 6+ years of systems, ML infrastructure, or high-performance backend engineering experience.
  • Hands-on experience with LLM serving runtimes such as vLLM, Dynamo, SGLang, TensorRT-LLM, TGI, or Triton.
  • Strong understanding of GPU memory, CUDA/NCCL basics, KV cache, batching, streaming, and distributed inference tradeoffs.
  • Proficient in Go or Python and comfortable reading runtime source code, profiling traces, and production metrics.
  • Experience operating production inference services with strict latency, availability, and cost targets.
  • Able to translate low-level performance work into customer-visible reliability, latency, and margin improvements.

What you will experience working with us:

  • A culture that values authenticity and diversity of thoughts and backgrounds;
  • An inclusive and respectable environment with open workspaces and exciting start-up spirit;
  • Fast-growing company with the chance to network with industrial pioneers and enthusiasts;
  • Ability to contribute directly and make an impact on the future of the digital asset industry;
  • Involvement in new projects, developing processes/systems;
  • Personal accountability, autonomy, fast growth, and learning opportunities;
  • Attractive welfare benefits and developmental opportunities such as training and mentoring.

--------------------------------------------------------------------

Bitdeer is committed to providing equal employment opportunities in accordance with country, state, and local laws. Bitdeer does not discriminate against employees or applicants based on conditions such as race, colour, gender identity and/or expression, sexual orientation, marital and/or parental status, religion, political opinion, nationality, ethnic background or social origin, social status, disability, age, indigenous status, and union.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
412,170 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
Singapore
$33k – $78k per year (Estimated) • In office • Moscow
Python
Go
TypeScript
Databases
PostgreSQL
Redis
Apache Kafka
Kafka
AI/ML
Model Context Protocol
LLM
RAG
DevOps
gRPC
OpenTelemetry
GitLab CI
CI/CD
Docker
Kubernetes
Grafana
GitLab
Amazon S3
Apply
$13k – $36k per year (Estimated) • In office • Almaty
Python
SQL
AI/ML
NLP
Mistral
TensorFlow
PyTorch
LLM
DevOps
Git
Docker
Apply
$11k – $23k per year (Estimated) • In office • 1+ year exp • Bachelor's Degree • Almaty
Python
Apply
FEA Analyst 5 hours ago
$14k – $29k per year (Estimated) • Remote/Hybrid • Pune
Python
Apply
Work-study contract 5 hours ago
$10k – $27k per year (Estimated) • In office • Contractor • Clermont-Ferrand
Python
Apply
$53k – $123k per year (Estimated) • In office • Full-Time • 3+ years exp • Singapore
DevOps
HPC
Web3
Bitcoin
Design
Axure RP
Apply
$90k – $223k per year (Estimated) • In office • Full-Time • 6+ years exp • Bachelor's Degree • Singapore
Python
TypeScript
AI/ML
LangGraph
LangChain
Claude
Model Context Protocol
AI Agents
CrewAI
LLM
RAG
OpenAI
Anthropic
LLM Guardrails
Multi-Agent Systems
DevOps
GCP
GitHub Actions
OpenTelemetry
Prometheus
Azure
CI/CD
Jenkins
AWS
Kubernetes
GitHub
HPC
Web3
Bitcoin
Apply
$17k – $42k per year (Estimated) • In office • Full-Time • 5+ years exp • Bachelor's Degree
Python
Go
AI/ML
vLLM
Triton Inference Server
TGI
Triton
DevOps
Terraform
Ansible
GCP
Helm
Prometheus
Azure
CI/CD
AWS
Docker
Kubernetes
Grafana
Platform Engineering
Self-Healing
Incident Management
HPC
Cybersecurity
ISO 27001
SOC 2
Zero Trust
Web3
Bitcoin
Apply
$95k – $241k per year (Estimated) • In office • Full-Time • 10+ years exp • Bachelor's Degree • Singapore
Verilog
SystemVerilog
DevOps
HPC
Web3
Bitcoin
Chips/EDA
UVM
Apply
$23k – $58k per year (Estimated) • In office • Full-Time • Bachelor's Degree
Verilog
Perl
DevOps
HPC
Web3
Bitcoin
Apply
$121k – $311k per year (Estimated) • In office • 10+ years exp • Singapore
Management
Stripe
Marketing
LinkedIn
Apply
In office • Singapore
Management
Outlook
Apply
$33k – $68k per year (Estimated) • In office • Full-Time • Singapore
Apply
Integration Manager 3 hours ago
In office • Full-Time • Singapore • Hong Kong
JavaScript
Frontend
Parcel
Apply
$88k – $205k per year (Estimated) • In office • Full-Time • 2+ years exp • Singapore
DevOps
SLI/SLO/SLA
Apply
See all jobs
This is one of many
412,170 more open roles from verified company boards, updated every day.