368,530open jobs
9,432companies
50,439added this week
Browse all
Location
In office (Dubai)
Seniority
Middle · 4+ years exp
Employment
Full-Time
Overview
Company
Impact
Profile match
Fuse Energy is a British household energy supplier headquartered in London, founded by former Revolut executives. It was the first new domestic energy supplier in Great Britain since the 2021 energy crisis and operates its own renewable generation assets.

Fuse Energy is a forward-thinking renewable energy startup on a mission to deliver a terawatt of renewable energy - fast. We're combining first-principles thinking with cutting-edge technology to build a radically better energy system. We raised $210M from top-tier investors including Multicoin, Balderton, Lakestar, Accel, Creandum, Lowercarbon, Ribbit, Box Group and strategic angels like Nico Rosberg, the Co-Founder of Solana and GPs behind Meta, Revolut, Spotify, Uber and more.

As data centres become one of the largest and fastest-growing sources of electricity demand, Fuse is expanding into high-performance compute infrastructure that sits at the intersection of energy and AI. We're building the GPU/CUDA performance layer and the inference serving layer at the same time, from scratch - and we're looking for the founding engineer to own the latter.

We're looking for a Founding AI Inference Engineer to define and build how Fuse serves AI inference workloads at scale, reporting directly to the CTO. Where our CUDA and GPU engineering hires own kernel-level and hardware performance, this role owns the layer above it: how models actually get served, scaled, and delivered against committed performance targets.

The Opportunity

Fuse is seeing significant demand for data centre capacity across the markets we operate in, primarily for inference. Few companies in the world can pair real power delivery with real compute the way Fuse can, which puts inference serving at the heart of how we turn that advantage into the best offering in the market. That's this role.

Responsibilities

  • Define Fuse's inference serving strategy and architecture from first principles

  • Design and build the serving stack: request routing, batching, scheduling, and autoscaling for high-throughput, latency-sensitive inference workloads

  • Own model-level optimisation strategy for serving - deciding where and how to apply quantisation, distillation, speculative decoding, and similar techniques to improve throughput and cost per token, partnering with the CUDA/GPU engineers

  • Make the core software architecture calls on serving frameworks and orchestration (e.g. vLLM, TensorRT-LLM, SGLang, Triton Inference Server, or equivalents)

  • Translate throughput, latency, and uptime commitments into concrete technical specifications and serving capacity plans

  • Act as a direct technical owner of inference performance and reliability

  • Work closely with the CUDA and GPU engineering teams to ensure custom kernels and hardware performance work are integrated cleanly into the serving layer

  • Set the standards, tooling, and benchmarks this function will run on as it grows

Requirements

  • 4+ years of experience building or operating large-scale inference serving systems, or equivalent strong project/industry experience

  • Deep, hands-on experience with inference serving frameworks and the techniques used to optimise them (batching, KV-cache management, quantisation, speculative decoding)

  • Strong systems thinking - able to reason about the full path from incoming request to served response across a large cluster

  • Comfortable working directly with GPU/CUDA engineers to integrate low-level performance work into a serving system

  • A track record of making high-stakes architecture calls and owning the outcome

  • Comfort operating without a playbook - this is a founding role shaping a new function around architecture that's still early-stage, not joining an established one

Nice to Have

  • Experience with Triton or custom ML inference/training frameworks

  • Experience with autoscaling or capacity planning for large-scale inference workloads

  • Exposure to multi-tenant serving or SLA-driven infrastructure

  • Background at a hyperscaler, frontier AI lab, or large-scale distributed inference system

  • Familiarity with Kubernetes/Slurm for cluster orchestration

  • Interest or experience in energy markets, grid systems, or sustainability-focused compute

Benefits

  • Competitive salary and an equity sign-on bonus

  • Biannual bonus scheme

  • Fully expensed tech to match your needs

  • Breakfast and dinner allowance for office based employees

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
368,530 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
Dubai
$108k – $242k per year (Estimated) • Remote • Full-Time • 5+ years exp
C#
SQL
C#
.NET
Databases
MS SQL
RabbitMQ
AI/ML
ChatGPT
Claude
Claude Code
Copilot
Cursor
LLM
RAG
Model Context Protocol
DevOps
CI/CD
Git
GitHub
Apply
$40k – $107k per year (Estimated) • Remote • 5+ years exp • Tbilisi
C#
SQL
C#
.NET
Databases
MS SQL
RabbitMQ
AI/ML
ChatGPT
Claude
Claude Code
Copilot
Cursor
LLM
RAG
Model Context Protocol
DevOps
CI/CD
Git
GitHub
Apply
Remote • Full-Time
JavaScript
Python
TypeScript
AI/ML
AI Agents
Function Calling
LLM
Prompt Engineering
Structured Outputs
Management
n8n
Zapier
Apply
$27k – $110k per year (Estimated) • Remote • Full-Time
JavaScript
Python
TypeScript
AI/ML
AI Agents
Function Calling
LLM
Prompt Engineering
Structured Outputs
Management
n8n
Zapier
Apply
$27k – $69k per year (Estimated) • Remote • Full-Time
JavaScript
Node JS
Python
TypeScript
AI/ML
Embeddings
LLM
RAG
LLM Guardrails
AI Agents
Function Calling
Model Context Protocol
Frontend
Angular
DevOps
AWS
Azure
CI/CD
Docker
GCP
Vector
Apply
$78k – $160k per year (Estimated) • In office • Full-Time • Bachelor's Degree • London
Python
DevOps
AWS
Git
IAM
Cybersecurity
Crowdstrike
FortiGate
ISO 27001
SentinelOne
SOC 2
Apply
$106k – $210k per year (Estimated) • In office • Full-Time • Bachelor's Degree
Apply
$91k – $179k per year (Estimated) • Remote • Full-Time
C++
Python
DevOps
AWS
Azure
CI/CD
GCP
RTOS
WebSockets
Cybersecurity
GDPR
ISO 27001
Web3
Solana
IoT
CoAP
FreeRTOS
MQTT
Zigbee
Apply
$75k – $180k per year (Estimated) • In office • Full-Time • Adelaide
Apply
HPC Network Engineer 1 month ago
$68k – $161k per year (Estimated) • In office • Full-Time • London
Python
JavaScript
AI/ML
InfiniBand
Frontend
Bootstrap
DevOps
Ansible
Configuration Management
Datadog
Grafana
Prometheus
HPC
Cybersecurity
FortiGate
Web3
Solana
Apply
Remote/Hybrid • Full-Time • Bachelor's Degree • Dubai
Apply
Remote/Hybrid • Full-Time • 15+ years exp • Master's Degree • Dubai
Apply
In office • Full-Time • 5+ years exp • Bachelor's Degree • Dubai
Apply
In office • Dubai
Apply
In office • 1+ year exp • Dubai
C++
JavaScript
Python
TypeScript
Apply
See all jobs
This is one of many
368,530 more open roles from verified company boards, updated every day.