368,746open jobs
9,444companies
47,506added this week
Browse all
Salary
$104k – $217k per year (Estimated)
Location
In office (London)
Employment
Full-Time
Overview
Company
Impact
Profile match
Fuse Energy is a British household energy supplier headquartered in London, founded by former Revolut executives. It was the first new domestic energy supplier in Great Britain since the 2021 energy crisis and operates its own renewable generation assets.

Fuse Energy is an energy startup on a mission to make energy abundant and affordable, fast. We combine first-principles thinking with cutting-edge technology to build a radically better energy system.

We've raised over $200M from top-tier investors including Balderton, Lakestar, Accel, Creandum, Lowercarbon, Ribbit, 20VC, Hummingbird and Collaborative Fund, alongside strategic angels including Nico Rosberg and GPs behind Meta, Revolut, Spotify and Uber.

We're building a fully integrated energy company: developing our own solar, batteries and other generation projects, building our own hardware, improving and developing grid infrastructure, trading power in real time, using AI across the business, and installing distributed energy in homes. By selling directly to consumers we cut out the middleman, lower costs and pass the savings on to our customers.

As data centres become one of the largest and fastest-growing sources of electricity demand, Fuse is expanding into high-performance compute infrastructure at the intersection of energy and AI. We're building the GPU/CUDA performance layer and the inference serving layer at the same time, from scratch, and we're looking for the founding engineer to own the latter. Reporting directly to the CTO, you'll own the layer above kernels and hardware: how models actually get served, scaled and delivered against committed performance targets. Few companies can pair real power delivery with real compute the way Fuse can, which puts inference serving at the heart of our offering.

Responsibilities

  • Define Fuse's inference serving strategy and architecture from first principles
  • Design and build the serving stack: request routing, batching, scheduling and autoscaling for high-throughput, latency-sensitive inference workloads
  • Own model-level optimisation strategy for serving, deciding where and how to apply quantisation, distillation, speculative decoding and similar techniques to improve throughput and cost per token, partnering with the CUDA/GPU engineers
  • Make the core software architecture calls on serving frameworks and orchestration (e.g. vLLM, TensorRT-LLM, SGLang, Triton Inference Server or equivalents)
  • Translate throughput, latency and uptime commitments into concrete technical specifications and serving capacity plans
  • Act as direct technical owner of inference performance and reliability
  • Work closely with the CUDA and GPU engineering teams to integrate custom kernels and hardware performance work cleanly into the serving layer
  • Set the standards, tooling and benchmarks this function will run on as it grows

Requirements

  • 4+ years building or operating large-scale inference serving systems, or equivalent strong project/industry experience
  • Deep, hands-on experience with inference serving frameworks and the techniques used to optimise them (batching, KV-cache management, quantisation, speculative decoding)
  • Strong systems thinking, able to reason about the full path from incoming request to served response across a large cluster
  • Comfortable working directly with GPU/CUDA engineers to integrate low-level performance work into a serving system
  • A track record of making high-stakes architecture calls and owning the outcome
  • Comfort operating without a playbook: this is a founding role shaping a new function, not joining an established one
  • Bonus: Triton or custom ML inference/training frameworks; autoscaling or capacity planning for large-scale inference; multi-tenant serving or SLA-driven infrastructure; background at a hyperscaler, frontier AI lab or large-scale distributed inference system; Kubernetes/Slurm; interest in energy markets, grid systems or sustainability-focused compute

Benefits

  • Competitive salary and eligibility for equity
  • Biannual bonus scheme
  • Fully expensed tech to match your needs
  • Private health insurance
  • Breakfast and dinner allowance for office-based employees

As we hire globally, benefits vary by location.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
368,746 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
London
$195k – $264k per year • In office • Full-Time • 15+ years exp • Master's Degree • United States
Python
AI/ML
Amazon SageMaker
Keras
Kubeflow
MLFlow
PyTorch
Scikit-learn
TensorFlow
Vertex AI
XGBoost
DevOps
AWS
Azure
CI/CD
CloudFormation
Docker
GCP
Kubernetes
Terraform
Cybersecurity
FedRAMP
NIST 800-53
Apply
Fullstack QA 10 hours ago
$14k – $32k per year (Estimated) • In office • 3+ years exp • Moscow
JavaScript
Python
SQL
TypeScript
Databases
Apache Kafka
PostgreSQL
DevOps
GitLab
Jenkins
Kubernetes
Management
Bitrix24
Confluence
Jira
Apply
$62k – $142k per year (Estimated) • In office • Full-Time • 6+ years exp • PhD • Madrid
Apex
JavaScript
Python
TypeScript
Databases
Databricks
Google BigQuery
Snowflake
AI/ML
Agentforce
AI Agents
Claude
Cursor
LangChain
LlamaIndex
LLM
Prompt Engineering
Marketing
Salesforce
Apply
$22k – $38k per year (net) • In office • Full-Time • 3+ years exp • Astana
Python
TypeScript
JavaScript
Python
Alembic
Celery
FastAPI
Litestar
Pydantic
SQLAlchemy
structlog
Databases
ClickHouse
ElasticSearch
MinIO
NATS
Neo4j
PostgreSQL
Redis
AI/ML
LLM
OpenAI
Pydantic AI
Frontend
Mantine
React Hook Form
React.js
Recharts
shadcn/ui
Tailwind CSS
TanStack Router
TanStack Table
Vite
Zustand
Radix UI
DevOps
Amazon S3
Ansible
Docker
Git
GitLab
GitLab CI
Grafana
gRPC
Terraform
QA
Playwright
Pytest
Apply
$18k – $28k per year (net) • In office • Saint Petersburg
AI/ML
ChatGPT
Claude
LLM
Marketing
AmoCRM
Apply
$78k – $160k per year (Estimated) • In office • Full-Time • Bachelor's Degree • London
Python
DevOps
AWS
Git
IAM
Cybersecurity
Crowdstrike
FortiGate
ISO 27001
SentinelOne
SOC 2
Apply
$106k – $210k per year (Estimated) • In office • Full-Time • Bachelor's Degree
Apply
$91k – $179k per year (Estimated) • Remote • Full-Time
C++
Python
DevOps
AWS
Azure
CI/CD
GCP
RTOS
WebSockets
Cybersecurity
GDPR
ISO 27001
Web3
Solana
IoT
CoAP
FreeRTOS
MQTT
Zigbee
Apply
$75k – $180k per year (Estimated) • In office • Full-Time • Adelaide
Apply
HPC Network Engineer 2 months ago
$68k – $161k per year (Estimated) • In office • Full-Time • London
Python
JavaScript
AI/ML
InfiniBand
Frontend
Bootstrap
DevOps
Ansible
Configuration Management
Datadog
Grafana
Prometheus
HPC
Cybersecurity
FortiGate
Web3
Solana
Apply
$65k – $155k per year (Estimated) • In office • Internship • Bachelor's Degree • London
Go
JavaScript
Ruby
Scala
Apply
In office • Internship • 1+ year exp • Bachelor's Degree • London
Go
JavaScript
Ruby
Scala
Apply
$221k – $370k per year • In office • Full-Time • 5+ years exp • London
AI/ML
OpenAI
Apply
$72k – $136k per year (Estimated) • Equity • Remote/Hybrid • 5+ years exp • London
Python
SQL
Databases
Snowflake
AI/ML
AI Agents
DevOps
Kibana
Marketing
Salesforce
Apply
$77k – $148k per year (Estimated) • Equity • In office • Master's Degree • London
JavaScript
Python
Scala
AI/ML
AI Agents
DevOps
GitHub
Apply
See all jobs
This is one of many
368,746 more open roles from verified company boards, updated every day.