703,818open jobs
41,492companies
99,424added this week
Browse all
Salary
$88k – $203k per year (Estimated)
Location
In office
Seniority
Staff · 10+ years exp
Overview
Company
Impact
Profile match
Nscale is a London-based AI infrastructure company that builds and operates GPU data centres and runs a full-stack AI cloud offering managed inference, Kubernetes and Slurm clusters, bare-metal instances and dedicated GPU capacity. Founded in 2024 by Josh Payne and Nathan Townsend, it develops sites in Norway, the UK, South Korea and North America, works with Microsoft and NVIDIA, and acquired Anyscale in July 2026 to extend its cloud platform. Its hiring spans data centre design and construction, electrical and infrastructure operations, HPC and storage engineering, networking, solutions architecture, legal, finance and marketing.

About Nscale

Nscale is the GPU cloud engineered for AI. We provide cost-effective, high-performance infrastructure for AI start-ups and large enterprise customers. Nscale enables AI-focused companies to achieve superior results by reducing the complexity of AI development. Our GPU cloud strengthens technical capabilities and directly supports strategic business outcomes, including cost management, rapid innovation, and environmental responsibility.

At Nscale, our Engineering team plays a critical role in deploying and operating the infrastructure and software platforms that power our customers.

We thrive on a culture of relentless innovation, ownership, and accountability, where every team member takes pride in their work and drives it with excellence and urgency. As an Nscaler, you’ll build trust through openness and transparency, where everyone is inspired to do their best work. If you join our team, you’ll be contributing to building the technology that powers the future.

About the Role

The Network Engineering Team is responsible for the design, validation, and ongoing operation of all networking services that underpin both the internal management platform and the customer-facing cloud infrastructure - including high-performance Ethernet fabrics, InfiniBand, WAN connectivity, and data center networking. The team also acts as a 3rd/4th line escalation point for the support organization.

As a Staff Network Engineer, you will be a technical authority for Nscale’s AI-optimized network fabrics. You will drive the technical direction across low-latency, high-bandwidth InfiniBand and Ethernet networks supporting large-scale training and inference workloads; own critical technical domains end to end; and raise the bar for architecture, automation, operational rigor, and engineering standards across the organization.

You will combine deep hands-on engineering with broad architectural influence. You’ll define reference architectures, drive consistency across sites, lead complex technical decisions and escalations, and mentor engineers while partnering closely with deployment, data center operations, platform engineering, and vendors.

What You'll Be Doing

  • Define, design, validate, and evolve large-scale InfiniBand/RoCE and Ethernet fabric architectures at rack, row, and data center scale, with tight integration to bare-metal provisioning and cluster management systems.
  • Own technical direction for high-performance Ethernet fabrics, including BGP, EVPN, VXLAN, LACP, and QoS, and establish reference architectures and standards implemented consistently across sites.
  • Design and engineer perimeter and security infrastructure - firewalls, NAT, VPN, and security policy architecture - across WAN and data center edge environments.
  • Lead network automation strategy in a GitOps model, building and guiding Python/Ansible tooling for provisioning, configuration validation, and compliance, with version-controlled configuration and CI/CD-driven change across multi-vendor environments.
  • Drive operational excellence by leading complex escalations and root-cause analysis for performance and stability issues, and systematically reducing reactive toil through runbooks, automation, and measurable SLOs.
  • Set the direction for network observability, telemetry, monitoring, and alerting to provide clear visibility into fabric health, performance, and traffic patterns.
  • Ensure the accuracy and reliability of source-of-truth network inventory and configuration data, with changes flowing through structured engineering and change-management practices.
  • Partner with deployment, data center operations, platform engineering, systems, storage, and vendors on new site delivery and platform evolution.
  • Act as a technical mentor and force multiplier across the team through architecture reviews, design reviews, incident leadership, documentation, and knowledge sharing.
  • Identify systemic risks and architectural gaps across sites and drive durable solutions that improve scalability, reliability, and operational simplicity.

About You (Skills / Qualifications)

  • 10+ years of network engineering experience, with significant depth in HPC, AI, hyperscale, or large-scale data center environments.
  • Extensive hands-on experience with RDMA-aware networking for AI/HPC workloads, including InfiniBand and/or RoCE, subnet managers such as OpenSM/UFM, and fabric orchestration.
  • Expert-level knowledge of modern data center routing and control planes, including BGP, EVPN-VXLAN, and Clos/spine-leaf architectures, with production experience on platforms such as Cumulus, Nokia, or Arista EOS.
  • Strong network automation expertise using Python and Ansible, Git-based workflows, and modern infrastructure-as-code and pipeline tooling such as Terraform, GitLab CI, or GitHub Actions; you treat the network as code rather than managing devices by hand.
  • Deep design and engineering experience with firewall platforms such as Juniper SRX and/or Palo Alto, including security policy architecture, high-availability design, and multi-tenant segmentation.
  • Experience designing network telemetry and observability for high-throughput, performance-sensitive environments.
  • Proven ability to lead complex technical decisions and incidents across networking, systems, storage, and HPC/AI workload teams, with the judgment to balance performance, reliability, operability, and delivery velocity.
  • Demonstrated experience defining architecture, standards, and technical strategy beyond a single project or site, and influencing engineering teams without relying on formal authority.
  • Strong communication and mentoring skills, with the ability to make complex technical trade-offs clear to engineering leaders, operators, and cross-functional partners.
  • Hands-on, adaptable, and comfortable operating with high ownership in a fast-paced environment building next-generation infrastructure for ML scale-out.

Equal Opportunities Statement

At Nscale, we are committed to fostering an inclusive, diverse, and equitable workplace. We believe that a variety of perspectives enriches our work environment, and we encourage applications from candidates of all backgrounds, experiences, and abilities. We strongly encourage applications from people of colour, the LGBTQ+ community, people with disabilities, neurodivergent people, parents, carers, and people from lower socio-economic backgrounds.

If there’s anything we can do to accommodate your specific situation, please let us know.

For information on how Nscale handles candidate personal data, please see our Employee & Candidate Privacy Notice:  Here.

Nscale does not accept unsolicited candidate submissions from recruitment agencies.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
703,818 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
In your city
DevOps 12 hours ago
$17k – $41k per year (Estimated) • Remote/Hybrid • Moscow
Python
Java
Bash
Java
Maven
Gradle
Databases
PostgreSQL
Apache Kafka
OpenSearch
DevOps
gRPC
Terraform
Ansible
Docker Compose
Helm
Loki
Prometheus
GitLab CI
CI/CD
Docker
Kubernetes
Grafana
GitLab
Linux
TCP/IP
Chips/EDA
PoC Library
Apply
In office • Full-Time • Bachelor's Degree • Pylaia
Python
Go
JavaScript
TypeScript
Python
Flask
FastAPI
Django
Frontend
React.js
DevOps
CI/CD
Git
Docker
Management
Agile
Apply
Data Scientist 12 hours ago
$21k – $59k per year (Estimated) • In office • 2+ years exp • Bengaluru
Python
SQL
AI/ML
AI Agents
Machine Learning
Analytics
A/B Testing
Apply
$34k – $82k per year (Estimated) • In office • Chennai
Python
SQL
AI/ML
LangChain
Vertex AI
AI Agents
NLP
AWS Bedrock
TensorFlow
PyTorch
RAG
OpenAI
LLMOps
DevOps
GCP
Azure
AWS
Docker
Kubernetes
Apply
$9k – $21k per year (Estimated) • In office • 1+ year exp • Bachelor's Degree • Dzerzhinsk
Python
C#
C++
C#
.NET
DevOps
Windows
Management
UML
Apply
$220k – $260k per year • In office • 5+ years exp • Bachelor's Degree • New York
AI/ML
InfiniBand
DevOps
Datadog
Prometheus
Grafana
OpenStack
Incident Management
Management
Jira
ServiceNow
ITSM
Apply
$100k – $180k per year • In office • Bachelor's Degree • Seattle
Apply
$157k – $281k per year (Estimated) • In office • 8+ years exp • Houston
AI/ML
Agentic Workflows
Apply
$199k – $431k per year (Estimated) • In office • Bachelor's Degree
AI/ML
Edge AI
Apply
$250k – $350k per year • In office • 15+ years exp
AI/ML
Edge AI
DevOps
SLI/SLO/SLA
Apply
See all jobs
This is one of many
703,818 more open roles from verified company boards, updated every day.