368,634open jobs
9,437companies
50,578added this week
Browse all
Salary
$93k – $220k per year (Estimated)
Location
Remote (United Kingdom)
Seniority
Staff
Employment
Full-Time
Overview
Company
Impact
Profile match
Jobgether is an AI-powered job platform focused on remote and flexible work. It matches candidates with relevant roles using skills and preference-based algorithms, and also offers career coaching and job-search guidance.

This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Staff Storage Platform Engineer (AI Storage) - Radian Arc based in United Kingdom.

This is a Staff-level opportunity to shape the storage architecture powering large-scale GPU and AI infrastructure across edge and core environments.

You will design, build, and operate high-performance storage systems supporting training, fine-tuning, and distributed inference workloads.

The role covers hyperconverged, local NVMe, and disaggregated storage architectures, with a strong focus on throughput, latency, resilience, and cost efficiency.

You will work at the intersection of storage, GPUs, networking, Kubernetes, and AI platform engineering, ensuring storage never becomes a bottleneck for compute.

As the primary storage specialist, you will combine architectural ownership with hands-on engineering, troubleshooting, performance optimization, and deployment.

You will also define reusable standards, influence long-term platform direction, and mentor engineers across adjacent infrastructure domains.

This is an ideal environment for a senior storage expert who wants significant technical ownership and direct impact on next-generation AI infrastructure.

Accountabilities

    • Storage architecture: Design scalable storage architectures for edge and core GPU deployments, covering hyperconverged platforms such as StorPool, local NVMe, and disaggregated systems such as VAST Data and Weka. Define reference architectures, reusable design patterns, fault domains, lifecycle strategies, and scaling approaches while balancing throughput, latency, resilience, data locality, operability, and cost.
    • AI workload optimization: Optimize storage for distributed training, fine-tuning, and inference workloads, including large dataset ingestion, model artifact distribution, checkpointing, and high-concurrency access. Establish realistic performance baselines and ensure storage architecture aligns with actual GPU workload behavior.
    • Distributed inference: Design storage architectures supporting inference platforms such as NVIDIA Dynamo, llm-d, or similar systems. Optimize model distribution, token-generation data paths, and KV-cache persistence and retrieval so infrastructure can scale efficiently across large GPU clusters without storage becoming a throughput or latency bottleneck.
    • High-performance data paths: Engineer efficient storage-to-GPU data paths using technologies such as GPU Direct Storage, RDMA/RoCE, NVMe-oF, and SPDK. Investigate and tune performance across hardware, networking, operating systems, filesystems, storage layers, and distributed workloads.
    • Platform integration: Integrate block, object, and shared file storage into Kubernetes and platform orchestration systems. Implement and maintain CSI integrations, support multi-tenant storage architectures, and define standards for storage integration across different deployment models.
    • Distributed storage: Contribute to large-scale storage platforms, including S3-compatible object storage, distributed file systems, and block storage. Design systems with clear operational boundaries, resilience models, scaling paths, and reusable operating patterns across multi-cluster and multi-site environments.
    • Performance and reliability: Lead storage benchmarking, capacity planning, performance investigations, incident response, and root-cause analysis. Establish measurable standards for throughput, latency consistency, recovery behavior, reliability, and operational maturity.
    • Engineering delivery: Own storage initiatives end to end, from architecture and validation through production rollout. Validate BOMs, topology decisions, node profiles, and deployment assumptions while ensuring changes are introduced safely with minimal customer impact.
    • Operational excellence: Improve storage observability, automation, runbooks, lifecycle management, and day-2 operations. Turn recurring incidents and operational pain points into durable engineering improvements and standardized practices.
    • Technical leadership: Act as the primary storage design authority, influencing platform architecture and roadmap decisions across compute, networking, DevOps, infrastructure, and operations. Communicate technical trade-offs clearly, mentor adjacent engineers, and raise the organization’s expertise in AI storage.
    • Requirements

      • Distributed storage expertise: Strong hands-on experience designing and operating distributed storage systems in high-performance computing, AI, GPU, or similarly demanding environments.
      • AI infrastructure experience: Proven experience designing storage architectures for large-scale AI training, fine-tuning, or inference, including dataset distribution, model artifacts, checkpointing, and high-concurrency data access.
      • AI storage knowledge: Deep understanding of how AI workload characteristics affect storage throughput, latency, concurrency, data locality, checkpoint recovery, and serving performance.
      • Storage technologies: Hands-on experience with technologies such as Weka, VAST Data, StorPool, local NVMe, distributed filesystems, S3-compatible object storage, block storage, and/or comparable enterprise storage platforms.
      • Linux and systems expertise: Strong knowledge of the Linux storage and I/O stack, storage hardware, NVMe devices, storage fabrics, and high-performance data paths.
      • Kubernetes: Familiarity with Kubernetes storage integrations, particularly CSI, and experience integrating storage into containerized or orchestrated platforms.
      • AI data paths: Practical knowledge of GPU Direct Storage, RDMA/RoCE, NVMe-oF, SPDK, and techniques for minimizing unnecessary data movement between storage and GPU compute.
      • Distributed inference: Experience with storage requirements for inference orchestration and model-serving environments, including model distribution and KV-cache persistence or retrieval, is highly valuable.
      • Troubleshooting: Ability to diagnose complex cross-layer issues involving storage hardware, networking, Linux kernels and I/O paths, filesystems, object/block storage, Kubernetes, and distributed workloads.
      • Automation: Strong Python and/or Bash skills, with experience applying software engineering practices to infrastructure automation, validation, lifecycle management, and operational tooling.
      • Observability: Experience designing or operating storage observability systems and using metrics and telemetry to identify performance, reliability, and capacity issues.
      • Technical leadership: Demonstrated ability to lead complex infrastructure initiatives, establish architectural standards, and influence multiple teams without relying on formal management authority.
      • Systems thinking: Ability to balance performance, scalability, reliability, operability, deployment complexity, and cost when making architecture decisions.
      • Communication and collaboration: Comfortable working with compute, networking, platform, DevOps, operations, deployment teams, vendors, and other technical stakeholders.
      • Ownership: Able to combine Staff-level strategic thinking with hands-on execution, particularly in a lean or fast-scaling environment where processes and standards are still being established.
      • Mentoring: Strong ability to share knowledge, guide engineers in adjacent domains, and raise the technical bar across the broader infrastructure organization.
      • Benefits

        • Attractive compensation package aligned with your expertise and experience.
        • Opportunity to play a foundational role in shaping a next-generation AI storage platform.
        • Significant architectural ownership and direct influence over long-term infrastructure strategy.
        • Exposure to cutting-edge GPU, AI inference, distributed storage, and high-performance data technologies.
        • International and diverse working environment with strong flexibility.
        • Remote-friendly work model across Europe.
        • Opportunity to join a fast-growing scale-up with an ambitious technology mission.
        • Broad cross-functional exposure across infrastructure, compute, networking, platform engineering, and operations.
        • Strong career growth potential as the infrastructure organization expands.
        • Inclusive environment committed to equal opportunity and professional development.
Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
368,634 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
In your city
$172k – $301k per year • Equity • In office • Full-Time • 8+ years exp • Minneapolis
C++
Go
Java
Python
AI/ML
AI Agents
Fine-tuning
Hybrid Search
LLM
Prompt Engineering
RAG
Semantic Search
Anthropic
Human-in-the-Loop
LLM Guardrails
OpenAI
Semantic Search
Structured Outputs
Function Calling
DevOps
Vector
Cybersecurity
Least Privilege
Management
ServiceNow
Apply
$82k – $110k per year • Remote/Hybrid • Full-Time • PhD • Vancouver
Python
SQL
Databases
Databricks
AI/ML
Anomaly Detection
Embeddings
Fine-tuning
Hallucination
Knowledge Distillation
LLM
RAG
Semantic Search
EU AI Act
Human-in-the-Loop
LLM Guardrails
Semantic Search
AI Agents
DevOps
Azure
Vector
Analytics
A/B Testing
Apply
$90k – $181k per year (Estimated) • In office • 5+ years exp • Bachelor's Degree • Long Beach
MATLAB
Python
Apply
$91k – $194k per year (Estimated) • Equity • In office • Full-Time • 2+ years exp • Bachelor's Degree • Brisbane
Python
SQL
Python
Django
FastAPI
Flask
Databases
Databricks
AI/ML
Function Calling
Hallucination
LLM
NumPy
Scikit-learn
Statsmodels
Streamlit
OpenAI Codex
AI Agents
RAG
DevOps
Azure
CI/CD
Git
Cybersecurity
HIPAA
Analytics
Plotly
QA
Pytest
Apply
In office • 3+ years exp • Bachelor's Degree • Phoenix
PowerShell
Python
DevOps
SLI/SLO/SLA
Cybersecurity
Okta
Tanium
Management
Google Workspace
Slack
ServiceNow
Apply
$126k – $201k per year • Equity • Remote • Full-Time • 5+ years exp • Bachelor's Degree
Analytics
A/B Testing
Apply
$84k – $166k per year (Estimated) • Remote • Full-Time • 7+ years exp • Bachelor's Degree
SQL
Apply
$80k – $190k per year • Remote • Full-Time • 2+ years exp
Apply
$134k – $223k per year (Estimated) • Remote • Full-Time • 5+ years exp • Bachelor's Degree
Bash
Python
AI/ML
Claude
Claude Code
Copilot
OpenAI Codex
DevOps
Azure
Azure DevOps
CI/CD
Gerrit
Git
Jenkins
KVM
QEMU
RTOS
VMWare
Xen
Cybersecurity
Tcpdump
Wireshark
IoT
FreeRTOS
Management
Confluence
Jira
Apply
$165k – $301k per year (Estimated) • Equity • Remote • Full-Time • 12+ years exp
AI/ML
AI Agents
Apply
See all jobs
This is one of many
368,634 more open roles from verified company boards, updated every day.