1,432,392open jobs
84,555companies
217,679added this week
Browse all
Salary
≈ $110k – $251k per year (Estimated)
Location
Remote (Europe)
Seniority
Staff

Confirmed on the employer's own hiring board on Oct 10, 2026. First seen by Alion on Jul 22, 2026. Radianarc scores B on the Alion truth index.

Overview
Company
Impact
Profile match
Turn GPU infrastructure into real-world intelligence with sovereign AI, intelligent orchestration, GPU-as-a-Service, and cloud gaming solutions from Radian Arc.
Backed by BeAngels

About Radian Arc.

We're specialists in outcome-optimized AI infrastructure - deploying, orchestrating and monetizing GPU compute where data, users and demand actually meet: inside telco networks, at the edge, and in core data centers. Not a generic AI platform. Not a consultancy. We’re the bridge between raw silicon and real-world results.

What impact you will have

Mission: Design, build, and operate the AI storage layer powering large-scale GPU infrastructure, enabling datasets, model artifacts, checkpoints, and inference state to be delivered to compute clusters with extremely high throughput and predictable latency.

You will play a key role in architecting and evolving the storage platform across edge and core deployments, supporting the full lifecycle of AI workloads including distributed inference, fine-tuning, and large-scale model training. The role spans multiple storage architectures used across the platform, including hyperconverged storage currently based on StorPool, local NVMe storage for latency-sensitive workloads and edge deployments, and disaggregated AI storage platforms such as VAST Data and Weka.

As the first dedicated storage platform role in the organization, this position combines Staff-level architectural ownership, technical direction, and cross-functional influence with hands-on execution across storage design, deployment, performance engineering, troubleshooting, platform integration, and operational improvement.

A key responsibility of this role is designing and optimizing the storage architecture underlying distributed inference stacks such as NVIDIA Dynamo, llm-d, or similar inference orchestration frameworks. This includes ensuring that storage systems efficiently support inference workloads through optimized dataset access, model artifact distribution, checkpoint handling, and KV-cache persistence. You will design scalable storage systems capable of feeding thousands of GPUs while balancing throughput, latency, resilience, and cost efficiency, and work closely with compute, networking, and platform engineering teams to ensure seamless integration with the platform orchestration layer.

Because this is currently the primary storage platform role in the company, the position is intentionally hybrid: you are expected to operate at L6 / Staff in terms of long-term design, standards, cross-team influence, and platform direction, while also directly executing critical storage work that, in a larger organization, would be distributed across multiple engineers.

This is a fully remote role, we will consider relevant candidates in all locations.

What you'll need

Core Experience

  • Strong hands-on experience designing and operating distributed storage systems for high-performance compute environments.
  • Proven experience designing storage architectures for large-scale AI inference or training platforms, including dataset distribution, checkpointing, and KV-cache storage patterns.
  • Deep knowledge of the Linux storage and I/O stack.
  • Strong understanding of AI workload data access patterns.
  • Experience optimizing storage for GPU-accelerated workloads.
  • WEKA Data Platform (an enterprise high-performance storage system for AI and HPC)
  • Familiarity with Kubernetes storage integrations such as CSI.
  • Experience operating large-scale storage clusters.
  • Experience owning both architecture and direct implementation in lean or fast-scaling environments is strongly preferred.

Advanced AI Storage Expertise

The candidate should have deep expertise in designing and operating storage platforms optimized for GPU-heavy environments and distributed AI workloads.

This includes a strong understanding of how training, fine-tuning, and inference systems interact with storage, and how storage architecture affects throughput, latency, concurrency, checkpoint recovery, dataset distribution, and serving performance.

Relevant expertise includes:

  • Strong understanding of storage access patterns for distributed inference and training.
  • Experience designing storage platforms that support large dataset ingestion and model artifact distribution at scale.
  • Practical experience tuning storage architectures for checkpointing, distributed file access, object access, and high-concurrency inference.
  • Familiarity with storage patterns for KV-cache persistence and retrieval.
  • Experience optimizing data locality and reducing unnecessary network movement between storage and compute.
  • Understanding of how storage performance affects large-scale AI frameworks, model-serving systems, and inference orchestration layers.

Systems & Troubleshooting

  • Ability to debug complex cross-layer issues spanning:
    • Storage hardware,
    • Networking,
    • Linux kernel and I/O paths,
    • Filesystems,
    • Object and block storage layers,
    • Kubernetes integrations,
    • Distributed workload behavior.
  • Strong knowledge of storage hardware, NVMe devices, storage fabrics, and high-performance data paths.
  • Experience designing storage observability systems.
  • Strong ability to act as the senior escalation point for ambiguous, high-impact, and multi-domain technical issues.

Automation

  • Strong automation skills using Python and/or Bash.
  • Experience applying software engineering practices to storage automation and operational tooling.
  • Experience building reusable tooling, standards, validation patterns, or lifecycle automation that increase leverage across teams.

Leadership

  • Proven ability to lead complex technical initiatives across teams.
  • Comfortable collaborating across engineering, operations, deployment teams, vendors, and platform stakeholders.
  • Strong systems-level thinking balancing performance, reliability, scalability, operability, and cost efficiency.
  • Demonstrated ability to set architectural direction and drive adoption of engineering standards across an organization.
  • Proven ability to lead through technical influence across multiple teams and domains, without relying on formal people management authority.
  • Strong mentoring capability and ability to raise the technical level of adjacent engineering teams.
  • Able to balance short-term execution needs with long-term platform design, operational sustainability, and cost efficiency.

What you’ll do

Storage Architecture

  • Design scalable AI storage architectures supporting both edge and core deployments.
  • Define storage strategies for distributed inference, fine-tuning, and training workloads.
  • Architect solutions across multiple storage models:
    • Hyperconverged infrastructure such as StorPool,
    • Local NVMe storage,
    • Disaggregated storage systems such as VAST, Weka, and related architectures.
  • Define reference architectures, design principles, and reusable patterns for storage platforms so future deployments follow standards rather than one-off implementations.
  • Evaluate trade-offs across throughput, latency, resilience, data locality, cost, and operability, and make clear recommendations to engineering and leadership.
  • Influence the long-term storage roadmap, including architecture choices for edge, core, hyperconverged, and disaggregated environments.

AI Workload Optimization

  • Optimize storage throughput and latency for GPU-heavy clusters.
  • Design data locality strategies to minimize dataset movement across the network.
  • Benchmark storage performance under real AI workloads.
  • Optimize I/O patterns for large dataset ingestion, checkpointing, and model artifact distribution.
  • Work directly with compute teams to ensure storage architecture matches the access patterns of distributed training, fine-tuning, and inference frameworks.
  • Establish performance baselines and validation methods so storage platforms are tested against realistic AI workload behavior rather than only synthetic benchmarks.

Platform Integration

  • Implement and maintain CSI drivers.
  • Integrate storage platforms with Kubernetes and orchestration systems.
  • Integrate block, object, and shared file storage into the platform.
  • Design multi-tenant storage architectures supporting isolated workloads.
  • Ensure storage capabilities are correctly exposed into platform services, workload orchestration, and lifecycle automation.
  • Define standards for how storage should be integrated into Kubernetes-based and platform-managed environments across different deployment models.

Distributed Storage Systems

  • Contribute to the design of exabyte-scale storage platforms.
  • Support S3-compatible object storage, distributed file systems, and block storage.
  • Integrate storage clusters into heterogeneous customer environments.
  • Design storage systems with clear fault domains, lifecycle management approaches, scaling paths, and operational boundaries.
  • Define reusable operating patterns for multi-cluster and multi-site storage environments.

Distributed Inference Storage Architecture

  • Design the storage architecture supporting distributed inference platforms such as NVIDIA Dynamo, llm-d, or similar frameworks.
  • Optimize storage performance for large-scale LLM inference workloads.
  • Design efficient strategies for KV-cache persistence and retrieval using distributed storage platforms such as VAST or Weka.
  • Optimize storage access patterns for token generation pipelines and high-concurrency inference workloads.
  • Ensure inference infrastructure scales efficiently across thousands of GPUs.
  • Partner with platform and inference teams to ensure storage design supports evolving inference architectures and avoids becoming a bottleneck in throughput, latency, or concurrency.

AI Data Path Optimization

  • Design high-performance data paths between GPU clusters and distributed storage.
  • Optimize performance using technologies such as:
    • GPU Direct Storage,
    • RDMA / RoCE,
    • NVMe-oF.
  • Ensure predictable latency for inference serving workloads.
  • Define architectural approaches for storage-to-GPU data movement that balance performance gains with operational complexity and deployment practicality.

Performance Engineering

  • Work with technologies such as the following, to maximize data throughput to GPU clusters.:
    • RDMA,
    • RoCE,
    • GPU Direct Storage,
    • SPDK,
    • NVMe-oF,
  • Lead storage performance investigations across hardware, network, OS, filesystem, and workload interaction points.
  • Drive systematic tuning of storage paths for large-scale GPU environments and define repeatable validation and benchmarking approaches for future deployments.

Reliability & Operations

  • Improve the reliability, durability, and observability of the storage stack.
  • Collaborate with operations teams to monitor storage systems using telemetry and metrics.
  • Optimize performance, latency, and resilience of storage infrastructure.
  • Lead incident response and root-cause analysis for major storage events and chronic performance issues.
  • Translate operational pain points and incidents into durable design changes, standards, runbooks, and architectural improvements.
  • Establish measurable benchmarks for storage reliability, performance consistency, recovery behavior, and operability across deployments.

Engineering Execution & Delivery

  • Lead end-to-end engineering delivery of storage infrastructure from architecture and validation through production rollout.
  • Support practical implementation of storage platforms in both new deployments and existing environments.
  • Validate storage BOMs and architecture assumptions together with infrastructure, compute, and deployment teams.
  • Contribute detailed input into datacenter layouts, node profiles, and storage topology decisions.
  • Drive scaling strategies, capacity planning, and storage lifecycle decisions.
  • Ensure storage changes are executed safely with minimal customer impact.
  • Act as both the architectural owner and the practical execution lead for critical storage initiatives during the build-out phase of the storage function.

Cross-Team Collaboration

  • Work closely with compute, networking, platform, DevOps, and operations teams.
  • Ensure storage integrates seamlessly into the AI platform architecture.
  • Act as the primary storage design authority across the organization, guiding adjacent teams on how storage constraints and capabilities should shape platform decisions.
  • Communicate architectural decisions, trade-offs, risks, and operational implications clearly to stakeholders.
  • Share knowledge and mentor engineers on high-performance storage design.
  • Raise the technical bar by helping adjacent teams better understand storage behavior in distributed AI environments.

Technical Stack

  • CSI.
  • NVMe / NVMe-oF.
  • Distributed file systems.
  • Object storage.
  • Linux storage stack.
  • RDMA / RoCE.
  • GPU Direct Storage.
  • SPDK.
  • StorPool.
  • VAST Data.
  • Weka.
  • MinioFS.
  • Rook Ceph.
  •  

What we offer

  • Attractive compensation package reflecting your expertise and experience.
  • A great work environment characterized by friendliness, international diversity, flexibility, and a hybrid-friendly approach.
  • You'll be part of a fast-growing scale-up with a mission to make a positive impact, offering an exciting career evolution.

Our job titles may span more than one job level. The actual base pay is dependent on a number of factors, such as transferable skills, work experience, business needs and market demands.

Our inclusive responsibility

Radian Arc is committed to creating a diverse and inclusive environment and is proud to be an equal opportunity employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, gender, gender identity or expression, sexual orientation, national origin, genetics, disability, age, veteran status, or any other protected category under applicable law.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
1,432,392 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

DevOps
Similar stack
Same company
In your city
Sr. DevOPS Engineer 7 hours ago
≈ $85k – $193k per year (Estimated) • Remote (United States) • Contractor
Python
Java
Java
Apache Tomcat
DevOps
Puppet
Ansible
Zabbix
Chef
Azure
CI/CD
Jenkins
AWS
Configuration Management
Nagios
TeamCity
Bamboo
Linux
Unix
Management
Agile
Apply
≈ $86k – $197k per year (Estimated) • Remote (United States) • Contractor
Python
DevOps
Terraform
Ansible
OpenShift
New Relic
Prometheus
GitLab CI
CI/CD
GitOps
Jenkins
AWS
Docker
Kubernetes
Grafana
Amazon EKS
GitLab
Linux
Apply
Sr. Cloud Engineer 6 hours ago
≈ $86k – $196k per year (Estimated) • Remote (United States) • Full-Time
Python
JavaScript
Java
SQL
C#
Node JS
C#
.NET
AI/ML
Pandas
DevOps
Kibana
CI/CD
AWS
Kubernetes
Amazon EKS
AWS Fargate
Amazon ECS
DNS
Management
Agile
Scrum
Apply
up to $200k per year • Remote (United States) • Contractor • 5+ years exp • Bachelor's Degree
Python
SQL
Databases
PostgreSQL
Snowflake
DevOps
GCP
Azure
Git
AWS
Analytics
ETL/ELT
Apply
≈ $35k – $89k per year (Estimated) • Remote (Hungary) • Hungary
Python
Go
AI/ML
LLM
DevOps
Terraform
Puppet
Ansible
Loki
VMWare
Prometheus
CI/CD
GitOps
AWS
Kubernetes
Grafana
Cybersecurity
FortiGate
Apply
≈ $84k – $188k per year (Estimated) • Remote (United States) • Full-Time • 2+ years exp • Bachelor's Degree • San Jose
Python
Java
Rust
C++
Databases
PostgreSQL
Redis
AI/ML
Fine-tuning
Multimodal AI
AI Agents
LLM
RAG
LLMOps
Edge AI
DevOps
Terraform
Helm
Istio
Prometheus
Docker
Kubernetes
Grafana
Platform Engineering
OpenStack
Apply
≈ $33k – $98k per year (Estimated) • In office • Beijing
Python
AI/ML
vLLM
SGLang
LLM
Apply
≈ $76k – $144k per year (Estimated) • In office • 5+ years exp • Bachelor's Degree • Germany
Python
JavaScript
ABAP
Python
FastAPI
ABAP
SAP BTP
AI/ML
LangChain
Model Context Protocol
LLM
DevOps
Azure DevOps
Azure
CI/CD
Docker
Management
Agile
BPMN
Apply
≈ $20k – $49k per year (Estimated) • Equity • In office • Full-Time • 10+ years exp • Bachelor's Degree • Hyderabad • Bengaluru
Python
C++
Lua
AI/ML
Copilot
ChatGPT
DevOps
CI/CD
Linux
BGP
OSPF
MPLS
Apply
后端研发工程师 6 hours ago
In office • Beijing
Java
C++
AI/ML
vLLM
SGLang
DevOps
Kubernetes
Apply
≈ $77k – $196k per year (Estimated) • In office • Full-Time
Python
Rust
Databases
ClickHouse
AI/ML
Anomaly Detection
Time Series Forecasting
NCCL
NVLink
Edge AI
KV Cache
Machine Learning
DevOps
OpenTelemetry
Prometheus
CI/CD
Kubernetes
Grafana
Platform Engineering
Incident Management
SLI/SLO/SLA
HPC
Linux
Apply
≈ $80k – $176k per year (Estimated) • Remote (Europe) • Freelance • 7+ years exp
AI/ML
Claude Code
Management
Agile
Scrum
Apply
≈ $10k – $24k per year (Estimated) • Remote (Malaysia) • 3+ years exp
Python
Bash
AI/ML
Machine Learning
DevOps
Ansible
Zabbix
Prometheus
Kubernetes
Grafana
KVM
KubeVirt
SLI/SLO/SLA
HPC
Linux
TCP/IP
DNS
Apply
≈ $99k – $237k per year (Estimated) • Remote (Europe) • Full-Time
Python
Bash
AI/ML
TensorFlow
PyTorch
NCCL
DevOps
HPC
Linux
VLAN
BGP
OSPF
Web3
Layer 2
Apply
≈ $94k – $218k per year (Estimated) • In office • Full-Time • 7+ years exp
Python
AI/ML
AI Agents
LLM
InfiniBand
LLM Guardrails
Machine Learning
DevOps
Prometheus
Kubernetes
Cortex
IAM
Cybersecurity
Keycloak
Snort
HashiCorp Vault
Wazuh
TheHive
Authentik
Zero Trust
Least Privilege
HashiCorp Boundary
PKI
SIEM
Apply
See all jobs
This is one of many
1,432,392 more open roles from verified company boards, updated every day.