1,119,188open jobs
64,547companies
192,307added this week
Browse all
Salary
$200k – $250k per year
Location
Hybrid (Santa Clara, United States)
Seniority
Staff · 10+ years exp
Employment
Full-Time

Confirmed on the employer's own hiring board on Oct 2, 2026. First seen by Alion on Oct 1, 2026.

Overview
Company
Impact
Profile match

DDN

DDN (DataDirect Networks) is a data storage company headquartered in Chatsworth, California that builds high-performance storage systems and a data intelligence platform for AI factories, cloud providers and supercomputing. Formed in 1998 through the merger of MegaDrive and ImpactData by co-founders Alex Bouzari and Paul Bloch, it has systems in 48 of the world's top 100 supercomputers and took a $300 million investment from Blackstone at a $5 billion valuation. It hires Rust, Go and Lustre software engineers, DevOps and build engineers, solutions architects, support engineers, manufacturing test technicians, and sales, tax and legal staff.

We are seeking a Staff Engineer with 10+ years of experience in distributed storage and Linux-based systems engineering. This is a hands-on senior technical role focused on design, debugging, performance, and operational excellence across LustreFS and adjacent stack components. The ideal candidate brings strong expertise in one or more Lustre subsystems, can independently drive complex investigations, and collaborates effectively across engineering, QE, support and release teams. Engineers who are comfortable using AI to accelerate triage, debugging, code comprehension and new feature design will be especially valuable.

Key Responsibilities

  • Design, develop and debug LustreFS features, fixes and enhancements across relevant subsystems such as llite, MDS/MDT, OSS/OST, LDLM and LNet.

  • Investigate customer and scale-related defects, drive root-cause analysis and implement high-quality fixes with strong attention to correctness and maintainability.

  • Contribute to performance tuning, failure analysis and reliability improvements for large-scale Lustre deployments.

  • Participate actively in code reviews, design reviews and subsystem discussions, bringing rigor to testing and operational readiness.

  • Work closely with QE and support to reproduce issues, improve diagnostic data quality and increase coverage for high-risk failure scenarios.

  • Help document subsystem behavior, debugging approaches, known failure patterns and operational best practices.

  • Use AI-assisted tools where appropriate to speed up issue triage, summarize logs, improve code understanding and capture reusable lessons learned.

Required Qualifications

  • 10+ years of experience in systems software, distributed systems, storage, Linux kernel or filesystem engineering.

  • Strong experience in LustreFS development, support or performance engineering with depth in at least one major subsystem.

  • Strong C programming and Linux systems debugging skills.

  • Working knowledge of Linux kernel internals, filesystem semantics, networking and performance analysis.

  • Experience with LNet and/or high-performance transports such as RDMA, InfiniBand, RoCE or TCP-based storage networking.

  • Ability to debug and resolve issues spanning multiple layers including client, server, network and backend storage.

  • Strong collaboration skills and the ability to work across functions in a fast-moving engineering environment.

Preferred Skills

  • Experience in HPC, AI infrastructure or large-scale parallel storage environments.

  • Exposure to metadata-heavy and throughput-heavy workload characterization and tuning.

  • Familiarity with ZFS, ldiskfs, NVMe-backed storage and related observability / performance tooling.

  • Experience creating test plans, reproducer frameworks, runbooks or diagnostic automation.

  • Comfort using AI tools to accelerate debugging, code reviews, triage, documentation and early-stage design ideation.

  • Experience mentoring junior engineers or leading focused technical efforts within a subsystem.

What You Will Work On

  • Hands-on development and debugging of LustreFS defects, performance issues and subsystem enhancements.

  • Customer-facing and scale-related issue investigation across llite, metadata, object storage, LNet and transport layers.

  • Collaborative design and implementation of reliability, observability and serviceability improvements.

  • Reviewing and validating fixes through targeted tests, failure injection, log analysis and performance characterization.

  • Using AI-assisted workflows to accelerate triage, debug loops, code understanding and documentation quality.

  • Contributing to team redundancy by strengthening documentation, code review quality and subsystem knowledge sharing.

Why This Role Matters

This role is central to building durable engineering redundancy in LustreFS: expanding deep subsystem ownership, reducing concentration risk, and accelerating next-generation delivery through strong engineering fundamentals and AI-enabled execution.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
1,119,188 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Industrial Engineering
Similar stack
Same company
Santa Clara
CNC Machinist 2 days ago
≈ $80k – $153k per year (Estimated) • In office • Full-Time • 5+ years exp • High School Diploma • United States
Apply
≈ $85k – $162k per year (Estimated) • In office • Full-Time • 7+ years exp • Bachelor's Degree • Somersworth
Management
Outlook
Microsoft Office
Apply
≈ $100k – $194k per year (Estimated) • In office • Full-Time • 9+ years exp • Bachelor's Degree • Manchester
Cybersecurity
CAPA
Management
Outlook
Microsoft Office
Apply
≈ $94k – $179k per year (Estimated) • In office • Full-Time • 5+ years exp • High School Diploma • Los Angeles
Apply
Lobby Supervisor 1 hour ago
≈ $87k – $167k per year (Estimated) • In office • Secret • 8+ years exp • High School Diploma • Saint Louis
Apply
≈ $61k – $159k per year (Estimated) • In office • Full-Time • 7+ years exp • Cairo
Python
Assembly
Databases
Oracle
AI/ML
InfiniBand
DevOps
Rest API
Terraform
Incident Management
IAM
Linux
Cybersecurity
CIS Benchmarks
Cryptography
Vault
Management
ITIL
Apply
In office • PhD • Singapore
Python
C++
AI/ML
Fine-tuning
Multimodal AI
Computer Vision
VLM
Vision-Language-Action
Machine Learning
DevOps
HPC
Linux
Robotics
ROS
Apply
≈ $29k – $78k per year (Estimated) • Remote (EAEU) • Saint Petersburg
JavaScript
PHP
SQL
1C
PHP
WordPress
Bitrix
Databases
MySQL
MariaDB
Frontend
Vue.js
DevOps
Rest API
CI/CD
Git
Gitflow
Trunk-Based Development
Linux
Apply
≈ $25k – $53k per year (Estimated) • Hybrid • Full-Time • 5+ years exp • Moscow
Python
Go
Bash
Databases
PostgreSQL
DevOps
gRPC
Terraform
Ansible
Zabbix
Helm
Loki
Podman
etcd
Packer
Prometheus
VictoriaMetrics
HAProxy
Docker
Kubernetes
Grafana
KVM
eBPF
SLI/SLO/SLA
Linux
Astra Linux
TCP/IP
DNS
DHCP
VLAN
BGP
OSPF
Cybersecurity
Wireshark
Keycloak
Tcpdump
HashiCorp Vault
Zero Trust
Apply
Equity • Remote (EU)
C++
Databases
PostgreSQL
ClickHouse
AI/ML
AI Agents
Time Series Forecasting
Hugging Face
DevOps
Rest API
gRPC
Terraform
GCP
Loki
containerd
CRI-O
Prometheus
Azure
CI/CD
GitOps
ArgoCD
AWS
Kubernetes
Grafana
Karpenter
Amazon EKS
Google GKE
Azure AKS
Akamai
eBPF
GitHub
GitLab
Linux
Apply
$95k – $110k per year • In office • Full-Time • 5+ years exp • Bachelor's Degree • Chatsworth
DevOps
HPC
Cybersecurity
CAPA
Apply
Staff Engineer 23 days ago
$185k – $275k per year • In office • Full-Time • Santa Clara
Python
C++
AI/ML
LLM
DevOps
OpenTelemetry
Prometheus
Grafana
Amazon S3
Linux
Cybersecurity
Tcpdump
Apply
≈ $125k – $242k per year (Estimated) • Hybrid • Full-Time • 7+ years exp • Bachelor's Degree • Santa Clara
Python
Bash
DevOps
Azure
AWS
HPC
Linux
Apply
≈ $132k – $255k per year (Estimated) • Hybrid • Full-Time • Santa Clara
Python
DevOps
Kubernetes
Platform Engineering
Apply
≈ $91k – $175k per year (Estimated) • In office • Full-Time • High School Diploma • Chatsworth
Analytics
Microsoft Excel
Apply
≈ $77k – $193k per year (Estimated) • In office • High School Diploma • Santa Clara
Apply
Field Supervisor 1 day ago
≈ $77k – $193k per year (Estimated) • In office • High School Diploma • Santa Clara
Apply
≈ $128k – $245k per year (Estimated) • Hybrid • Full-Time • 6+ years exp • Bachelor's Degree • Santa Clara
Python
MATLAB
Chips/EDA
KLayout
Apply
$98k – $129k per year • In office • Full-Time • 3+ years exp • Bachelor's Degree • Santa Clara • Santa Rosa • San Francisco
Apply
Sales Associate 1 day ago
$36k – $50k per year • In office • Santa Clara
Apply
See all jobs
This is one of many
1,119,188 more open roles from verified company boards, updated every day.