549,821open jobs
20,529companies
75,324added this week
Browse all
Salary
$180k – $260k per year
Location
Remote (United States)
Seniority
Staff · 8+ years exp
Employment
Full-Time
Overview
Company
Impact
Profile match
Jobgether is a Belgian recruitment platform built entirely around remote and flexible work, aggregating openings from thousands of employers that allow work from outside an office. Its matching engine ranks roles against a candidate's skills, seniority and stated preferences on location and flexibility, rather than leaving people to filter a keyword search, and it verifies how genuinely remote each posting is. The company also runs an AI screening layer that shortlists applicants for employers, and publishes research and guidance on distributed work practices alongside the job marketplace itself.

This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a HPC Storage Engineer - West Coast based in the United States.

This is a senior, hands-on infrastructure engineering role focused on building and scaling a multi-region storage platform for demanding AI workloads.

You’ll own critical storage systems spanning network volumes, local NVMe, and S3-compatible object storage at petabyte scale.

Your work will directly influence training, fine-tuning, and inference performance, including cold-start speed, data streaming, and workload reliability.

You’ll operate at the intersection of storage, networking, hardware, and software, with substantial ownership from architecture through production operations.

The role offers significant latitude to automate manual processes, establish SLOs, optimize performance, and shape long-term storage strategy.

You’ll collaborate closely with SRE, network engineering, supply chain teams, and infrastructure partners in a fast-moving remote environment.

This opportunity is ideal for an engineer who enjoys solving complex production problems and building infrastructure that serves millions of developers.

Accountabilities

    • Own the capacity, durability, availability, and performance of network volumes, local NVMe, and S3-compatible object storage.
    • Tune the complete I/O path, including device and filesystem configuration, caching, read-ahead strategies, replication, erasure coding, and client-side mount behavior.
    • Diagnose complex storage and performance issues end to end, identifying root causes and implementing durable solutions.
    • Lead capacity expansions, hardware refreshes, migrations, and data rebalancing while minimizing or eliminating customer-visible disruption.
    • Design and optimize the networking infrastructure supporting storage workloads, including high-throughput east-west fabrics, MTU and jumbo-frame configuration, congestion and flow control, multipath, and NIC/offload settings.
    • Optimize storage traffic across RDMA/RoCE and high-speed InfiniBand or Ethernet environments, collaborating with network engineering on topology, oversubscription, and cross-region data movement.
    • Develop and ship production software in Go, Python, Rust, or similar languages for storage control-plane services, provisioning, data movement, and monitoring.
    • Build and extend integrations with internal control-plane services, S3-compatible interfaces, CSI drivers, Kubernetes APIs, vendor platforms, and cloud-provider APIs.
    • Replace manual operational procedures with reliable automation and infrastructure-as-code, while participating fully in code reviews, testing, and CI.
    • Instrument storage infrastructure with meaningful metrics covering IOPS, throughput, latency, errors, retries, capacity utilization, and tenant consumption.
    • Build dashboards, SLOs, and alerts that identify degradation proactively and support reliable production operations.
    • Participate in an on-call rotation and lead blameless post-incident follow-through, ensuring lessons learned translate into measurable system improvements.
    • Requirements

      • 8+ years of experience in infrastructure, storage, or systems engineering, including substantial ownership of production storage environments at scale.
      • Deep practical experience with at least one distributed storage platform such as Ceph, MinIO, Lustre, GPFS/Spectrum Scale, MooseFS, WekaFS, VAST, ZFS-based systems, or a comparable technology.
      • Strong knowledge of Linux internals and the storage stack, including block devices, filesystems, NVMe, page cache, I/O schedulers, NFS/SMB, iSCSI, and NVMe-oF.
      • Hands-on experience building or operating S3-compatible object storage services.
      • Strong networking fundamentals and demonstrated experience tuning networks specifically for storage workloads.
      • Proven ability to write and ship production-quality software using Go, Python, Rust, or a similar programming language, beyond scripting alone.
      • Experience with observability platforms such as Prometheus, Grafana, Datadog, or equivalent, including designing meaningful metrics and monitoring strategies.
      • Demonstrated ability to analyze and resolve performance problems under real production pressure.
      • Self-directed approach, with the ability to take broad infrastructure goals, develop an options analysis, diagnose problems, and execute solutions with minimal supervision.
      • Strong continuous-improvement mindset, with a track record of eliminating operational toil and replacing recurring manual work with automation.
      • High ownership and accountability, including the willingness to follow problems across team boundaries through to resolution.
      • Collaborative, low-ego communication style combined with confidence in technical decision-making.
      • Experience with AI/ML storage workloads, including checkpointing, dataset streaming, model-weight distribution, GPU-adjacent data locality, or GPUDirect Storage, is preferred.
      • Familiarity with Kubernetes storage internals, including CSI drivers, PV/PVC lifecycles, StatefulSets, and local persistent volumes, is a plus.
      • Bare-metal or colocation experience, including hardware selection, vendor management, firmware, and physical failure domains, is beneficial.
      • Experience operating multi-tenant environments where isolation, fairness, and quality of service are critical is preferred.
      • Background in a rapidly scaling cloud or infrastructure provider is advantageous.
      • Benefits

        • Base salary: $180,000-$260,000, with the final range determined based on career level, experience, qualifications, and location.
        • Meaningful equity through stock options, giving employees an opportunity to share in the company’s growth.
        • Generous medical, dental, and vision coverage.
        • Flexible paid time off.
        • Remote-first work environment with collaborative teams and Slack as a primary internal communication channel.
        • $1,200 home office and equipment stipend to help create an effective remote workspace.
        • Opportunity to work on cutting-edge AI infrastructure with a strong emphasis on ownership, learning, and technical impact.
        • Inclusive workplace committed to equal opportunity and respect for people from diverse backgrounds.
        • Candidates must be legally authorized to work in the United States; employment visa sponsorship is not currently available.
Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
549,821 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
In your city
Remote/Hybrid • 8+ years exp
Python
JavaScript
TypeScript
Python
Flask
FastAPI
Django
Databases
Snowflake
AI/ML
LangChain
LlamaIndex
LLM
RAG
Frontend
GraphQL
React.js
DevOps
Helm
Prometheus
WebSockets
Docker
Kubernetes
Cortex
QA
Cypress
Playwright
Pytest
Apply
$17k – $42k per year (Estimated) • Remote • Full-Time • Moscow
Python
Bash
Databases
OpenSearch
AI/ML
Model Context Protocol
vLLM
Ollama
LLM
DevOps
Terraform
Docker Compose
Helm
Prometheus
Yandex Cloud
GitLab CI
CI/CD
GitOps
ArgoCD
Docker
Kubernetes
Grafana
Harbor
GitLab
Cybersecurity
SBOM
Apply
$16k – $41k per year (Estimated) • In office • Full-Time • 3+ years exp • Bachelor's Degree • Hyderabad
Python
PowerShell
AI/ML
AI Agents
LLM Guardrails
NIST AI RMF
DevOps
GCP
Azure
AWS
Cybersecurity
OWASP Top 10
Apply
$13k – $37k per year (Estimated) • In office • 3+ years exp • Bengaluru
Python
JavaScript
SQL
DevOps
CI/CD
QA
Cypress
Playwright
Appium
BrowserStack
Apply
Engineering Manager 11 min ago
$23k – $64k per year (Estimated) • In office • 10+ years exp • Pune
Python
JavaScript
TypeScript
Node JS
Python
pySpark
Databases
MySQL
PostgreSQL
Redis
RabbitMQ
Apache Kafka
AI/ML
Cursor
LangChain
Hadoop
Spark
ChatGPT
YOLO
AI Agents
LLM
RAG
NVIDIA NeMo
Roboflow
Frontend
React.js
DevOps
GCP
AWS
Docker
Kubernetes
Nginx
Management
Agile
Apply
$196k – $270k per year • Equity • Remote • Full-Time • 5+ years exp • Bachelor's Degree
Apply
$150k – $220k per year • Equity • Remote • Full-Time • 5+ years exp
JavaScript
TypeScript
Node JS
Frontend
React.js
Mobile
React Native
DevOps
AWS
Kubernetes
Platform Engineering
QA
Selenium
Cypress
Playwright
Appium
Apply
$57k – $65k per year • Remote • Full-Time
Apply
$99k – $223k per year (Estimated) • Remote • Full-Time • 5+ years exp
Marketing
Salesforce
Apply
R&D Engineer 1 hour ago
$101k – $211k per year (Estimated) • Equity • Remote • Full-Time • 5+ years exp
Python
AI/ML
CUDA Toolkit
YOLO
Fine-tuning
Quantization
Knowledge Distillation
Computer Vision
PyTorch
CUDA
Edge AI
Model Distillation
Robotics
Localization
Apply
See all jobs
This is one of many
549,821 more open roles from verified company boards, updated every day.