654,929open jobs
38,112companies
91,902added this week
Browse all
Salary
$82k – $201k per year (Estimated)
Location
Remote (Australia)
Seniority
Senior · 8+ years exp
Employment
Full-Time
Overview
Company
Impact
Profile match
Jobgether is a Belgian recruitment platform built entirely around remote and flexible work, aggregating openings from thousands of employers that allow work from outside an office. Its matching engine ranks roles against a candidate's skills, seniority and stated preferences on location and flexibility, rather than leaving people to filter a keyword search, and it verifies how genuinely remote each posting is. The company also runs an AI screening layer that shortlists applicants for employers, and publishes research and guidance on distributed work practices alongside the job marketplace itself.

This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Senior Storage Production Engineer - DGX Cloud based in Australia.

Join a high-impact production engineering team responsible for large-scale storage infrastructure supporting demanding AI and machine learning workloads.

You’ll design, operate, and continuously improve distributed storage systems where reliability, availability, performance, and data integrity are critical.

Your work will help ensure GPU cloud services deliver predictable, low-latency access to data at significant scale.

You’ll combine storage engineering, systems development, automation, observability, and cloud-native technologies to solve complex infrastructure challenges.

The role spans the full storage service lifecycle, from architecture and deployment through capacity planning, incident response, optimization, and continuous improvement.

You’ll work across engineering teams to automate operations, improve efficiency, and build resilient systems capable of supporting rapidly evolving AI workloads.

This is an opportunity for an experienced storage engineer to tackle technically challenging problems at the intersection of distributed systems, cloud infrastructure, and AI.

Accountabilities

    • Design, implement, operate, and continuously improve large-scale storage clusters with strong scalability, availability, reliability, and data integrity.

    • Develop and maintain monitoring, logging, alerting, and observability systems to proactively identify storage health and performance issues.

    • Optimize storage architectures for AI/ML workloads, focusing on low-latency data access, efficient caching, high throughput, and predictable performance.

    • Manage the complete lifecycle of storage services, from architecture and initial design through deployment, production operation, and ongoing optimization.

    • Support storage services before production launch through system-build consultation, automation frameworks, capacity planning, and launch readiness reviews.

    • Monitor production infrastructure for availability, latency, capacity, and overall system health.

    • Automate storage operations and implement intelligent mechanisms for fault detection, remediation, data placement, and system optimization.

    • Improve storage efficiency through compression, deduplication, tiering strategies, and intelligent workload placement.

    • Scale storage infrastructure using automation, policy-based tiering, and dynamic data migration techniques.

    • Implement appropriate encryption, access controls, auditing, and other mechanisms to maintain storage security and compliance.

    • Participate in incident response, sustainable operational practices, and blameless root-cause analysis.

    • Participate in an on-call rotation supporting critical storage and production systems.

    • Diagnose and resolve complex storage, networking, performance, reliability, and availability issues.

    • Contribute to infrastructure-as-code, CI/CD, configuration management, and automation practices.

    • Collaborate with engineering and infrastructure teams to improve storage performance, reliability, and operational efficiency.

    • Continuously evaluate emerging storage technologies and approaches that can improve large-scale AI infrastructure.

    • Requirements

      • Bachelor’s degree or equivalent practical experience in Computer Science, Storage Systems, or a related technical discipline.

      • 8+ years of practical experience working with production infrastructure, storage systems, or closely related technologies.

      • Strong experience with distributed and high-performance storage solutions, including clustered and parallel file systems, distributed object storage, and enterprise storage platforms.

      • Deep understanding of block, file, and object storage technologies, including their scalability, reliability, performance characteristics, and operational requirements.

      • Experience with storage networking protocols such as NFS, SMB, iSCSI, S3, Fibre Channel, RDMA, and NVMe over Fabrics.

      • Strong knowledge of algorithms, data structures, complexity analysis, software design, and automation of large-scale Linux-based storage environments.

      • Professional programming experience in one or more relevant languages, such as C/C++, Java, Python, Go, NodeJS, or Bash.

      • Hands-on experience with infrastructure configuration and automation tools such as Ansible, Chef, Puppet, or Terraform.

      • Experience with observability and monitoring technologies such as InfluxDB, Prometheus, Grafana, or the Elastic stack.

      • Strong understanding of distributed storage architectures, replication strategies, erasure coding, capacity planning, performance tuning, and troubleshooting.

      • Experience analyzing and improving distributed storage performance at scale.

      • Strong debugging and systematic problem-solving abilities, particularly when diagnosing complex storage and infrastructure failures.

      • Good understanding of network protocols, architectures, and troubleshooting techniques as they relate to storage performance, stability, and availability.

      • Experience with Git, code review, CI/CD pipelines, and infrastructure-as-code practices is advantageous.

      • Experience operating storage solutions in private or public cloud environments based on Kubernetes, OpenStack, or hybrid cloud architectures is a strong plus.

      • Experience designing automated storage migration, backup, and disaster recovery strategies is beneficial.

      • Strong written and verbal communication skills, with the ability to collaborate effectively across technical teams.

      • Strong work ethic, attention to quality, ownership, and commitment to completing work reliably.

      • Comfortable working collaboratively across teams while adapting to different working styles and evolving technologies.

      • Benefits

        • High-impact infrastructure: Work on large-scale storage systems supporting demanding AI and machine learning workloads.

        • Technical depth: Solve complex challenges across distributed storage, cloud infrastructure, networking, automation, observability, and performance engineering.

        • AI infrastructure exposure: Help optimize storage architectures and operations for advanced AI/ML workloads.

        • Scale and complexity: Work with production systems where availability, performance, reliability, and data integrity are critical.

        • Automation focus: Build intelligent automation for monitoring, fault detection, remediation, data migration, and storage optimization.

        • Remote work: Full-time remote position based in Australia.

        • Cross-functional collaboration: Work with experienced engineers across storage, infrastructure, cloud, networking, and AI-related domains.

        • Continuous innovation: Explore emerging storage technologies and approaches for next-generation cloud infrastructure.

        • Professional growth: Develop expertise in distributed systems, high-performance storage, cloud-native infrastructure, and production engineering.

        • Operational ownership: Take meaningful responsibility for the reliability, scalability, and continuous improvement of critical production systems.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
654,929 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
In your city
SRE Sênior 1 day ago
$44k – $107k per year (Estimated) • Remote • Full-Time
Python
Bash
Databases
Apache Kafka
OpenSearch
AI/ML
Copilot
Cursor
Claude
ChatGPT
Edge AI
DevOps
Terraform
Helm
GitHub Actions
Istio
Prometheus
GitLab CI
CI/CD
GitOps
ArgoCD
Jenkins
AWS
Docker
Kubernetes
Grafana
Platform Engineering
Service Mesh
FinOps
Incident Management
Apply
$52k – $130k per year (Estimated) • Remote • Full-Time • 7+ years exp
Python
JavaScript
Java
Kotlin
TypeScript
SQL
Scala
Groovy
Java
Spring Boot
Databases
PostgreSQL
Snowflake
Databricks
Amazon Redshift
AI/ML
AI Agents
LLM
Frontend
React.js
DevOps
AWS
Kubernetes
Incident Management
Apply
$17k – $45k per year (Estimated) • Remote • Full-Time • 5+ years exp • Bachelor's Degree
Python
Java
SQL
C#
Databases
Oracle
QA
Selenium
Robot Framework
Apply
$178k – $326k per year (Estimated) • In office • 10+ years exp • Bachelor's Degree • Seattle
Python
Java
Scala
Databases
Apache Kafka
AI/ML
Spark
DevOps
AWS
Apply
Remote • Full-Time
JavaScript
Java
Kotlin
TypeScript
Node JS
Java
Spring Framework
Spring Boot
Kotlin
Mockito
Node JS
Nest.JS
Frontend
React.js
Mobile
JUnit
DevOps
Rest API
CI/CD
Git
Docker
Management
Agile
Scrum
Kanban
QA
Jest
Apply
SRE Sênior 1 day ago
$44k – $107k per year (Estimated) • Remote • Full-Time
Python
Bash
Databases
Apache Kafka
OpenSearch
AI/ML
Copilot
Cursor
Claude
ChatGPT
Edge AI
DevOps
Terraform
Helm
GitHub Actions
Istio
Prometheus
GitLab CI
CI/CD
GitOps
ArgoCD
Jenkins
AWS
Docker
Kubernetes
Grafana
Platform Engineering
Service Mesh
FinOps
Incident Management
Apply
$52k – $130k per year (Estimated) • Remote • Full-Time • 7+ years exp
Python
JavaScript
Java
Kotlin
TypeScript
SQL
Scala
Groovy
Java
Spring Boot
Databases
PostgreSQL
Snowflake
Databricks
Amazon Redshift
AI/ML
AI Agents
LLM
Frontend
React.js
DevOps
AWS
Kubernetes
Incident Management
Apply
$47k – $96k per year (Estimated) • Remote • Full-Time • 9+ years exp
JavaScript
Java
Java
Spring Boot
Frontend
Vue.js
DevOps
AWS
Apply
Remote • Full-Time • 10+ years exp • Bachelor's Degree
Apply
$30k – $75k per year (Estimated) • Remote • Full-Time • 12+ years exp • Bachelor's Degree
Analytics
ETL/ELT
Apply
See all jobs
This is one of many
654,929 more open roles from verified company boards, updated every day.