This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Senior Storage Production Engineer - DGX Cloud based in Australia.
Join a high-impact production engineering team responsible for large-scale storage infrastructure supporting demanding AI and machine learning workloads.
You’ll design, operate, and continuously improve distributed storage systems where reliability, availability, performance, and data integrity are critical.
Your work will help ensure GPU cloud services deliver predictable, low-latency access to data at significant scale.
You’ll combine storage engineering, systems development, automation, observability, and cloud-native technologies to solve complex infrastructure challenges.
The role spans the full storage service lifecycle, from architecture and deployment through capacity planning, incident response, optimization, and continuous improvement.
You’ll work across engineering teams to automate operations, improve efficiency, and build resilient systems capable of supporting rapidly evolving AI workloads.
This is an opportunity for an experienced storage engineer to tackle technically challenging problems at the intersection of distributed systems, cloud infrastructure, and AI.
Accountabilities
Design, implement, operate, and continuously improve large-scale storage clusters with strong scalability, availability, reliability, and data integrity.
Develop and maintain monitoring, logging, alerting, and observability systems to proactively identify storage health and performance issues.
Optimize storage architectures for AI/ML workloads, focusing on low-latency data access, efficient caching, high throughput, and predictable performance.
Manage the complete lifecycle of storage services, from architecture and initial design through deployment, production operation, and ongoing optimization.
Support storage services before production launch through system-build consultation, automation frameworks, capacity planning, and launch readiness reviews.
Monitor production infrastructure for availability, latency, capacity, and overall system health.
Automate storage operations and implement intelligent mechanisms for fault detection, remediation, data placement, and system optimization.
Improve storage efficiency through compression, deduplication, tiering strategies, and intelligent workload placement.
Scale storage infrastructure using automation, policy-based tiering, and dynamic data migration techniques.
Implement appropriate encryption, access controls, auditing, and other mechanisms to maintain storage security and compliance.
Participate in incident response, sustainable operational practices, and blameless root-cause analysis.
Participate in an on-call rotation supporting critical storage and production systems.
Diagnose and resolve complex storage, networking, performance, reliability, and availability issues.
Contribute to infrastructure-as-code, CI/CD, configuration management, and automation practices.
Collaborate with engineering and infrastructure teams to improve storage performance, reliability, and operational efficiency.
Continuously evaluate emerging storage technologies and approaches that can improve large-scale AI infrastructure.
Bachelor’s degree or equivalent practical experience in Computer Science, Storage Systems, or a related technical discipline.
8+ years of practical experience working with production infrastructure, storage systems, or closely related technologies.
Strong experience with distributed and high-performance storage solutions, including clustered and parallel file systems, distributed object storage, and enterprise storage platforms.
Deep understanding of block, file, and object storage technologies, including their scalability, reliability, performance characteristics, and operational requirements.
Experience with storage networking protocols such as NFS, SMB, iSCSI, S3, Fibre Channel, RDMA, and NVMe over Fabrics.
Strong knowledge of algorithms, data structures, complexity analysis, software design, and automation of large-scale Linux-based storage environments.
Professional programming experience in one or more relevant languages, such as C/C++, Java, Python, Go, NodeJS, or Bash.
Hands-on experience with infrastructure configuration and automation tools such as Ansible, Chef, Puppet, or Terraform.
Experience with observability and monitoring technologies such as InfluxDB, Prometheus, Grafana, or the Elastic stack.
Strong understanding of distributed storage architectures, replication strategies, erasure coding, capacity planning, performance tuning, and troubleshooting.
Experience analyzing and improving distributed storage performance at scale.
Strong debugging and systematic problem-solving abilities, particularly when diagnosing complex storage and infrastructure failures.
Good understanding of network protocols, architectures, and troubleshooting techniques as they relate to storage performance, stability, and availability.
Experience with Git, code review, CI/CD pipelines, and infrastructure-as-code practices is advantageous.
Experience operating storage solutions in private or public cloud environments based on Kubernetes, OpenStack, or hybrid cloud architectures is a strong plus.
Experience designing automated storage migration, backup, and disaster recovery strategies is beneficial.
Strong written and verbal communication skills, with the ability to collaborate effectively across technical teams.
Strong work ethic, attention to quality, ownership, and commitment to completing work reliably.
Comfortable working collaboratively across teams while adapting to different working styles and evolving technologies.
High-impact infrastructure: Work on large-scale storage systems supporting demanding AI and machine learning workloads.
Technical depth: Solve complex challenges across distributed storage, cloud infrastructure, networking, automation, observability, and performance engineering.
AI infrastructure exposure: Help optimize storage architectures and operations for advanced AI/ML workloads.
Scale and complexity: Work with production systems where availability, performance, reliability, and data integrity are critical.
Automation focus: Build intelligent automation for monitoring, fault detection, remediation, data migration, and storage optimization.
Remote work: Full-time remote position based in Australia.
Cross-functional collaboration: Work with experienced engineers across storage, infrastructure, cloud, networking, and AI-related domains.
Continuous innovation: Explore emerging storage technologies and approaches for next-generation cloud infrastructure.
Professional growth: Develop expertise in distributed systems, high-performance storage, cloud-native infrastructure, and production engineering.
Operational ownership: Take meaningful responsibility for the reliability, scalability, and continuous improvement of critical production systems.

