435,297open jobs
15,143companies
67,158added this week
Browse all
Salary
$107k – $246k per year (Estimated)
Location
Remote/Hybrid (Singapore)
Seniority
Staff · 8+ years exp
Employment
Full-Time
Overview
Company
Impact
Profile match
Cloudera provides a hybrid data platform for analytics, machine learning and data engineering. Its lineage traces to early enterprise Hadoop distributions and a later merger with Hortonworks. Large regulated enterprises use it to run analytics across on-premise and cloud estates.

Business Area:

Professional Services

Seniority Level:

Mid-Senior level

Job Description:

At Cloudera, we empower people to transform complex data into clear and actionable insights. With as much data under management as the hyperscalers, we're the preferred data partner for the top companies in almost every industry. Powered by the relentless innovation of the open source community, Cloudera advances digital transformation for the world’s largest enterprises.

As adoption of AI grows across our data and AI platform, customers are using Cloudera AI for increasingly diverse workloads across on-premises, private, sovereign, and public-cloud environments. We are seeking a Staff Systems Engineer - AI Infrastructureto understand the infrastructure requirements and constraints behind these workloads and drive hands-on engineering solutions that improve our product.

This is a hands-on system engineering role. You will investigate ambiguous technical scenarios, identify underlying systems constraints, and develop and validate solutions through experimentation, prototyping, benchmarking, and engineering work. Success means turning diverse workload and infrastructure scenarios into validated technical solutions and, where appropriate, scalable product capabilities and improvements.

As a Staff System Engineer, you will:

  • Work across AI Infrastructure & Systems Engineering: Investigate and design solutions for AI workloads across heterogeneous GPU environments, on-premises datacenters, private infrastructure, and public clouds; reproduce complex scenarios and validate solutions through hands-on experimentation and proof-of-concepts.

  • Work across AI Workload & Inference Engineering: Develop and optimize infrastructure solutions for production AI workloads, including inference and serving, considering workload characteristics, performance, GPU capacity, resource utilization, and deployment constraints.

  • Systems & Performance Engineering: Diagnose issues across Linux, GPU runtimes and drivers, containers, networking, storage, and hardware; identify root causes and validate solutions to performance, reliability, and scalability challenges.

  • GPU Resource Efficiency: Investigate approaches for efficiently allocating and utilizing GPU resources across AI workloads, including workload-aware sharing and partitioning where appropriate.

  • Infrastructure Tooling & Validation: Build diagnostic, benchmarking, deployment, and validation tooling to reproduce complex infrastructure scenarios and evaluate product performance across different environments.

  • Product & Engineering Collaboration: Translate infrastructure findings into technical requirements, product improvements, performance optimizations, and reusable platform capabilities in partnership with product and engineering teams.

We’re excited about you if you have:

  • 8+ years of experience in systems software, distributed infrastructure, platform engineering, performance engineering, or a related field, with a track record of independently solving complex, ambiguous engineering problems.

  • Hands-on experience with production GPU-based infrastructure supporting AI workloads, with a strong understanding of the infrastructure characteristics and constraints that affect them.

  • Ability to take ambiguous problems, develop hypotheses, investigate root causes, and build or validate solutions through experimentation, debugging, prototyping, and measurement without requiring step-by-step direction.

  • Strong understanding of distributed systems, Linux, containers, and production infrastructure, with the ability to reason across multiple layers of the technology stack.

  • Demonstrated ability to diagnose and optimize bottlenecks involving GPU utilization, compute, memory, networking, I/O, or workload/runtime behavior.

  • Hands-on experience designing or operating production infrastructure in on-premises, private-cloud, and/or public-cloud environments, with an understanding of the practical constraints of heterogeneous environments.

You may also have:

  • Experience with modern AI inference and serving technologies such as NVIDIA NIM, vLLM, SGLang, Triton, or equivalent.

  • Experience with Kubernetes or distributed AI workload orchestration.

  • Experience with GPU resource management, sharing, or partitioning, including technologies such as NVIDIA MIG.

  • Experience diagnosing high-performance GPU networking or distributed communication issues.

  • Experience building infrastructure diagnostics, benchmarks, or proof-of-concept systems to evaluate new architectures or technologies.

What you can expect from us:

  • Generous PTO Policy

  • Support work life balance with Unplugged Days

  • Flexible WFH Policy

  • Mental & Physical Wellness programs

  • Phone and Internet Reimbursement program

  • Access to Continued Career Development

  • Comprehensive Benefits and Competitive Packages

  • Paid Volunteer Time

  • Employee Resource Groups

EEO/VEVRAA

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
435,297 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
Singapore
$63k – $129k per year (Estimated) • Remote • 10+ years exp
Java
Kotlin
C++
Java
Spring Boot
Databases
PostgreSQL
DynamoDB
ElasticSearch
DevOps
CI/CD
AWS
Docker
Kubernetes
Platform Engineering
Cybersecurity
PCI DSS
Apply
$152k – $313k per year (Estimated) • Remote • 10+ years exp
Java
Kotlin
C++
Java
Spring Boot
Databases
PostgreSQL
DynamoDB
ElasticSearch
DevOps
CI/CD
AWS
Docker
Kubernetes
Platform Engineering
Cybersecurity
PCI DSS
Apply
$55k – $113k per year (Estimated) • Remote • 10+ years exp
Java
Kotlin
C++
Java
Spring Boot
Databases
PostgreSQL
DynamoDB
ElasticSearch
DevOps
CI/CD
AWS
Docker
Kubernetes
Platform Engineering
Cybersecurity
PCI DSS
Apply
$125k – $148k per year • Equity • Remote/Hybrid • Confidential • Full-Time • Halifax • Manchester • Leeds
DevOps
GCP
Azure
Platform Engineering
Incident Management
Apply
Software Engineer 4 hours ago
$64k – $96k per year • Remote/Hybrid • Confidential • Full-Time • Edinburgh • Manchester • Bristol
Python
Bash
DevOps
CI/CD
Kubernetes
Apply
$109k – $155k per year • Remote/Hybrid • Full-Time • 10+ years exp • Bachelor's Degree • Denver
Analytics
Tableau
Marketing
Salesforce
Apply
$61k – $125k per year (Estimated) • Remote/Hybrid • Full-Time • 8+ years exp • Bachelor's Degree • Budapest • Barcelona • Prague • Madrid
Python
JavaScript
Scala
Databases
Apache Kafka
AI/ML
Flink
DevOps
GCP
Azure
CI/CD
AWS
Docker
Kubernetes
Apply
$300k – $350k per year • Remote/Hybrid • Full-Time • 10+ years exp • Bachelor's Degree • Washington
DevOps
Kubernetes
Apply
$171k – $205k per year • Remote/Hybrid • Full-Time • 8+ years exp • Bachelor's Degree • New York
DevOps
GCP
Azure
AWS
Apply
$35k – $83k per year (Estimated) • Remote/Hybrid • Full-Time • 6+ years exp • Bachelor's Degree • Prague
Python
DevOps
Terraform
Ansible
GCP
Azure
CI/CD
AWS
Kubernetes
QA
Playwright
Swagger
Postman
Rest-Assured
k6
Locust
Apply
In office • Full-Time • Singapore
Marketing
Salesforce
Apply
Remote/Hybrid • Full-Time • PhD • Singapore
Apply
$63k – $160k per year (Estimated) • In office • Internship • Singapore
Python
Apply
$96k – $255k per year (Estimated) • In office • Internship • Singapore
Python
Apply
$66k – $161k per year (Estimated) • In office • Full-Time • 3+ years exp • Singapore
Management
Confluence
Apply
See all jobs
This is one of many
435,297 more open roles from verified company boards, updated every day.