376,649open jobs
9,801companies
47,951added this week
Browse all
Salary
$144k – $245k per year (Estimated)
Location
Remote (United States)
Seniority
Senior · 5+ years exp
Employment
Full-Time
Overview
Company
Impact
Profile match
Mirantis is a B2B open-source cloud computing, container management, and AI infrastructure company headquartered in Campbell, California. Originally known as a core contributor to OpenStack, Mirantis now provides open-cloud software and managed services centered on Kubernetes, multi-cloud platforms, and enterprise AI workloads.

About Mirantis

Mirantis, an IREN company, is the Kubernetes-native AI infrastructure company, enabling organizations to build and operate scalable, secure, and sovereign infrastructure for modern AI, machine learning, and data-intensive applications. By combining open source innovation with deep expertise in Kubernetes orchestration, Mirantis empowers platform engineering teams to deliver composable, production-ready developer platforms across any environment-on-premises, in the cloud, at the edge, or in sovereign data centers. As enterprises navigate the growing complexity of AI-driven workloads, Mirantis delivers the automation, GPU orchestration, and policy-driven control needed to manage infrastructure with confidence and agility. Committed to open standards and freedom from lock-in, Mirantis ensures that customers retain full control of their infrastructure strategy. https://www.mirantis.com/

Job Summary

Mirantis is looking for a Technical Product Manager to own observability for k0rdent AI, our control plane for GPU infrastructure and distributed AI workloads. In this role, you will define the observability strategy, roadmap, and feature priorities that determine how operators gain visibility into the health, performance, and resource utilization of GPU clusters running large-scale training and inference. You will shape how k0rdent AI handles everything from GPU-level metrics and distributed tracing across AI workloads, to multi-tenant log aggregation and intelligent alerting - powered by the OpenTelemetry ecosystem, and Prometheus-compatible metrics pipelines.

The ideal candidate brings strong technical fluency in observability tooling and the AI infrastructure stack. You will work directly with engineering to shape requirements, with marketing to define positioning, and with customers to help ensure their success.

Responsibilities

  • Own the vision, roadmap, and priorities for k0rdent AI observability across the full stack: GPU compute, east-west fabric (InfiniBand, RoCE), high-performance storage, DPU/SmartNIC telemetry, workload schedulers, inference serving, and data services

  • Translate requirements from NeoClouds, GPU clouds, telcos, sovereign clouds, and enterprise platform teams into clear product direction; partner with engineering to define requirements and evaluate trade-offs

  • Manage the observability backlog using feedback from production deployments and design partners to refine priorities

  • Track and shape our response to emerging observability standards and technologies, including OpenTelemetry (OTel), DCGM GPU metrics, InfiniBand/RoCE fabric counters, storage platform telemetry APIs, and AI workload profiling

  • Define integration strategies for vendor telemetry sources across the ecosystem - NVIDIA compute and BlueField DPUs, storage platforms (VAST, Weka, DDN), workload managers (SLURM), inference stacks, and vector and relational databases - into a unified, operator-facing observability plane

  • Partner with product marketing and field teams on positioning, technical briefs, and reference architectures; represent Mirantis with customers, analysts, and ecosystem partners

  • 5+ years in product management or a senior technical role owning an observability product or operating large-scale monitoring infrastructure

  • Working knowledge of Prometheus, OpenTelemetry, distributed tracing (Jaeger, Tempo), and log aggregation (Loki, Elasticsearch/OpenSearch)

  • Fluency in Kubernetes observability, cloud-native monitoring, or metrics and alerting pipeline architecture

  • Ability to work directly with engineering on technical trade-offs and with field teams in competitive GPU cloud and NeoCloud deals

Strongly Preferred:

  • Exposure to GPU observability, including DCGM metrics, AI workload profiling and performance analysis

  • Familiarity with east-west fabric telemetry - InfiniBand counters, RoCEv2 congestion metrics (ECN, PFC, DCQCN), or switch-level fabric health

  • Experience with high-performance storage telemetry from platforms such as VAST Data, Weka, or DDN, including IOPS, latency, and throughput instrumentation at scale

  • Familiarity with NVIDIA BlueField DPU telemetry, SR-IOV, or offload pipeline observability

  • Exposure to workload-level visibility for SLURM job scheduling, inference serving stacks (vLLM, Triton, TensorRT-LLM), or data service telemetry from vector databases (Milvus, Qdrant) and relational databases in AI pipelines

Why you’ll love Mirantis

  • Build the observability foundation for the AI cloud era, working directly with leading GPU cloud operators, NeoClouds, sovereign clouds, and AI-first enterprises

  • Collaborate with a world-class, distributed team committed to openness and technical excellence

  • Shape the product narrative and influence go-to-market success

It is understood that Mirantis, Inc. may use automated decision-making technology (ADMT) for specific employment-related decisions. Opting out of ADMT use is requested for decisions about evaluation and review connected with the specific employment decision for the position applied for. You also have the right to appeal any decisions made by ADMT by sending your request to [email protected]

By submitting your resume, you consent to the processing and storage of your personal data in accordance with applicable data protection laws, for the purposes of considering your application for current and future job opportunities.

We are a Leader for Container Management in G2 (#2 after AWS)!

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
376,649 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
United States
DevOps инженер 3 hours ago
$16k – $40k per year (Estimated) • In office • Bachelor's Degree • Saint Petersburg
PHP
Python
Databases
Apache Kafka
ElasticSearch
MySQL
PostgreSQL
RabbitMQ
Redis
DevOps
Ansible
AWS
CI/CD
Docker
Git
GitHub
GitHub Actions
GitLab
Grafana
Helm
Jenkins
Kibana
Kubernetes
Logstash
Prometheus
Terraform
Zabbix
Apply
In office • Full-Time • 4+ years exp • Bachelor's Degree • Kathmandu
Java
Python
Scala
SQL
JavaScript
Java
Spring Boot
Databases
PostgreSQL
Snowflake
AI/ML
ChatGPT
Claude
Copilot
Cursor
Frontend
React.js
Mobile
State Management
DevOps
AWS
Azure
CI/CD
Docker
Git
GitHub
Kubernetes
Rest API
Apply
$22k – $60k per year (Estimated) • In office • 4+ years exp • Bengaluru
Python
Python
FastAPI
DevOps
AWS
Docker
GCP
Kubernetes
Terraform
Apply
AI Solution Architect 3 hours ago
$32k – $76k per year (Estimated) • In office • 10+ years exp • Delhi
Python
Databases
Google BigQuery
Milvus
Pinecone
PostgreSQL
Qdrant
Weaviate
AI/ML
AI Agents
Amazon SageMaker
AWS Bedrock
Fine-tuning
Hugging Face
LangChain
LangGraph
LLM
OpenAI
PyTorch
RAG
TensorFlow
Transformers
Vertex AI
Frontend
GraphQL
DevOps
AWS
AWS Lambda
Azure
Azure AKS
GCP
IAM
WebSockets
Kubernetes
Apply
$15k – $37k per year (Estimated) • In office • Full-Time • 7+ years exp • Bachelor's Degree • Pune
DevOps
Docker
Kubernetes
Rest API
Management
Confluence
Jira
Apply
$127k – $248k per year (Estimated) • Remote/Hybrid • Full-Time • United States
Go
Python
AI/ML
InfiniBand
NVLink
DevOps
AWS
Grafana
Kubernetes
OpenTelemetry
Platform Engineering
Prometheus
VictoriaMetrics
Apply
$72k – $114k per year (Estimated) • Remote/Hybrid • Full-Time • 5+ years exp • Bachelor's Degree • Warsaw
Go
Rust
SQL
Rust
Axum
SQLx
Tonic
Databases
PostgreSQL
AI/ML
InfiniBand
NVLink
DevOps
AWS
Cilium
gRPC
Kubernetes
KubeVirt
KVM
OpenTelemetry
Platform Engineering
Rest API
Splunk
Cybersecurity
Calico
Keycloak
Apply
Remote/Hybrid • Full-Time • 5+ years exp • Bachelor's Degree • Prague
Go
Rust
SQL
Rust
Axum
SQLx
Tonic
Databases
PostgreSQL
AI/ML
InfiniBand
NVLink
DevOps
AWS
Cilium
gRPC
Kubernetes
KubeVirt
KVM
OpenTelemetry
Platform Engineering
Rest API
Splunk
Cybersecurity
Calico
Keycloak
Apply
Remote/Hybrid • Full-Time • Riga
Go
Python
Rust
Databases
ElasticSearch
OpenSearch
AI/ML
Anomaly Detection
DevOps
AIOps
AWS
Cortex
eBPF
Grafana
Incident Management
Jaeger
Kubernetes
Loki
Mimir
OpenTelemetry
Platform Engineering
Prometheus
SLI/SLO/SLA
Splunk
Thanos
HPC
Apply
$185k – $352k per year (Estimated) • Remote/Hybrid • Full-Time • United States
AI/ML
CUDA
CUDA Toolkit
Fine-tuning
LLM
PyTorch
TensorRT
TensorRT-LLM
Triton
DevOps
AWS
Kubernetes
KubeVirt
Platform Engineering
SLURM
Apply
$129k – $203k per year • Remote/Hybrid • Full-Time • 3+ years exp • Bachelor's Degree • Cambridge
Python
Apply
$81k – $162k per year (Estimated) • Equity • Remote • Full-Time • 5+ years exp • Charlotte
Analytics
Power BI
Apply
$125k – $226k per year • In office • Full-Time • 12+ years exp • Bachelor's Degree • Princeton • Austin
JavaScript
SQL
TypeScript
Java
Java
Spring Boot
Spring Cloud
Databases
Apache Kafka
Cassandra
DynamoDB
PostgreSQL
Redis
Frontend
Angular
React.js
DevOps
AWS
Azure
Azure DevOps
CI/CD
CloudFormation
Docker
GCP
GitHub
GitHub Actions
GitLab
Jenkins
Kubernetes
Rest API
Terraform
Apply
Fraud Data Analyst 1 day ago
$98k – $163k per year • In office • Full-Time • 3+ years exp • Bachelor's Degree • Arlington
Python
SQL
DevOps
Git
GitLab
Analytics
Tableau
Apply
$133k – $338k per year • Remote/Hybrid • Full-Time • 8+ years exp • Associate's Degree • Chicago • Milwaukee • Dallas • Columbus • Kirkland
AI/ML
Knowledge Graph
DevOps
AWS
Azure
GCP
Platform Engineering
SLI/SLO/SLA
Apply
See all jobs
This is one of many
376,649 more open roles from verified company boards, updated every day.