694,567open jobs
40,696companies
99,873added this week
Browse all
Salary
$138k – $238k per year (Estimated)
Location
Remote (United States)
Seniority
Senior · 5+ years exp
Employment
Full-Time
Overview
Company
Impact
Profile match
Mirantis is an American infrastructure software company founded in 1999 that helps enterprises run container and cloud platforms on their own hardware rather than on public cloud. It began as an OpenStack specialist, acquired Docker Enterprise in 2019 and now sells Kubernetes distributions, container runtimes and managed services aimed at organisations with regulatory, cost or latency reasons to keep workloads in their own data centres. Headquartered in Campbell, California with a substantial engineering presence in eastern Europe, it competes with Red Hat and the cloud providers' own managed offerings.

About Mirantis

Mirantis, an IREN company, is the Kubernetes-native AI infrastructure company, enabling organizations to build and operate scalable, secure, and sovereign infrastructure for modern AI, machine learning, and data-intensive applications. By combining open source innovation with deep expertise in Kubernetes orchestration, Mirantis empowers platform engineering teams to deliver composable, production-ready developer platforms across any environment-on-premises, in the cloud, at the edge, or in sovereign data centers. As enterprises navigate the growing complexity of AI-driven workloads, Mirantis delivers the automation, GPU orchestration, and policy-driven control needed to manage infrastructure with confidence and agility. Committed to open standards and freedom from lock-in, Mirantis ensures that customers retain full control of their infrastructure strategy. https://www.mirantis.com/

Job Summary

Mirantis is looking for a Technical Product Manager to own observability for k0rdent AI, our control plane for GPU infrastructure and distributed AI workloads. In this role, you will define the observability strategy, roadmap, and feature priorities that determine how operators gain visibility into the health, performance, and resource utilization of GPU clusters running large-scale training and inference. You will shape how k0rdent AI handles everything from GPU-level metrics and distributed tracing across AI workloads, to multi-tenant log aggregation and intelligent alerting - powered by the OpenTelemetry ecosystem, and Prometheus-compatible metrics pipelines.

The ideal candidate brings strong technical fluency in observability tooling and the AI infrastructure stack. You will work directly with engineering to shape requirements, with marketing to define positioning, and with customers to help ensure their success.

Responsibilities

  • Own the vision, roadmap, and priorities for k0rdent AI observability across the full stack: GPU compute, east-west fabric (InfiniBand, RoCE), high-performance storage, DPU/SmartNIC telemetry, workload schedulers, inference serving, and data services

  • Translate requirements from NeoClouds, GPU clouds, telcos, sovereign clouds, and enterprise platform teams into clear product direction; partner with engineering to define requirements and evaluate trade-offs

  • Manage the observability backlog using feedback from production deployments and design partners to refine priorities

  • Track and shape our response to emerging observability standards and technologies, including OpenTelemetry (OTel), DCGM GPU metrics, InfiniBand/RoCE fabric counters, storage platform telemetry APIs, and AI workload profiling

  • Define integration strategies for vendor telemetry sources across the ecosystem - NVIDIA compute and BlueField DPUs, storage platforms (VAST, Weka, DDN), workload managers (SLURM), inference stacks, and vector and relational databases - into a unified, operator-facing observability plane

  • Partner with product marketing and field teams on positioning, technical briefs, and reference architectures; represent Mirantis with customers, analysts, and ecosystem partners

  • 5+ years in product management or a senior technical role owning an observability product or operating large-scale monitoring infrastructure

  • Working knowledge of Prometheus, OpenTelemetry, distributed tracing (Jaeger, Tempo), and log aggregation (Loki, Elasticsearch/OpenSearch)

  • Fluency in Kubernetes observability, cloud-native monitoring, or metrics and alerting pipeline architecture

  • Ability to work directly with engineering on technical trade-offs and with field teams in competitive GPU cloud and NeoCloud deals

Strongly Preferred:

  • Exposure to GPU observability, including DCGM metrics, AI workload profiling and performance analysis

  • Familiarity with east-west fabric telemetry - InfiniBand counters, RoCEv2 congestion metrics (ECN, PFC, DCQCN), or switch-level fabric health

  • Experience with high-performance storage telemetry from platforms such as VAST Data, Weka, or DDN, including IOPS, latency, and throughput instrumentation at scale

  • Familiarity with NVIDIA BlueField DPU telemetry, SR-IOV, or offload pipeline observability

  • Exposure to workload-level visibility for SLURM job scheduling, inference serving stacks (vLLM, Triton, TensorRT-LLM), or data service telemetry from vector databases (Milvus, Qdrant) and relational databases in AI pipelines

Why you’ll love Mirantis

  • Build the observability foundation for the AI cloud era, working directly with leading GPU cloud operators, NeoClouds, sovereign clouds, and AI-first enterprises

  • Collaborate with a world-class, distributed team committed to openness and technical excellence

  • Shape the product narrative and influence go-to-market success

It is understood that Mirantis, Inc. may use automated decision-making technology (ADMT) for specific employment-related decisions. Opting out of ADMT use is requested for decisions about evaluation and review connected with the specific employment decision for the position applied for. You also have the right to appeal any decisions made by ADMT by sending your request to [email protected]

By submitting your resume, you consent to the processing and storage of your personal data in accordance with applicable data protection laws, for the purposes of considering your application for current and future job opportunities.

We are a Leader for Container Management in G2 (#2 after AWS)!

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
694,567 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
United States
In office • Moscow
Python
JavaScript
TypeScript
Python
SQLAlchemy
FastAPI
Alembic
Databases
PostgreSQL
Qdrant
AI/ML
Model Context Protocol
vLLM
AI Agents
LLM
Triton
Frontend
Svelte
DevOps
gRPC
Helm
Prometheus
GitLab CI
CI/CD
Docker
Kubernetes
Grafana
Management
Confluence
Jira
QA
Sentry
Apply
$18k – $43k per year (Estimated) • Remote/Hybrid
Python
PowerShell
Bash
AI/ML
Anomaly Detection
DevOps
Terraform
Ansible
GCP
OpenShift
Loki
CloudFormation
Prometheus
Azure
CI/CD
AWS
Kubernetes
Grafana
Bitbucket
GitHub
GitLab
Amazon CloudWatch
Linux
Apply
$22k – $56k per year (Estimated) • In office • Full-Time • 4+ years exp • Bachelor's Degree • Cape Town
Python
SQL
Python
FastAPI
pySpark
Poetry
Mypy
Databases
Databricks
Apache Iceberg
Delta Lake
MS SQL
Apache Hudi
Trino
AI/ML
llama.cpp
Spark
vLLM
LM Studio
Ollama
LLM
DevOps
Rest API
OpenTelemetry
Azure
CI/CD
Git
Docker
Cybersecurity
GDPR
QA
Pytest
Apply
$80k – $90k per year • Remote • Full-Time • 5+ years exp • United States
Python
JavaScript
SQL
Bash
Databases
ElasticSearch
OpenSearch
DevOps
Ansible
Red Hat
Zabbix
OpenShift
Kibana
Loki
Fluent Bit
Fluentd
Logstash
Prometheus
CI/CD
Kubernetes
Nginx
Grafana
SLI/SLO/SLA
Linux
Apache HTTP Server
Apply
$157k – $190k per year • Equity • Remote • 8+ years exp
AI/ML
Anthropic
DevOps
Splunk
Terraform
GCP
Loki
Datadog
Prometheus
Azure
AWS
Docker
Kubernetes
Grafana
Linux
IoT
Matter
Apply
$124k – $240k per year (Estimated) • Remote/Hybrid • Full-Time • 10+ years exp • Bachelor's Degree • United States
Go
Databases
PostgreSQL
Redis
Apache Kafka
AI/ML
Flink
Machine Learning
DevOps
Terraform
GCP
Azure
GitOps
ArgoCD
AWS
Kubernetes
Platform Engineering
IAM
Cybersecurity
Okta
ISO 27001
SOC 2
Active Directory
LDAP
Apply
$142k – $256k per year (Estimated) • Remote/Hybrid • Full-Time • United States
Go
AI/ML
Machine Learning
DevOps
Rest API
gRPC
Terraform
OpenTofu
GitOps
ArgoCD
AWS
Kubernetes
Platform Engineering
Apply
$58k – $107k per year (Estimated) • Remote/Hybrid • Full-Time • 5+ years exp • Bachelor's Degree • Barcelona
Go
Rust
SQL
Rust
Axum
SQLx
Tonic
Databases
PostgreSQL
AI/ML
InfiniBand
NVLink
Machine Learning
DevOps
Rest API
gRPC
Splunk
Cilium
OpenTelemetry
AWS
Kubernetes
Platform Engineering
KVM
KubeVirt
Linux
DNS
DHCP
BGP
Cybersecurity
Keycloak
Calico
PKI
Apply
$142k – $256k per year (Estimated) • Remote/Hybrid • Full-Time • 5+ years exp • Bachelor's Degree • United States
Go
Rust
SQL
Rust
Axum
SQLx
Tonic
Databases
PostgreSQL
AI/ML
InfiniBand
NVLink
Machine Learning
DevOps
Rest API
gRPC
Cilium
OpenTelemetry
AWS
Kubernetes
Platform Engineering
KVM
KubeVirt
Linux
DNS
DHCP
BGP
Cybersecurity
Keycloak
Calico
PKI
Apply
$151k – $263k per year (Estimated) • Remote • Full-Time • 5+ years exp • United States
Go
TypeScript
AI/ML
Claude Code
Quantization
LLM
OpenAI Codex
Machine Learning
DevOps
Helm
CI/CD
AWS
Kubernetes
Platform Engineering
API Gateway
Apply
General Cardiologist 4 hours ago
$82k – $220k per year (Estimated) • Remote/Hybrid • Full-Time • United States
Apply
Finance Director 4 hours ago
$178k – $329k per year (Estimated) • Remote/Hybrid • Full-Time • 7+ years exp • Bachelor's Degree • United States
AI/ML
Computer Use
Apply
$56k – $113k per year (Estimated) • Remote/Hybrid • Full-Time • 3+ years exp • Bachelor's Degree • United States
Apply
$104k – $209k per year (Estimated) • Remote/Hybrid • Full-Time • 9+ years exp • Bachelor's Degree • United States
Management
Microsoft Office
Apply
$52k – $113k per year (Estimated) • Remote/Hybrid • Full-Time • 1+ year exp • United States
AI/ML
Computer Use
Apply
See all jobs
This is one of many
694,567 more open roles from verified company boards, updated every day.