579,193open jobs
24,952companies
79,803added this week
Browse all
Salary
$143k – $262k per year (Estimated)
Location
Remote (United States)
Seniority
Staff · 6+ years exp
Employment
Full-Time
Overview
Company
Impact
Profile match

Nexxa is building the best AI systems for heavy industries - enabling machines, systems, and operations to think, decide, and act autonomously across manufacturing, large-scale infrastructure, logistics, and legacy environments.

Our mission is to translate deep technical breakthroughs into operational reality, solving some of the hardest systems-level problems in industry.

About the Role

We're looking for a Senior/Staff DevOps Engineer who has spent the last several years building and operating the infrastructure that lets AI and industrial systems run reliably at scale. You understand what it takes to keep production ML and data workloads fast, observable, and resilient - from GPU-backed training and inference clusters to the pipelines that connect them to real-world industrial environments.

This role is ideal for candidates who want deep infrastructure ownership at a company where uptime, latency, and reliability directly affect physical operations - not just software. You'll partner closely with AI, data, and product engineering teams to make sure the systems they build can actually run in production, safely and at scale.

What You'll Do

  • Own and evolve Nexxa's core infrastructure - compute, networking, storage, and deployment systems - end-to-end

  • Design and operate CI/CD pipelines that support fast, safe iteration across AI, data, and product engineering teams

  • Build and maintain infrastructure-as-code (e.g., Terraform, Pulumi) for reproducible, auditable environments across cloud and on-prem/edge deployments

  • Architect and manage Kubernetes-based platforms for training, inference, and application workloads, including GPU scheduling and autoscaling

  • Partner with data and AI teams to support the infrastructure behind:

    • Data warehouses and lakehouse architectures (e.g., Snowflake, BigQuery, Redshift, Databricks)

    • Feature stores, embedding indices, and retrieval pipelines

    • Model training, evaluation, and serving infrastructure

  • Define and drive observability practices - metrics, logging, tracing, and alerting - across distributed systems

  • Establish and enforce reliability practices: SLOs/SLIs, incident response, postmortems, and on-call rotations

  • Design for security and compliance across cloud infrastructure, secrets management, and access control, particularly relevant to industrial and legacy-environment integrations

  • Make pragmatic tradeoffs across cost, latency, reliability, and developer velocity

  • Collaborate with engineering leadership to define infrastructure roadmap and platform strategy

  • Mentor engineers on infrastructure best practices and raise the bar for operational excellence across the org

Required Qualifications

  • 6+ years of experience in DevOps, Site Reliability Engineering, Platform Engineering, or infrastructure-focused software engineering roles

  • Deep hands-on experience with:

    • Cloud platforms (AWS, GCP, or Azure) at production scale

    • Kubernetes in production, including GPU workload scheduling

    • Infrastructure-as-code tooling (Terraform, Pulumi, or equivalent)

    • CI/CD systems (e.g., GitHub Actions, GitLab CI, CircleCI, Jenkins, ArgoCD)

  • Strong track record designing and operating observability stacks (e.g., Prometheus, Grafana, Datadog, OpenTelemetry)

  • Experience supporting ML/AI infrastructure - training clusters, model serving, data pipelines - a strong plus

  • Excellent scripting/programming skills (Python, Go, or Bash) for automation and tooling

  • Proven ability to independently scope and lead infrastructure projects from design through production rollout

  • Strong incident management instincts - you can lead through an outage calmly and drive toward root cause

Preferred Qualifications

  • Experience operating infrastructure that bridges cloud and edge/on-prem environments, especially in industrial or manufacturing contexts

  • Familiarity with data warehouse/lakehouse platforms (Snowflake, BigQuery, Redshift, Databricks)

  • Experience with service mesh, zero-trust networking, or compliance frameworks relevant to industrial/critical infrastructure (e.g., SOC 2, IEC 62443)

  • History of building internal developer platforms or self-service infrastructure tooling

  • Experience scaling infrastructure teams or setting technical direction at a Staff level

What Success Looks Like

  • You can own ambiguous, high-stakes infrastructure problems end-to-end

  • Systems you build stay reliable as usage and scale grow - you design for the next order of magnitude, not just today

  • You bring strong technical judgment on tradeoffs between reliability, cost, and speed

  • You raise the bar for operational rigor and engineering discipline across the team

  • You help define what's next for the platform, not just execute what's known

Why Join Nexxa.ai?

  • Innovative Environment: Play a critical role in transforming heavy industries through groundbreaking AI and automation technologies

  • Collaborative Culture: Be part of a team that values innovation, discipline, and continuous improvement

  • Professional Growth: Benefit from significant opportunities for career development and advancement

  • Competitive Compensation: Enjoy a comprehensive salary and equity package reflective of your expertise and contributions

If you're passionate about building the infrastructure that powers advanced AI solutions in the real world, we'd love to connect.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
579,193 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
In your city
$28k – $66k per year (Estimated) • In office • 4+ years exp • Bengaluru
Python
Go
Rust
DevOps
GCP
AWS
Incident Management
Apply
$122k – $354k per year (Estimated) • Remote/Hybrid • Full-Time • Bachelor's Degree • Edinburgh
Python
Go
JavaScript
Java
TypeScript
Java
Spring Boot
Frontend
Angular
DevOps
Helm
Docker
Kubernetes
Management
Agile
Scrum
Apply
$65k – $156k per year (Estimated) • In office • Full-Time • Master's Degree • Edinburgh
Python
MATLAB
MATLAB
Simulink
Apply
$25k – $63k per year (Estimated) • Equity • In office • Full-Time • 3+ years exp • High School Diploma • Hanoi • Ho Chi Minh City
DevOps
GCP
Apply
$15k per year • In office • Yekaterinburg
Python
SQL
Python
FastAPI
Django
DevOps
Git
Apply
Product Team 1 day ago
$122k – $273k per year (Estimated) • Remote/Hybrid • Full-Time • San Francisco
Design
Figma
Apply
$88k – $243k per year (Estimated) • Equity • Remote • Full-Time • Toronto
Python
Bash
AI/ML
EU AI Act
ISO 42001
DevOps
Terraform
GCP
Pulumi
CI/CD
AWS
Kubernetes
Platform Engineering
Amazon EKS
IAM
Amazon ECS
Cybersecurity
ISO 27001
SOC 2
Least Privilege
Management
Google Workspace
Apply
$129k – $287k per year (Estimated) • Equity • Remote/Hybrid • Full-Time • San Francisco
Python
Bash
AI/ML
EU AI Act
ISO 42001
DevOps
Terraform
GCP
Pulumi
CI/CD
AWS
Kubernetes
Platform Engineering
Amazon EKS
IAM
Amazon ECS
Cybersecurity
ISO 27001
SOC 2
Least Privilege
Management
Google Workspace
Apply
Staff DevOps Engineer 16 days ago
$140k – $266k per year (Estimated) • Equity • Remote • Full-Time • 6+ years exp • Toronto
Python
Go
Databases
Snowflake
Databricks
Google BigQuery
Amazon Redshift
BigQuery
AI/ML
Feature Store
DevOps
Terraform
GCP
GitHub Actions
OpenTelemetry
CircleCI
Datadog
Prometheus
Pulumi
GitLab CI
Azure
CI/CD
ArgoCD
Jenkins
AWS
Kubernetes
Grafana
Platform Engineering
Service Mesh
Incident Management
Cybersecurity
SOC 2
Zero Trust
Apply
$64k – $171k per year (Estimated) • Equity • Remote/Hybrid • Full-Time • Toronto
Python
AI/ML
Weights & Biases
Function Calling
AI Agents
Arize Phoenix
DeepEval
Langfuse
LangSmith
Promptfoo
LLM
RAG
Hallucination
Human-in-the-Loop
LLM Evaluation
DevOps
CI/CD
Apply
See all jobs
This is one of many
579,193 more open roles from verified company boards, updated every day.