368,611open jobs
9,439companies
50,719added this week
Browse all
Salary
$179k – $309k per year (Estimated)
Location
In office (Santa Clara)
Seniority
Staff · 8+ years exp
Employment
Full-Time
Overview
Company
Impact
Profile match
FlexAI is an artificial intelligence infrastructure company headquartered in Paris, France, and founded in 2023 by former Apple, Intel, and NVIDIA engineers. The company offers workload as a service, abstracting away the specific accelerator so training and inference jobs can run across mixed GPU and alternative silicon without rewriting code. It targets European AI teams that want compute portability and sovereignty rather than lock-in to a single cloud or chip vendor.

About FlexAI

Build and Deploy AI the right way, anywhere.

The FlexAI Compute Infrastructure Platform provides an "end-to-end AI compute layer" for running and managing workloads across any cloud, any GPU, and any deployment model (public, hybrid, or on-prem). It brings together "1-click simplicity" for users with "enterprise-grade orchestration, security, and automation" under the hood.

Founded byBrijesh Tripathi, who bring experience from Nvidia, Apple, Tesla, Intel and Zoox, FlexAI is notjustbuilding a product - we’re shaping the future of AI. Our teams are strategically distributed across Silicon Valley and Bengaluru, united by a shared mission: to deliver more compute with less complexity.

If you're passionate about shaping the future of artificial intelligence, driving innovation, and contributing to a sustainable and inclusive AI ecosystem,FlexAI is the place for you !

Role Overview

At FlexAI, we’re building a high-performance, cloud-agnostic AI compute platform designed for next-generation training and inference workloads. As a Staff AI Runtime Engineer, you’ll play a pivotal role in the design, development, and optimization of the core runtime infrastructure that powers distributed training and deployment of large AI models (LLMs and beyond).

This is a hands-on leadership role - perfect for a systems-minded software engineer who thrives at the intersection of AI workloads, runtimes, and performance-critical infrastructure. You’ll own critical components of our PyTorch-based stack, lead technical direction, and collaborate across engineering, research, and product to push the boundaries of elastic, fault-tolerant, high-performance model execution.

What You'll Do

Lead Runtime Design & Development:

  • Own the core runtime architecture supporting AI training and inference at scale.
  • Design resilient and elastic runtime features (e.g. dynamic node scaling, job recovery) within our custom PyTorch stack.
  • Optimize distributed training reliability, orchestration, and job-level fault tolerance.

Drive Performance at Scale:

  • Profile and enhance low-level system performance across training and inference pipelines.
  • Improve packaging, deployment, and integration of customer models in production environments.
  • Ensure consistent throughput, latency, and reliability metrics across multi-node, multi-GPU setups.

Build Internal Tooling & Frameworks:

  • Design and maintain libraries and services that support model lifecycle: training, checkpointing, fault recovery, packaging, and deployment.
  • Implement observability hooks, diagnostics, and resilience mechanisms for deep learning workloads.
  • Champion best practices in CI/CD, testing, and software quality across the AI Runtime stack.

Collaborate & Mentor:

  • Work cross-functionally with Research, Infrastructure, and Product teams to align runtime development with customer and platform needs.
  • Guide technical discussions, mentor junior engineers, and help scale the AI Runtime team’s capabilities.

What You’ll Need to Be Successful

  • 8+ years of experiencein systems/software engineering, with deep exposure to AI runtime, distributed systems, or compiler/runtime interaction.
  • Experience in delivering PaaS services.
  • Proven experience optimizing and scaling deep learning runtimes(e.g. PyTorch, TensorFlow, JAX) for large-scale training and/or inference.
  • Strong programming skills in Pythonand C++(Go or Rust is a plus).
  • Familiarity with distributed training frameworks, low-level performance tuning, and resource orchestration.
  • Experience working with multi-GPU, multi-node, or cloud-native AI workloads.
  • Solid understanding of containerized workloads, job scheduling, and failure recovery in production environments.

Nice to Have

  • Contributions to PyTorch internalsor open-source DL infrastructure projects.
  • Familiarity with LLM training pipelines, checkpointing, or elastic training orchestration.
  • Experience with Kubernetes, Ray, TorchElastic, or custom AI job orchestrators.
  • Background in systems research, compilers, or runtime architecturefor HPC or ML.
  • Start up previous experience

This position is In-Person and located at our Santa Clara, CA Office.

What We Offer

  • A competitive salary and benefits package
  • Work on cutting-edge AI infrastructure
  • Build products used by developers and enterprises
  • High ownership, fast execution, real impact
  • Collaborative, high-caliber team
Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
368,611 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
Santa Clara
$67k – $160k per year (Estimated) • In office • Full-Time • France
C++
C++
PyTorch C++
TensorFlow C++
AI/ML
Computer Vision
Multimodal AI
OpenCV
PyTorch
TensorFlow
Edge AI
Apply
$123k – $251k per year (Estimated) • In office • Full-Time • 5+ years exp • Bachelor's Degree • Dallas • Denver • Birmingham
Java
SQL
Java
Gradle
Hibernate
Maven
Spring Boot
Spring Framework
Databases
Apache Kafka
MySQL
Redis
DevOps
CI/CD
Dynatrace
Jenkins
Kubernetes
OpenShift
Cybersecurity
SonarQube
Apply
$28k – $58k per year (Estimated) • Remote/Hybrid • Full-Time • 7+ years exp • Moscow
DevOps
Ansible
CI/CD
FinOps
Kubernetes
OpenStack
SLI/SLO/SLA
Terraform
VMWare
Apply
Platform Engineer 1 day ago
$87k – $140k per year • In office • Full-Time • 3+ years exp • Berlin
Databases
PostgreSQL
Redis
DevOps
AWS
Azure
Bicep
CI/CD
Docker
GCP
GitHub Actions
Kubernetes
OpenShift
Terraform
GitHub
Apply
Founding Engineer 1 day ago
$81k – $116k per year • In office • Full-Time • Bachelor's Degree • Munich
JavaScript
Python
TypeScript
Databases
MySQL
PostgreSQL
Frontend
Next.js
React.js
Tailwind CSS
DevOps
AWS
Azure
CI/CD
Docker
GCP
Grafana
Kubernetes
OpenTelemetry
Prometheus
Apply
$78k – $181k per year (Estimated) • In office • Full-Time • Bachelor's Degree • San Jose
JavaScript
Python
AI/ML
Claude
Edge AI
OpenAI Codex
Apply
$109k – $238k per year (Estimated) • In office • Full-Time • 5+ years exp • Santa Clara
Go
Python
SQL
Databases
Apache Kafka
Cassandra
DynamoDB
PostgreSQL
Redis
AI/ML
PyTorch
TensorFlow
Edge AI
DevOps
AWS
Azure
CI/CD
Docker
GCP
Grafana
gRPC
Kubernetes
OpenTelemetry
Platform Engineering
Prometheus
Apply
$107k – $235k per year (Estimated) • In office • Full-Time • 5+ years exp • Bachelor's Degree • Santa Clara
Node JS
Python
JavaScript
Python
FastAPI
AI/ML
Edge AI
Frontend
React.js
DevOps
CI/CD
Kubernetes
Apply
$35k – $84k per year (Estimated) • In office • Full-Time • 8+ years exp • Bengaluru
Go
Python
AI/ML
Edge AI
DevOps
AWS
Azure
CI/CD
GCP
GitOps
Grafana
Kubernetes
OpenTelemetry
Platform Engineering
Prometheus
Pulumi
Self-Healing
VictoriaMetrics
Apply
$26k – $68k per year (Estimated) • In office • Full-Time • 5+ years exp • Bengaluru
Node JS
Python
JavaScript
Python
FastAPI
SQLModel
AI/ML
Edge AI
Frontend
React Query
React.js
DevOps
CI/CD
Kubernetes
Apply
$100k – $137k per year • Equity • In office • Full-Time • 2+ years exp • Bachelor's Degree • Santa Clara
Apply
$72k – $99k per year • Equity • In office • Full-Time • Santa Clara
Apply
$166k – $290k per year • Equity • In office • Full-Time • 8+ years exp • Santa Clara
Management
ServiceNow
Apply
$133k – $272k per year (Estimated) • In office • Santa Clara
Go
Python
AI/ML
Edge AI
LLM
RAG
Apply
$80k – $110k per year • In office • Full-Time • 10+ years exp • Bachelor's Degree • Santa Clara
MATLAB
Python
Apply
See all jobs
This is one of many
368,611 more open roles from verified company boards, updated every day.