368,657open jobs
9,442companies
50,883added this week
Browse all
Salary
$35k – $84k per year (Estimated)
Location
In office (Bengaluru)
Seniority
Staff · 8+ years exp
Employment
Full-Time
Overview
Company
Impact
Profile match
FlexAI is an artificial intelligence infrastructure company headquartered in Paris, France, and founded in 2023 by former Apple, Intel, and NVIDIA engineers. The company offers workload as a service, abstracting away the specific accelerator so training and inference jobs can run across mixed GPU and alternative silicon without rewriting code. It targets European AI teams that want compute portability and sovereignty rather than lock-in to a single cloud or chip vendor.

About FlexAI

Build and Deploy AI the right way, anywhere.

The FlexAI Compute Infrastructure Platform provides an "end-to-end AI compute layer" for running and managing workloads across any cloud, any GPU, and any deployment model (public, hybrid, or on-prem). It brings together "1-click simplicity" for users with "enterprise-grade orchestration, security, and automation" under the hood.

Founded byBrijesh Tripathi, who bring experience from Nvidia, Apple, Tesla, Intel and Zoox, FlexAI is notjustbuilding a product - we’re shaping the future of AI. Our teams are strategically distributed across Silicon Valley and Bengaluru, united by a shared mission: to deliver more compute with less complexity.

If you're passionate about shaping the future of artificial intelligence, driving innovation, and contributing to a sustainable and inclusive AI ecosystem,FlexAI is the place for you !

Role Overview

FlexAI is looking for a Staff DevOps / SRE Engineer to define our infrastructure strategy, establish SRE best practices, and build systems capable of running large-scale AI workloads across distributed, multi-cloud environments.

You’ll work closely with developers to ensure our platform is reliable, performant, and scalable - without slowing down product velocity.

What You’ll Do

Own Reliability & Architecture:

  • Design and evolve the infrastructure backbone for our AI and PaaS platform
  • Build highly available, fault-tolerant, and scalable systems
  • Define and drive SRE practices (SLIs, SLOs, error budgets)

Build Infrastructure at Scale:

  • Lead Infrastructure as Code using Pulumi
  • Own and scale Kubernetes clusters and containerized workloads
  • Standardize and automate infrastructure for global deployments

CI/CD & Automation:

  • Design and scale CI/CD pipelines for fast, reliable releases
  • Build self-healing systems and automated remediation workflows
  • Drive GitOps and platform engineering practices

Observability & Performance:

  • Implement end-to-end observability using VictoriaMetrics and Grafana (metrics, logs, traces)
  • Identify and resolve performance bottlenecks (latency, throughput, cost)
  • Lead incident response, root cause analysis, and postmortems

Leadership & Collaboration:

  • Partner with backend, AI, runtime, and security teams
  • Guide infrastructure decisions and scaling strategy
  • Mentor engineers and raise the bar on reliability and engineering standards

Security & Resilience:

  • Embed security into infrastructure and deployment workflows
  • Design for resilience (disaster recovery, chaos testing, capacity planning)

What You'll Need to Be Successful

  • 8+ years of experience in DevOps, SRE, or Infrastructure Engineering
  • Proven experience operating large-scale, distributed systems in production
  • Deep expertise in:
    • Kubernetes & container orchestration
    • Pulumi (or similar IaC tools)
    • Cloud or hybrid environments (AWS, GCP, Azure, or on-prem)
    • Observability stacks (Prometheus, Grafana, OpenTelemetry)
  • Strong experience with CI/CD, automation, and release engineering
  • Proficiency in Python, Go, or Bash
  • Strong systems thinking and debugging skills in high-scale environments
  • Experience defining and operating with SLOs / SLAs
  • Experience in startup environments
  • Comfortable leveraging AI coding tools and agents to move faster

Nice to Have

  • Experience with AI/ML infrastructure or GPU workloads
  • Familiarity with distributed or high-performance compute systems
  • Exposure to platform engineering / internal developer platforms
  • Experience scaling systems from Beta to production

Why FlexAI

  • Work on cutting-edge AI infrastructure
  • Build systems that power developers and enterprises
  • High ownership, fast execution, real impact
  • Collaborative, high-caliber team
Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
368,657 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
Bengaluru
$93k – $126k per year • Remote • Full-Time • 5+ years exp • Bachelor's Degree • United States
Node JS
SQL
TypeScript
JavaScript
Databases
MS SQL
Oracle
Mobile
JUnit
DevOps
AWS
Azure
CI/CD
GCP
Git
GitLab
GitLab CI
Jenkins
Management
Jira
QA
JMeter
Playwright
Postman
Rest-Assured
TestNG
Apply
$195k – $264k per year • In office • Full-Time • 15+ years exp • Master's Degree • United States
Python
AI/ML
Amazon SageMaker
Keras
Kubeflow
MLFlow
PyTorch
Scikit-learn
TensorFlow
Vertex AI
XGBoost
DevOps
AWS
Azure
CI/CD
CloudFormation
Docker
GCP
Kubernetes
Terraform
Cybersecurity
FedRAMP
NIST 800-53
Apply
$96k – $218k per year (Estimated) • Equity • In office • Full-Time • 8+ years exp • Toronto
Python
Databases
Databricks
Snowflake
AI/ML
AI Agents
AWS Bedrock
AWS Bedrock AgentCore
LLM
LLM Evaluation
DevOps
AWS
CI/CD
GCP
Apply
$68k – $142k per year (Estimated) • In office • Full-Time • 10+ years exp • Bachelor's Degree • Marseille
AI/ML
Knowledge Graph
DevOps
AWS
Azure
GCP
Management
ServiceNow
Apply
$16k – $60k per year (Estimated) • In office • Full-Time • PhD • Mumbai
Python
SQL
Python
pySpark
Databases
Presto
Snowflake
AI/ML
Dagster
Prefect
Spark
DevOps
Amazon S3
AWS
CI/CD
Analytics
ETL/ELT
Power BI
Tableau
Apply
$78k – $181k per year (Estimated) • In office • Full-Time • Bachelor's Degree • San Jose
JavaScript
Python
AI/ML
Claude
Edge AI
OpenAI Codex
Apply
$179k – $309k per year (Estimated) • In office • Full-Time • 8+ years exp • Santa Clara
C++
Rust
C++
PyTorch C++
TensorFlow C++
AI/ML
JAX
LLM
PyTorch
Ray
TensorFlow
Edge AI
DevOps
CI/CD
Kubernetes
HPC
Apply
$109k – $238k per year (Estimated) • In office • Full-Time • 5+ years exp • Santa Clara
Go
Python
SQL
Databases
Apache Kafka
Cassandra
DynamoDB
PostgreSQL
Redis
AI/ML
PyTorch
TensorFlow
Edge AI
DevOps
AWS
Azure
CI/CD
Docker
GCP
Grafana
gRPC
Kubernetes
OpenTelemetry
Platform Engineering
Prometheus
Apply
$107k – $235k per year (Estimated) • In office • Full-Time • 5+ years exp • Bachelor's Degree • Santa Clara
Node JS
Python
JavaScript
Python
FastAPI
AI/ML
Edge AI
Frontend
React.js
DevOps
CI/CD
Kubernetes
Apply
$26k – $68k per year (Estimated) • In office • Full-Time • 5+ years exp • Bengaluru
Node JS
Python
JavaScript
Python
FastAPI
SQLModel
AI/ML
Edge AI
Frontend
React Query
React.js
DevOps
CI/CD
Kubernetes
Apply
$31k – $82k per year (Estimated) • In office • Full-Time • 3+ years exp • Hyderabad • Bengaluru
Apply
$31k – $73k per year (Estimated) • In office • Full-Time • 5+ years exp • Bengaluru
Apply
$16k – $34k per year (Estimated) • Remote/Hybrid • Full-Time • 2+ years exp • Bachelor's Degree • Mumbai • Bengaluru
JavaScript
PowerShell
SQL
C#
C#
.NET
Databases
Azure SQL Database
MS SQL
DevOps
Azure
Rest API
Cybersecurity
Microsoft Entra ID
QA
Postman
Swagger
Apply
$37k – $73k per year (Estimated) • In office • Internship • 4+ years exp • Bachelor's Degree • Bengaluru
Python
Scala
SQL
Databases
Apache Kafka
Databricks
AI/ML
ChatGPT
Copilot
Cursor
Spark
DevOps
AWS
Azure
CI/CD
GCP
Git
GitHub
Terraform
Apply
$41k – $89k per year (Estimated) • Remote/Hybrid • Full-Time • 8+ years exp • Bengaluru
C#
TypeScript
JavaScript
C#
.NET
Databases
Apache Kafka
AI/ML
Copilot
LLM
OpenAI
Frontend
Angular
GraphQL
DevOps
Azure
Azure AKS
Azure DevOps
CI/CD
Docker
GitHub
GitHub Actions
Grafana
Kubernetes
Prometheus
Rest API
Apply
See all jobs
This is one of many
368,657 more open roles from verified company boards, updated every day.