1,469,448open jobs
87,948companies
231,463added this week
Browse all
Salary
$175k – $250k per year
Location
In office (San Francisco)
Seniority
Middle · 3+ years exp
Employment
Full-Time

Confirmed on the employer's own hiring board on Oct 11, 2026. First seen by Alion on Mar 3, 2026. Blaxel scores B on the Alion truth index.

Overview
Company
Impact
Profile match
Blaxel is the infrastructure foundation for autonomous agents: isolated microVMs that boot in milliseconds and resume in ~25ms, persistent shared memory, and programmable networking.

The role

We're looking for a world-class Site Reliability Engineer to ensure the reliability, performance, and scalability of our AI infrastructure platform.

You’ll be building and operating the core systems that power agentic AI at scale. Your mission: keep our ultra-low-latency, stateful, serverless compute engine rock-solid as we serve billions of agent requests for the most sophisticated AI teams in the world.

This role is highly technical and execution-heavy. You’ll own our reliability posture end-to-end-observability, performance tuning, incident ops, infrastructure health, and the automation systems that keep everything running smoothly. We want you to design new reliability systems, push the boundaries of automation, and continuously evolve the platform to meet the demands of next-generation AI workloads. If you're a builder who thrives on owning critical infrastructure at scale, this role is for you.

What you'll do

Collaborating closely with the founders, the infra team, and the dev team-and leveraging AI wherever it creates leverage-you will architect and operate the systems that keep Blaxel fast, resilient, and secure.

  • Architect, operate, and continuously improve the core infrastructure powering our 25ms cold-start compute engine.

  • Build and evolve our observability stack (metrics, traces, logs), ensuring we detect issues before users do.

  • Define, monitor, and drive SLOs/SLIs across key system surfaces to maintain world-class reliability.

  • Lead incident response with rigor: root cause analysis, post-mortems, and driving systemic fixes.

  • Design and implement self-healing, automated operational systems to eliminate toil and scale ops.

  • Work across compute, networking, storage, and sandboxed execution layers to tune performance under extreme workloads.

  • Build automation and tooling-often with AI agents-to streamline operations, debugging, capacity planning, and failure prediction.

  • Stress-test and push our systems to the edge: load testing, chaos engineering, and performance benchmarking.

  • Own security best practices at the infrastructure layer, from sandboxed compute to network isolation.

  • Partner with platform engineers to ensure reliability is designed into new features from day one.

Who you are

  • Deeply technical by default: Fluent across systems, cloud, networking, and distributed computing. You love debugging real failures, not theoretical ones.

  • AI-fluent operator: You understand how AI systems behave under scale, their unique resource patterns, and the infrastructure challenges of agentic frameworks.

  • Builder at heart: You want to invent new reliability systems-not just maintain existing ones. You thrive in a zero-to-one infra environment.

  • High-velocity execution: You have a strong bias for action and a track record of shipping reliable systems quickly with excellent judgment.

  • Automation-first mindset: You hate repeated manual work and instinctively reach for automation or AI-driven ops to scale yourself.

  • Calm under pressure: When incidents hit, you operate with clarity, precision, and ownership.

  • Data-driven engineer: You measure everything-latency, tail behavior, resource efficiency, reliability trends-and let data guide your decisions.

Required skills

  • 3+ years in SRE, DevOps, or infrastructure engineering roles

  • Strong proficiency in at least one programming language such as Go, Rust, or Python

  • Hands-on experience with a major cloud provider (AWS, GCP)

  • Solid knowledge of Linux systems, networking fundamentals, and distributed systems

  • Experience with bare-metal servers and datacenter operations (PXE/iPXE provisioning, IPMI/BMC, RAID/NVMe, SR-IOV, high-throughput networking)

  • Experience with Kubernetes or similar orchestrators

  • Familiarity with observability stacks (Prometheus, Grafana, ELK, Datadog)

  • Experience building and maintaining CI/CD pipelines (GitHub Actions, GitLab CI, Jenkins)

  • Strong debugging, problem-solving, and incident-management skills

Preferred

  • Experience with infrastructure-as-code tools such as Terraform or Pulumi

  • Knowledge of service mesh or API gateway technologies

  • Exposure to chaos engineering or resiliency-testing frameworks

  • Background in security best practices for cloud environments

  • Prior experience in high-growth or high-availability environments

Bonus

Experience with any of the following is a plus (not required):

  • Serverless compute systems

  • Sandboxed execution environments

  • Ultra-low-latency runtime engineering

  • Distributed key-value stores and databases

  • Chaos engineering

  • Rust, Go, or systems-level programming

  • Deep generative AI infrastructure

About Blaxel

Blaxel is AWS for AI agents. We’re a new kind of cloud computing infrastructure optimized for the unique demands of agentic AI, leveraging a purpose-built 25ms cold-start serverless compute engine.

Now processing billions of agent requests, we power the coding agents and background AI tasks infrastructure for top AI startups. Founders choose us when they hit the limits of general-purpose clouds. We solve the hard infrastructure problems-statefulness, ultra-low latency, and secure sandboxed code execution-so they can focus on building their core AI products.

We raised a $7.3M seed round led by First Round Capital.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
1,469,448 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

DevOps
Similar stack
Same company
San Francisco
$145k – $180k per year • Remote (United States) • Full-Time • 6+ years exp • Bachelor's Degree • Alexandria
Python
JavaScript
SQL
Bash
Databases
Snowflake
Databricks
DynamoDB
ElasticSearch
Amazon Redshift
DevOps
Splunk
Azure DevOps
Kibana
OpenTelemetry
Azure
CI/CD
AWS
Kubernetes
Grafana
minikube
Amazon EKS
Honeycomb
Amazon CloudWatch
API Gateway
Linux
Cybersecurity
HIPAA
NIST 800-53
Management
Agile
Apply
Production Engineer 10 days ago
$89k – $111k per year • In office • 2+ years exp • Bachelor's Degree • Havre de Grace
Analytics
Microsoft Excel
Apply
Production Engineer 9 days ago
$81k – $101k per year • In office • Bachelor's Degree • Marietta
Management
Microsoft Office
Apply
$84k – $123k per year • In office • 2+ years exp • Winona
Apply
≈ $93k – $181k per year (Estimated) • In office • Full-Time • 5+ years exp • Bachelor's Degree • Austin
Apply
≈ $16k – $39k per year (Estimated) • Equity • Remote (India) • Full-Time • 5+ years exp • India
Python
JavaScript
SQL
Databases
PostgreSQL
AI/ML
AI Agents
LLM
Machine Learning
Frontend
Next.js
React.js
DevOps
gRPC
Ansible
GCP
Helm
VMWare
Azure
CI/CD
ArgoCD
Jenkins
AWS
Kubernetes
Bamboo
AWX
Windows
Cybersecurity
Crowdstrike
Management
Agile
Apply
In office • Full-Time
Python
JavaScript
TypeScript
Python
FastAPI
AI/ML
Claude
ChatGPT
Claude Code
LLM
Frontend
Next.js
React.js
DevOps
Terraform
AWS
GitHub
Management
Slack
Apply
In office • Full-Time
Python
TypeScript
AI/ML
Claude
Claude Code
Langfuse
LLM
OpenAI
LLMOps
DevOps
Terraform
Docker Compose
Kali Linux
AWS CDK
CI/CD
AWS
Docker
Kubernetes
Grafana
GitHub
Amazon S3
Amazon ECS
Linux
Cybersecurity
Metasploit
Nmap
Trivy
Semgrep
Analytics
Metabase
Fivetran
Apply
In office • Full-Time • Tokyo
Python
JavaScript
TypeScript
Python
FastAPI
Databases
PostgreSQL
Azure Cosmos DB
AI/ML
Cursor
ChatGPT
Claude Code
Gemini
LLM
OpenAI
Frontend
GraphQL
Next.js
React.js
DevOps
Rest API
Terraform
Azure
Bicep
GitHub
Chips/EDA
PoC Library
Management
Slack
Apply
In office • Full-Time • Tokyo
Python
JavaScript
TypeScript
Python
FastAPI
Databases
PostgreSQL
Azure Cosmos DB
AI/ML
Cursor
ChatGPT
Claude Code
Gemini
LLM
OpenAI
Frontend
GraphQL
Next.js
React.js
DevOps
Rest API
Terraform
Azure
Bicep
GitHub
Chips/EDA
PoC Library
Management
Slack
Apply
Developer Relations 3 months ago
$140k – $190k per year • In office • Full-Time • San Francisco
AI/ML
AI Agents
LLM
Management
Discord
Apply
$140k – $190k per year • In office • Full-Time • San Francisco
TypeScript
AI/ML
AI Agents
Design
Figma
Apply
$150k – $250k per year • In office • Full-Time • San Francisco
Python
Go
Rust
TypeScript
AI/ML
AI Agents
LLM
DevOps
GCP
Azure
CI/CD
AWS
Docker
Kubernetes
Linux
Apply
≈ $231k – $453k per year (Estimated) • In office • 10+ years exp • PhD • San Francisco
Python
Java
C++
Apply
Automation Engineer 6 hours ago
≈ $198k – $384k per year (Estimated) • In office • 5+ years exp • San Francisco
Python
AI/ML
Machine Learning
DevOps
CI/CD
Wi-Fi
Apply
$100k – $150k per year • Equity 0.5–1.5% • In office • Full-Time • San Francisco
Python
DevOps
Linux
Robotics
ROS
Apply
Software Engineer II 5 hours ago
$119k – $160k per year • In office • Full-Time • 3+ years exp • Bachelor's Degree • Glendale • Seattle • San Francisco
JavaScript
TypeScript
Node JS
Node JS
Express
Frontend
Next.js
React.js
Sass
DevOps
CI/CD
Git
Management
Confluence
Jira
Agile
QA
Jest
Mocha
Vitest
Apply
Social Media Manager 6 hours ago
$60k – $110k per year • In office • Full-Time • 1+ year exp • San Francisco
Marketing
YouTube
Instagram
Apply
See all jobs
This is one of many
1,469,448 more open roles from verified company boards, updated every day.