787,996open jobs
49,912companies
122,545added this week
Browse all
Salary
$250k – $325k per year
Location
Hybrid (Foster City, United States)
Seniority
Staff
Employment
Full-Time

Confirmed on the employer's own hiring board on Sep 25, 2026. First seen by Alion on Sep 25, 2026. Replit scores A on the Alion truth index.

Overview
Company
Impact
Profile match
Replit is a cloud-based software creation platform and integrated development environment (IDE) that enables users to write, execute, and deploy code directly within a web browser. The platform supports a wide array of programming languages and offers real-time collaborative coding, automated environment setup, and integrated AI-powered development tools.

Replit is the agentic software creation platform that enables anyone to build applications using natural language. With millions of users worldwide, Replit is democratizing software development by removing traditional barriers to application creation.

About the Role

Replit enables people to build software with AI. The systems underneath that experience must support safe production changes, measurable reliability, and predictable performance as usage grows.

This Engineering Manager will lead SRE across observability, incident management, load testing, performance engineering, cloud cost and capacity, and rollout infrastructure. You'll lead and grow an existing team that builds and operates production platforms and works hands-on across application and infrastructure boundaries.

This is a software-building leadership role, not simply an incident-management function. You'll help teams ship safely, understand production behavior, and remove performance bottlenecks through concrete engineering improvements. You should be comfortable going deep on a rollout failure or performance investigation while developing technical leaders and sustainable ownership across a distributed team.

What You'll Do

  • Observability. Build and operate metrics, logs, traces, and alerting capabilities. Help teams establish meaningful SLOs and use production telemetry to diagnose problems and verify improvements.

  • Incident Management. Own incident tooling and practices, coordinate cross-team response, and turn incident reviews into engineering improvements that reduce recovery time and repeat failures.

  • Load Testing. Build and maintain load/failure testing capabilities. Validate critical paths under expected demand, quantify headroom, and test recovery and production readiness with service owners.

  • Performance Engineering. Lead deep engagements with internal teams on SLOs and end-to-end performance. Use profiling, telemetry, and load tests to identify bottlenecks and deliver improvements with service owners-not just recommendations.

  • Stay technically engaged. Review designs and production changes, debug difficult failure modes, and use AI coding tools-including Replit-to prototype and automate. Apply rigorous review and verification to AI-generated changes.

  • Build and grow a high-ownership engineering team. Coach engineers, develop technical leaders, manage performance, and hire against agreed needs. Make distributed collaboration, mentoring, and backup coverage deliberate rather than relying on a few permanent escalation points.

  • Measure outcomes and close the loop. Track rollout safety, recovery time, repeat incidents, critical-path latency/throughput, test coverage, and improvements arising from cost/capacity analysis. Agree success measures and continuing ownership with partner teams.

What You'll Bring

  • Demonstrated engineering management. You have led and developed engineers, made prioritization and performance decisions, hired thoughtfully, and delivered through a team-not only acted as its strongest individual contributor.

  • Software-oriented production systems depth. You have built and operated distributed systems or reliability platforms and can reason across deployment behavior, Kubernetes, telemetry, service dependencies, and recovery mechanisms.

  • Safe-change and performance judgment. You have led consequential migrations or incidents and used measurement to diagnose reliability or performance problems. You can distinguish symptoms from causes and validate fixes under realistic conditions.

  • Platform-product and cross-team judgment. You can build capabilities other teams adopt, lead hands-on engagements without absorbing every service's operations, and make clear tradeoffs among reliability, performance, engineering effort, and cost.

Nice to Have

  • Experience with GitOps or progressive-delivery platforms such as Harness, ArgoCD, or Kargo.

  • Experience with observability, profiling, load-testing, and failure-testing systems, including OpenTelemetry or comparable tooling.

  • Experience with cloud cost attribution, capacity planning, and provider coordination, particularly on GCP.

  • Experience growing distributed teams and using AI tools to increase engineering output while preserving production safeguards.

Full-Time Employee Benefits Include:

Competitive Salary & Equity

401(k) Program with a 4% match (US Only)

Health, Dental, Vision and Life Insurance

Short Term and Long Term Disability

Paid Parental, Medical, Caregiver Leave

Flexible Time Off (FTO) + Holidays

Commuter Benefits (In-Office & US Only)

Monthly Wellness Stipend

Autonomous Work Environment

In Office Set-Up Reimbursement (In-Office Only)

Quarterly Team Gatherings

In Office Amenities (In-Office Only)

Want to learn more about what we are up to?

Interviewing + Culture at Replit

To achieve our mission of making programming more accessible around the world, we need our team to be representative of the world. We welcome your unique perspective and experiences in shaping this product. We encourage people from all kinds of backgrounds to apply, including and especially candidates from underrepresented and non-traditional backgrounds.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
787,996 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Leadership
Similar stack
Same company
Foster City
$159k – $255k per year • In office • TS/SCI • Full-Time • 10+ years exp • Bachelor's Degree • McLean • Washington
Apply
≈ $153k – $285k per year (Estimated) • Hybrid • Full-Time • 10+ years exp • Bachelor's Degree • Omaha
DevOps
Azure
Platform Engineering
Management
Agile
Apply
$110k – $150k per year • Remote (United States) • Full-Time • 3+ years exp
Apex
Apex
Lightning Web Components
MuleSoft
Visualforce
Databases
Databricks
AI/ML
Copilot
DevOps
Platform Engineering
Analytics
ETL/ELT
Management
Agile
Scrum
Apply
$199k – $348k per year • Equity • Hybrid • Full-Time • Waltham
AI/ML
AI Agents
LLM
RAG
LLM Guardrails
Multi-Agent Systems
Management
ServiceNow
Apply
$199k – $348k per year • Equity • Hybrid • Full-Time • San Diego
AI/ML
AI Agents
LLM
RAG
LLM Guardrails
Multi-Agent Systems
Management
ServiceNow
Apply
≈ $88k – $170k per year (Estimated) • Hybrid • Full-Time • 5+ years exp • Bachelor's Degree • Dublin
AI/ML
Hadoop
Model Context Protocol
AI Agents
LLM
Machine Learning
DevOps
CI/CD
Git
AWS
Docker
Kubernetes
Cybersecurity
GDPR
HIPAA
Apply
$121k – $171k per year • Hybrid • Full-Time • 8+ years exp • Bachelor's Degree • Mississauga
JavaScript
Java
TypeScript
SQL
Java
Maven
Spring Boot
Spring Cloud
Gradle
Databases
Oracle
ElasticSearch
Apache Kafka
AI/ML
Spark
Frontend
Angular
React.js
DevOps
GCP
Kibana
Logstash
Jenkins
AWS
Docker
Kubernetes
Analytics
ETL/ELT
Informatica
Talend
Management
Agile
Scrum
Apply
$177k – $265k per year • Hybrid • Full-Time • 10+ years exp • Jersey City • New York
Python
Java
C++
AI/ML
LangChain
Prompt Engineering
AI Agents
RAG
Agentic Workflows
DevOps
Rest API
CI/CD
Management
Agile
Apply
$107k – $161k per year • Hybrid • Full-Time • 5+ years exp • Bachelor's Degree • Irving • Tampa
JavaScript
Java
Databases
Oracle
MS SQL
Frontend
React.js
DevOps
Rest API
CI/CD
Docker
Kubernetes
SOAP
QA
Mocha
Apply
≈ $167k – $311k per year (Estimated) • Hybrid • Full-Time • 12+ years exp • Bachelor's Degree • Dallas • Malvern
AI/ML
AI Agents
DevOps
SRE
Platform Engineering
Cybersecurity
Threat Modeling
Management
Confluence
Agile
Apply
$250k – $325k per year • Hybrid • Full-Time • Foster City
AI/ML
AI Agents
Replit
DevOps
Terraform
GCP
Istio
Envoy
Kubernetes
Cloudflare
Service Mesh
Google GKE
Apply
$200k – $300k per year • Hybrid • Full-Time • 4+ years exp • Foster City
JavaScript
AI/ML
AI Agents
Replit
Frontend
React.js
Design
Figma
Apply
$130k – $290k per year • Hybrid • Full-Time • Foster City
Go
Rust
AI/ML
AI Agents
Replit
DevOps
GCP
Kubernetes
Google GKE
Google Cloud Run
Linux
Apply
$250k – $360k per year • Hybrid • Full-Time • Foster City
AI/ML
Function Calling
AI Agents
LLM
Replit
Recommender Systems
Agentic Workflows
Tool Use
Apply
$110k – $140k per year • Hybrid • Full-Time • New York
AI/ML
Claude
ChatGPT
AI Agents
Replit
DevOps
Self-Healing
Management
Slack
Apply
$250k – $325k per year • Hybrid • Full-Time • Foster City
AI/ML
AI Agents
Replit
DevOps
Terraform
GCP
Istio
Envoy
Kubernetes
Cloudflare
Service Mesh
Google GKE
Apply
$243k – $315k per year • In office • Full-Time • 14+ years exp • Bachelor's Degree • Foster City
AI/ML
Model Context Protocol
Multimodal AI
Function Calling
AI Agents
RAG
Semantic Search
LLMOps
A2A
Human-in-the-Loop
Structured Outputs
Semantic Search
Knowledge Graph
LLM Evaluation
LLM Guardrails
Tool Use
Machine Learning
DevOps
CI/CD
Platform Engineering
Cybersecurity
Least Privilege
Apply
$204k – $280k per year • In office • Full-Time • 10+ years exp • Bachelor's Degree • Foster City
Python
SQL
AI/ML
Anomaly Detection
Machine Learning
Apply
$137k – $213k per year • In office • Full-Time • 3+ years exp • Bachelor's Degree • Foster City
Python
AI/ML
Copilot
Spark
ChatGPT
Scikit-learn
TensorFlow
PyTorch
NeuralProphet
Time Series Forecasting
Recommender Systems
Machine Learning
Apply
$219k – $351k per year • In office • Full-Time • 12+ years exp • Bachelor's Degree • Foster City
DevOps
Kubernetes
Platform Engineering
Incident Management
Apply
See all jobs
This is one of many
787,996 more open roles from verified company boards, updated every day.