368,530open jobs
9,432companies
50,439added this week
Browse all
Salary
$142k – $255k per year (Estimated)
Location
Remote (United States)
Seniority
Staff
Employment
Contractor
Overview
Company
Impact
Profile match
Tech Holding is a global technology solutions provider headquartered in Los Angeles, California, and founded in 2021. The company offers a range of services including professional consulting, managed IT services, and technical staffing solutions to support digital transformation and software engineering projects. It operates internationally, serving diverse industries by providing specialized talent and technological strategies to optimize business operations.

About us:

Working at Tech Holding isn't just a job, it's an opportunity to be a part of something bigger. We are a full-service consulting firm that was founded on the premise of delivering predictable outcomes and high-quality solutions to our clients.  Our founders and team members have industry experience and have held senior positions in a wide variety of companies - from emerging startups to large Fortune 50 firms - and we have taken our combined experiences and developed a unique approach that is supported by the principles of deep expertise, integrity, transparency, and dependability.

The Role:

We are looking for a hands-on Lead Site Reliability Engineer for a project based assignment to establish and continuously improve the performance, reliability, and scalability of our platform.This role will define what the platform can reliably sustain today, identify where constraints will emerge, and ensure the organization is prepared to scale before demand arrives.

This is not a traditional DevOps role or an advisory architecture position. You will work directly across application services, infrastructure, databases, networking, caching, queues, external dependencies, and operational processes to identify bottlenecks, validate system limits, and lead remediation.

You will partner closely with engineering, product, and leadership to provide clear, evidence-based answers around capacity, performance, reliability, risk, and the cost of scaling.

Key Responsibilities:

  • Establish performance, throughput, latency, and capacity baselines for critical customer and platform workflows
  • Define and maintain SLOs, error budgets, performance budgets, dashboards, alerts, and reliability thresholds
  • Instrument and analyze the full request path across application services, compute, storage, networking, databases, caches, queues, DNS, registry dependencies, and third-party services
  • Identify system bottlenecks and lead cross-functional remediation efforts with engineering teams
  • Build capacity models that show what the platform can sustain, where constraints will emerge, and what additional scale will cost
  • Lead load, stress, soak, spike, failure, and recovery testing in representative environments
  • Develop realistic demand scenarios for major customers, partnerships, pilots, and high-volume events
  • Drive architecture hardening, graceful degradation, dependency-failure planning, and resilience improvements
  • Partner with Test Automation and Scalability Engineering to establish automated performance testing, regression coverage, and production release gates
  • Own technical readiness assessments for major pilots, partnerships, and production launches
  • Create operational runbooks for scale-up events, incidents, rollback, recovery, and dependency failures
  • Lead performance and reliability investigations during incidents and ensure lessons are incorporated into future engineering work
  • Make infrastructure cost, performance, and reliability tradeoffs visible to engineering and executive leadership
  • Recommend capacity and reliability investments before they become production constraints

Required Skills:

  • Significant experience in Site Reliability Engineering, performance engineering, platform engineering, distributed systems, or a closely related engineering discipline
  • Experience supporting production systems with meaningful scale, traffic, latency, or availability requirements
  • Deep understanding of observability, performance analysis, capacity planning, and reliability engineering
  • Strong hands-on experience with cloud infrastructure and production distributed systems
  • Deep knowledge of databases, networking, caching, queueing, compute, storage, and common distributed-system failure modes
  • Experience defining and operating against SLOs, SLIs, error budgets, and production reliability metrics
  • Hands-on experience performing load, stress, soak, scalability, and resilience testing
  • Ability to profile systems, diagnose bottlenecks, tune architecture, and work directly with engineering teams to implement improvements
  • Experience designing for graceful degradation, dependency failures, recovery, and high-demand scenarios
  • Strong incident management and root-cause analysis experience
  • Ability to translate technical performance and reliability risks into clear business implications for senior leadership
  • Strong judgment around when systems genuinely require optimization versus when additional complexity is premature

Nice to have:

  • Experience operating high-scale SaaS, identity, DNS, registry, infrastructure, or other highly distributed platforms
  • Experience creating capacity-cost models and forecasting infrastructure requirements
  • Experience building performance and reliability gates into CI/CD pipelines
  • Experience preparing platforms for significant increases in traffic associated with enterprise customers or strategic partnerships
  • Experience leading reliability or performance initiatives that span multiple engineering teams

What Success Looks Like

Within your first several months, you will have:

  • Established measurable throughput, latency, and capacity baselines for critical platform journeys
  • Defined initial SLOs, error budgets, dashboards, alerts, and performance thresholds
  • Identified the platform's most significant scalability and reliability constraints and created an actionable remediation roadmap
  • Validated representative high-scale scenarios through load, soak, stress, failure, and recovery testing
  • Developed a capacity and cost model showing how the platform can support significant increases in demand
  • Established production-readiness criteria and clear go/no-go evidence for major launches and partnerships
  • Created repeatable scale-up, incident, rollback, and dependency-failure runbooks
  • Given leadership a clear, evidence-based understanding of the platform's current capacity envelope and future scaling requirements

Employment type:

  • Contract

* Applicants must be authorized to work for ANY employer in the U.S. We are unable to sponsor or take over sponsorship of an employment Visa at this time

Tech Holding is proud to be an Equal Opportunity Employer and is committed to fostering a diverse and inclusive workplace. We welcome applicants from all backgrounds and experiences, and we consider qualified applicants without regard to race, color, religion, gender, sexual orientation, gender identity, national origin, disability, veteran status, or any other legally protected characteristic. If you require accommodation in the application process, please contact our HR 

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
368,530 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
In your city
$122k – $270k per year (Estimated) • Remote • Full-Time • 8+ years exp • Bachelor's Degree • Toronto
AI/ML
Anomaly Detection
DevOps
AIOps
Dynatrace
Incident Management
OpenTelemetry
Platform Engineering
Management
ServiceNow
Apply
$20k – $54k per year (Estimated) • In office • Full-Time • Pune
Java
SQL
DevOps
GCP
Incident Management
Kubernetes
IAM
QA
JMeter
Selenium
Apply
Data Analyst 5B 1 day ago
$11k – $26k per year (Estimated) • In office • Full-Time • 3+ years exp • Bachelor's Degree • Noida
Databases
Databricks
Oracle
AI/ML
Spark
AI Agents
Edge AI
DevOps
Incident Management
Apply
$42k – $104k per year (Estimated) • In office • Internship • 5+ years exp
Bash
PowerShell
Python
DevOps
Amazon EC2
Amazon EKS
Ansible
AWS
AWS Lambda
Azure
CentOS Stream
Chef
CI/CD
CloudFormation
Configuration Management
Datadog
FinOps
GCP
Git
GitHub Actions
GitLab CI
Grafana
Hyper-V
Istio
Jenkins
Kubernetes
KVM
Linkerd
Platform Engineering
Prometheus
Puppet
Service Mesh
Splunk
Terraform
Ubuntu
VMWare
Windows Server
Amazon CloudWatch
Amazon ECS
Amazon S3
API Gateway
AWS Step Functions
GitHub
GitLab
IAM
Cybersecurity
GDPR
ISO 27001
SOC 2
Apply
$54k – $159k per year (Estimated) • In office • Full-Time • 4+ years exp • Bachelor's Degree • Singapore
Bash
PowerShell
Python
DevOps
Ansible
AWS
Azure
Chef
Docker
Hyper-V
Incident Management
Kubernetes
Puppet
Red Hat
Terraform
VMWare
Windows Server
Apply
$28k – $55k per year (Estimated) • In office • 5+ years exp • Bachelor's Degree • Ahmedabad
Python
SQL
Databases
Amazon Redshift
Apache Kafka
Snowflake
DevOps
AWS
CI/CD
GCP
Amazon Kinesis
Amazon S3
GitHub
GitLab
Analytics
Tableau
ETL/ELT
Apply
$30k – $73k per year (Estimated) • Remote/Hybrid • 5+ years exp • Gurgaon
Node JS
TypeScript
JavaScript
Java
Java
Spring Boot
Spring Cloud
Databases
Apache Kafka
Frontend
Angular
DevOps
AWS
Azure
Docker
Jenkins
Kubernetes
WebRTC
Management
ServiceNow
Marketing
Salesforce
Apply
$33k – $80k per year (Estimated) • Remote/Hybrid • 5+ years exp • Bengaluru
Node JS
TypeScript
JavaScript
Java
Java
Spring Boot
Spring Cloud
Databases
Apache Kafka
Frontend
Angular
DevOps
AWS
Azure
Docker
Jenkins
Kubernetes
WebRTC
Management
ServiceNow
Marketing
Salesforce
Apply
$137k – $279k per year (Estimated) • In office • Full-Time • Los Angeles
Node JS
TypeScript
JavaScript
AI/ML
Claude
LLM
DevOps
Azure
Apply
$23k – $57k per year (Estimated) • In office • 7+ years exp • Ahmedabad
Python
DevOps
Amazon EKS
Ansible
AWS
Azure
Azure AKS
CI/CD
CloudFormation
Configuration Management
Docker
GCP
GitHub Actions
Google GKE
Jenkins
Kubernetes
Terraform
Platform Engineering
Amazon ECS
GitHub
Apply
See all jobs
This is one of many
368,530 more open roles from verified company boards, updated every day.