368,530open jobs
9,432companies
50,439added this week
Browse all
Salary
$109k – $211k per year (Estimated)
Location
In office (Washington)
Seniority
Senior
Overview
Company
Impact
Profile match
Tiger Analytics is a global data science and AI consulting firm headquartered in Silicon Valley, California. The company specializes in building customized data engineering, machine learning, and advanced analytics solutions for major enterprises across industries such as financial services, healthcare, retail, and manufacturing. With a global presence spanning the US, India, the UK, and Singapore, it helps Fortune 1000 companies transform complex data into actionable business value at scale.

Role Overview

We are seeking a high-caliber Site Reliability Engineer (SRE) to join our Forward Engineering team. You will be the guardian of our production ecosystems, ensuring that our complex, data-driven AI platforms remain resilient, scalable, and highly performant. This role is a hybrid of software engineering and systems architecture, with a specialized focus on MLOps -bridging the gap between model development and production-grade reliability.

Key Responsibilities

1. Reliability & Performance Engineering

  • SLA/SLO Management: Define, monitor, and maintain Service Level Objectives (SLOs) and Service Level Indicators (SLIs) for critical AI/ML services.
  • Error Budgeting: Manage error budgets to balance the velocity of feature releases from the ML team with the stability of the production environment.
  • Scalability: Architect and manage auto-scaling strategies for Kubernetes (GKE) to handle fluctuating workloads during model training and high-volume inference.

2. MLOps & AI Infrastructure

  • Model Serving Reliability: Ensure the high availability of Vertex AI endpoints and custom inference services.
  • GPU/TPU Optimization: Monitor and optimize compute resource utilization (accelerators) to ensure cost-efficient performance for Large Language Models (LLMs).
  • Pipeline Resilience: Support and stabilize ML pipelines (Vertex AI Pipelines/Kubeflow) to ensure seamless data flow from ingestion to model retraining.

3. Automation & Orchestration (Eliminating "Toil")

  • Infrastructure as Code (IaC): Use Terraform or Pulumi to provision and manage consistent, version-controlled cloud environments.
  • CI/CD & GitOps: Design and optimize robust deployment pipelines for both application code and ML models using GitHub Actions, Cloud Build, or ArgoCD.
  • Task Automation: Develop custom Python or Go scripts to automate repetitive operational tasks, self-healing mechanisms, and resource cleanup.

4. Monitoring, Alerting & Incident Response

  • Observability: Build and manage comprehensive dashboards using Prometheus, Grafana, or Google Cloud Operations Suite (Stackdriver).
  • Incident Management: Act as a primary responder in on-call rotations, leading the technical resolution of production outages.
  • Blameless Post-Mortems: Conduct deep-dive root cause analysis (RCA) to ensure systemic issues are identified and permanently remediated through code.

Requirements

Orchestration: Expert-level knowledge of Kubernetes (K8s) and Docker.

MLOps Stack: Familiarity with tools such as Kubeflow, Vertex AI, MLflow, or DVC.

Scripting: Strong proficiency in Python (for automation) and Bash; knowledge of Go is a plus.

Data Systems: Experience managing the reliability of data-heavy services (BigQuery, Pub/Sub, or Vector Databases like Pinecone/Milvus).

Networking: Solid understanding of VPCs, Load Balancers, DNS, and secure service mesh (Istio/Anthos).

Benefits

Benefits

Significant career development opportunities exist as the company grows. The position offers a unique opportunity to be part of a small, fast-growing, challenging and entrepreneurial environment, with a high degree of individual responsibility.

Tiger Analytics provides equal employment opportunities to applicants and employees without regard to race, color, religion, age, sex, sexual orientation, gender identity/expression, pregnancy, national origin, ancestry, marital status, protected veteran status, disability status, or any other basis as protected by federal, state, or local law.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
368,530 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
Washington
$20k – $49k per year (Estimated) • Remote • Full-Time • Bachelor's Degree
Python
Ruby
SQL
Databases
Amazon Neptune
Neo4j
AI/ML
Hallucination
LangChain
LangGraph
LLM
Model Context Protocol
Spark
AI Agents
LLM Guardrails
DevOps
Ansible
AWS
Azure
CI/CD
Docker
GCP
GitHub Actions
GitLab CI
Jenkins
Kubernetes
Rest API
Terraform
GitHub
GitLab
QA
Playwright
Postman
Selenium
Swagger
Apply
$41k – $103k per year (Estimated) • Remote • Full-Time • 5+ years exp
Go
Python
AI/ML
Reinforcement Learning
Edge AI
DevOps
Ansible
AWS
Azure
Chef
CI/CD
Docker
GCP
GitHub Actions
Google GKE
Jenkins
Kubernetes
Platform Engineering
Puppet
Terraform
GitHub
Apply
$80k – $169k per year (Estimated) • Equity • Remote • Full-Time • 5+ years exp
Go
Python
Databases
PostgreSQL
RabbitMQ
DevOps
Alertmanager
Ansible
Atlantis
Backstage
Chef
CI/CD
containerd
Docker
GCP
GitOps
Google GKE
Grafana
Helm
Incident Management
Kubernetes
Loki
Platform Engineering
Prometheus
Puppet
Terraform
Thanos
IAM
Cybersecurity
Checkov
SOC 2
Least Privilege
Apply
$112k – $217k per year (Estimated) • Equity • Remote • Full-Time • 5+ years exp
Go
Python
Databases
PostgreSQL
RabbitMQ
DevOps
Alertmanager
Ansible
Atlantis
Backstage
Chef
CI/CD
containerd
Docker
GCP
GitOps
Google GKE
Grafana
Helm
Incident Management
Kubernetes
Loki
Platform Engineering
Prometheus
Puppet
Terraform
Thanos
IAM
Cybersecurity
Checkov
SOC 2
Least Privilege
Apply
$42k – $106k per year (Estimated) • Equity • Remote • Full-Time • 5+ years exp
Go
Python
Databases
PostgreSQL
RabbitMQ
DevOps
Alertmanager
Ansible
Atlantis
Backstage
Chef
CI/CD
containerd
Docker
GCP
GitOps
Google GKE
Grafana
Helm
Incident Management
Kubernetes
Loki
Platform Engineering
Prometheus
Puppet
Terraform
Thanos
IAM
Cybersecurity
Checkov
SOC 2
Least Privilege
Apply
$109k – $220k per year (Estimated) • In office • Full-Time • Toronto
Databases
Amazon Neptune
Azure Cosmos DB
Neo4j
AI/ML
LLM
RAG
Knowledge Graph
Analytics
ETL/ELT
Apply
MLOps Lead Engineer 6 days ago
$128k – $273k per year (Estimated) • In office • Full-Time • 8+ years exp • Saint Louis
Python
Python
pySpark
Databases
Databricks
Delta Lake
AI/ML
MLFlow
Spark
DevOps
AWS
Azure
Azure DevOps
CI/CD
GitHub Actions
Platform Engineering
GitHub
Apply
$132k – $238k per year (Estimated) • In office • Full-Time • 10+ years exp
DevOps
CI/CD
SLI/SLO/SLA
Analytics
ETL/ELT
Apply
$130k – $259k per year (Estimated) • In office • Full-Time • 12+ years exp • Master's Degree • Jersey City
Python
SQL
Analytics
Power BI
Apply
$131k – $286k per year (Estimated) • In office • Full-Time
Python
SQL
Databases
Google Cloud Spanner
Milvus
pgvector
Pinecone
PostgreSQL
AI/ML
AutoGen
Kubeflow
LangChain
LlamaIndex
LLM
Prompt Engineering
PyTorch
RAG
Triton
Triton Inference Server
Vertex AI
vLLM
Google AI Studio
Hugging Face
LLMOps
TPU
TGI
DevOps
GCP
Google GKE
Kubernetes
Platform Engineering
Terraform
IAM
Apply
$81k – $122k per year • Remote/Hybrid • Full-Time • PhD • Atlanta • Washington
JavaScript
SQL
AI/ML
AI Agents
Agentforce
Marketing
Salesforce
Apply
$171k – $273k per year • In office • Full-Time • 8+ years exp • PhD • San Francisco • Washington
AI/ML
A2A
Agentforce
AI Agents
Model Context Protocol
DevOps
AWS
GCP
Marketing
Salesforce
Apply
Data Scientist 1 hour ago
$113k – $188k per year • In office • Full-Time • 5+ years exp • Bachelor's Degree • Arlington • Washington
Python
Databases
Databricks
DevOps
AWS
Azure
Analytics
ETL/ELT
Power BI
Apply
Data Scientist 1 hour ago
$98k – $163k per year • In office • Full-Time • 3+ years exp • Bachelor's Degree • Arlington • Washington
SQL
Databases
Databricks
DevOps
AWS
Azure
Analytics
ETL/ELT
Power BI
Apply
$99k – $150k per year • Remote • Full-Time • 5+ years exp • Bachelor's Degree • Washington
JavaScript
AI/ML
Agentforce
AI Agents
Marketing
HubSpot
Marketo
Salesforce
Apply
See all jobs
This is one of many
368,530 more open roles from verified company boards, updated every day.