599,985open jobs
30,608companies
86,410added this week
Browse all
Salary
$161k – $308k per year (Estimated)
Location
Remote/Hybrid (San Francisco, United States)
Seniority
Senior · 5+ years exp
Employment
Full-Time
Overview
Company
Impact
Profile match

ML Ops Engineer - Agentic AI Lab (Founding Team)

Location: San Francisco Bay Area

Type: Full-Time

Compensation: Competitive salary + meaningful equity (founding tier)

Backed by 8VC, we're building a world-class team to tackle one of the industry’s most critical infrastructure problems.

About the Role

Our AI Lab is pioneering the future of intelligent infrastructure through open-source LLMs, agent-native pipelines, retrieval-augmented generation (RAG), and knowledge-graph-grounded models.

We’re hiring an ML Ops Engineer to be the glue between ML research and production systems - responsible for automating the model training, deployment, versioning, and observability pipelines that power our agents and AI data fabric.

You’ll work across compute orchestration, GPU infrastructure, fine-tuned model lifecycle management, model governance, and security e

Responsibilities

  • Build and maintain secure, scalable, and automated pipelines for:

  • LLM fine-tuning, SFT, LoRA, RLHF, DPO training

  • RAG embedding pipelines with dynamic updates

  • Model conversion, quantization, and inference rollout

  • Manage hybrid compute infrastructure (cloud, on-prem, GPU clusters) for training and

    inference workloads using Kubernetes, Ray, and Terraform

  • Containerize models and agents using Docker, with reproducible builds and CI/CD via

    GitHub Actions or ArgoCD

  • Implement and enforce model governance: versioning, metadata, lineage, reproducibility,

    and evaluation capture

  • Create and manage evaluation and benchmarking frameworks (e.g. OpenLLM-Evals,

    RAGAS, LangSmith)

  • Integrate with security and access control layers (OPA, ABAC, Keycloak) to enforce

    model policies per tenant

  • Instrument observability for model latency, token usage, performance metrics, error

    tracing, and drift detection

  • Support deployment of agentic apps with LangGraph, LangChain, and custom inference

    backends (e.g. vLLM, TGI, Triton)

Desired Experience

Model Infrastructure:

  • 4+ years in MLOps, ML platform engineering, or infra-focused ML roles

  • Deep familiarity with model lifecycle management tools: MLflow, Weights & Biases, DVC,

  • HuggingFace Hub

  • Experience with large model deployments (open-source LLMs preferred): LLaMA,

  • Mistral, Falcon, Mixtral

  • Comfortable with tuning libraries (HuggingFace Trainer, DeepSpeed, FSDP, QLoRA)

  • Familiarity with inference serving: vLLM, TGI, Ray Serve, Triton Inference Server

Automation + Infra:

  • Proficient with Terraform, Helm, K8s, and container orchestration

  • Experience with CI/CD for ML (e.g. GitHub Actions + model checkpoints)

  • Managed hybrid workloads across GPU cloud (Lambda, Modal, HuggingFace Inference,

  • Sagemaker)

  • Familiar with cost optimization (spot instance scaling, batch prioritization, model sharding)

Agent + Data Pipeline Support:

Familiarity with LangChain, LangGraph, LlamaIndex or similar RAG/agent orchestration tools

Built embedding pipelines for multi-source documents (PDF, JSON, CSV, HTML)

Integrated with vector databases (Weaviate, Qdrant, FAISS, Chroma)

Security & Governance:

Implemented model-level RBAC, usage tracking, audit trails

Integrated with API rate limits, tenant billing, and SLA observability

Experience with policy-as-code systems (OPA, Rego) and access layers

Preferred Stack

  • LLM Ops: HuggingFace, DeepSpeed, MLflow, Weights & Biases, DVC

  • Infra: Kubernetes (GKE/EKS), Ray, Terraform, Helm, GitHub Actions, ArgoCD

  • Serving: vLLM, TGI, Triton, Ray Serve

  • Pipelines: Prefect, Airflow, Dagster

  • Monitoring: Prometheus, Grafana, OpenTelemetry, LangSmith

  • Security: OPA (Rego), Keycloak, Vault

  • Languages: Python (primary), Bash, optionally Rust or Go for tooling

Mindset & Culture Fit

  • Builder's mindset with startup autonomy: you automate what slows you down

  • Obsessive about reproducibility, observability, and traceability

  • Comfortable with a hybrid team of AI researchers, DevOps, and backend engineers

  • Interested in aligning ML systems to product delivery, not just papers

  • Bonus: experience with SOC2, HIPAA, or GovCloud-grade model operations

What We’re Looking For

Experience:

  • 5+ years as a full stack or backend engineer

  • Experience owning and delivering production systems end-to-end

  • Prior experience with modern frontend frameworks (React, Next.js)

  • Familiarity with building APIs, databases, cloud infrastructure, or deployment workflows at scale

  • Comfortable working in early-stage startups or autonomous roles, prior experience as a founder, founding engineer, or a 0-1 pre-seed startup is a big plus

Mindset:

  • Comfortable with ambiguity, eager to prototype and iterate quickly

  • Strong sense of ownership - prefers to build systems rather than wait for tickets

  • Enjoys thinking about architecture, performance, and tradeoffs at every level

  • Clear communicator and pragmatic team player

  • Values equity and impact over prestige or hierarchy

  • Prior startup or founding team experience

Why This Role Matters

Your work will enable models and agents to be trained, evaluated, deployed, and governed at

scale - across many tenants, models, and tasks. This is the backbone of a secure, reliable,

and scalable AI-native enterprise system. If you dream about using AI to solve some really hard

real world problems - we would love to hear from you.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
599,985 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
San Francisco
$60k – $109k per year (Estimated) • In office
Python
Go
Java
Scala
Databases
ElasticSearch
AI/ML
Prompt Engineering
LLM
DevOps
Terraform
CI/CD
Docker
Kubernetes
Platform Engineering
Apply
$44k – $101k per year (Estimated) • In office
Python
Go
Java
Scala
Databases
ElasticSearch
AI/ML
Prompt Engineering
LLM
DevOps
Terraform
CI/CD
Docker
Kubernetes
Platform Engineering
Apply
$51k – $117k per year (Estimated) • In office
Python
Go
Java
Scala
Databases
ElasticSearch
AI/ML
Prompt Engineering
LLM
DevOps
Terraform
CI/CD
Docker
Kubernetes
Platform Engineering
Apply
$35k – $79k per year (Estimated) • In office
Python
Go
Java
Scala
Databases
ElasticSearch
AI/ML
Prompt Engineering
LLM
DevOps
Terraform
CI/CD
Docker
Kubernetes
Platform Engineering
Apply
$23k – $54k per year (Estimated) • In office • Full-Time • 5+ years exp • Bachelor's Degree • Bengaluru
Python
Frontend
Lighthouse
Apply
$180k – $377k per year (Estimated) • In office • Full-Time • PhD • San Francisco
AI/ML
Fine-tuning
Reinforcement Learning
AI Agents
PyTorch
Tokenization
Hugging Face
Knowledge Graph
LLM Evaluation
Apply
Founding Designer 9 months ago
$144k – $271k per year (Estimated) • In office • Full-Time • 5+ years exp • San Francisco
AI/ML
AI Agents
Apply
$93k – $193k per year (Estimated) • In office • Full-Time • 3+ years exp • San Francisco
Python
SQL
AI/ML
AI Agents
Knowledge Graph
Apply
$156k – $337k per year (Estimated) • In office • Full-Time • San Francisco
Python
Rust
SQL
Databases
PostgreSQL
Weaviate
Neo4j
Chroma
Delta Lake
DuckDB
Pinecone
FAISS
Qdrant
AI/ML
LangGraph
LangChain
DeepSpeed
LlamaIndex
LoRA
vLLM
Fine-tuning
RLHF
Reinforcement Learning
Quantization
AI Agents
AutoGPT
Cohere SDK
LangSmith
PEFT
QLoRA
Ragas
TGI
AgentOps
Falcon
Mistral
Transformers
LLM
RAG
TruLens
BabyAGI
Ray
Reranking
Mixtral
Hugging Face
Amazon SageMaker
DPO
SFT
PPO
FSDP
Knowledge Graph
LLM Evaluation
Interpretability
DevOps
Kubernetes
Vector
Apply
$157k – $317k per year (Estimated) • In office • Full-Time • 10+ years exp • San Francisco
Databases
Weaviate
Pinecone
Qdrant
AI/ML
LangGraph
LangChain
vLLM
AI Agents
TGI
LLM
RAG
Multi-Agent Systems
DevOps
Terraform
GCP
Helm
GitHub Actions
OpenTelemetry
Prometheus
Pulumi
Azure
CI/CD
GitOps
ArgoCD
AWS
Docker
Kubernetes
Grafana
Platform Engineering
Amazon EKS
Google GKE
Vector
Progressive Delivery
IAM
QA
Sentry
Apply
Senior Data Analyst 6 hours ago
$120k – $243k per year (Estimated) • In office • Full-Time • 9+ years exp • Bachelor's Degree • San Francisco
Python
SQL
Analytics
Tableau
Metabase
Looker
Apply
$109k – $230k per year (Estimated) • In office • Full-Time • 3+ years exp • San Francisco
Design
Figma
Sketch
Adobe XD
Apply
$150k – $250k per year • In office • Full-Time • 2+ years exp • Bachelor's Degree • San Francisco
TypeScript
Databases
PostgreSQL
AI/ML
AI Agents
DevOps
AWS
Kubernetes
Apply
$100k – $180k per year • In office • Full-Time • San Francisco
Python
TypeScript
AI/ML
Computer Vision
Frontend
Tailwind CSS
Apply
$150k – $200k per year • In office • Full-Time • 8+ years exp • San Francisco
DevOps
Azure
Cybersecurity
Okta
Management
Slack
ServiceNow
Microsoft Teams
Marketing
LinkedIn
Apply
See all jobs
This is one of many
599,985 more open roles from verified company boards, updated every day.