368,634open jobs
9,437companies
50,578added this week
Browse all
Salary
$139k – $284k per year (Estimated)
Location
In office (Palo Alto)
Employment
Full-Time
Overview
Company
Impact
Profile match
Ollama is the easiest way to automate your work using open models, while keeping your data safe.

Ollama is the most popular way for developers to access open models. What started as an open-source, local-first runtime is now the largest developer network in the open-model ecosystem: 8.9 million monthly active developers and over 67,000+ community-built integrations. We're backed by Y Combinator, Benchmark, 8VC, and Theory Ventures.

Our team is small and talent dense. We're flat, low-ego, and fast-moving. We like people who are truth-seeking, passionate, design-driven, and who enjoy shipping code.

About the role

You'll build Ollama’s cloud, a scalable inference platform that lets developers run large, capable open models in their workflow. You'll work on high-throughput, low-latency distributed systems - inference serving, GPU fleet management, routing, metering, and the platform that Pro, Max, Team, and Enterprise customers rely on to process trillions of tokens.

What you'll do

  • Build and scale the inference platform that serves every request from ollama.com.

  • Design the routing and capacity layer that places workloads across GPUs and regions for cost, latency, and availability.

  • Own multi-tenant infrastructure: isolation, quotas, usage metering, billing, and Pro/Max/team/enterprise tiering.

  • Build the reliability, observability, and cost controls for our team and customers

You may be a fit if

  • You have deep experience with high-throughput, low-latency distributed systems - inference serving, traffic routing, real-time data pipelines, or large-scale APIs.

  • You're comfortable with cost/performance tradeoffs at scale and have owned a production service end-to-end.

  • You've worked with Kubernetes, GPU scheduling, or inference infrastructure.

  • You think in terms of reliability, SLOs, and honest capacity planning.

  • Bonus: experience building an inference platform, GPU fleet management, or billing/metering for an AI service.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
368,634 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
Palo Alto
$195k – $264k per year • In office • Full-Time • 15+ years exp • Master's Degree • United States
Python
AI/ML
Amazon SageMaker
Keras
Kubeflow
MLFlow
PyTorch
Scikit-learn
TensorFlow
Vertex AI
XGBoost
DevOps
AWS
Azure
CI/CD
CloudFormation
Docker
GCP
Kubernetes
Terraform
Cybersecurity
FedRAMP
NIST 800-53
Apply
Fullstack QA 7 hours ago
$14k – $32k per year (Estimated) • In office • 3+ years exp • Moscow
JavaScript
Python
SQL
TypeScript
Databases
Apache Kafka
PostgreSQL
DevOps
GitLab
Jenkins
Kubernetes
Management
Bitrix24
Confluence
Jira
Apply
Senior Data Scientist 7 hours ago
$54k – $88k per year (Estimated) • Equity • Remote/Hybrid • Full-Time • 4+ years exp • Master's Degree • Warsaw
Python
SQL
Databases
Databricks
Google BigQuery
AI/ML
Spark
DevOps
Azure
GCP
GitHub
Kubernetes
Analytics
Power BI
Tableau
Management
Confluence
Jira
Apply
$18k – $48k per year (Estimated) • In office • Full-Time • 7+ years exp • Mumbai
JavaScript
SQL
TypeScript
Java
Java
Spring Boot
Frontend
Angular
DevOps
AWS
Azure
Kubernetes
Rest API
Apply
$89k – $212k per year (Estimated) • Remote/Hybrid • Full-Time • 6+ years exp • Rehovot
Bash
Python
DevOps
Amazon CloudWatch
Amazon EKS
AWS
CI/CD
Datadog
GitLab
GitLab CI
GitOps
Grafana
IAM
Kubernetes
PagerDuty
Pulumi
Terraform
Terragrunt
VictoriaMetrics
Apply
$142k – $290k per year (Estimated) • In office • Full-Time • Palo Alto
Go
AI/ML
Ollama
DevOps
GitHub
Apply
GTM Engineer 1 month ago
$131k – $235k per year (Estimated) • In office • Full-Time • Palo Alto
AI/ML
Ollama
Apply
$99k – $184k per year (Estimated) • In office • Full-Time • 3+ years exp • Palo Alto
AI/ML
Ollama
Apply
$141k – $288k per year (Estimated) • In office • Full-Time • Palo Alto
C++
Go
AI/ML
CUDA
CUDA Toolkit
MLX ML
Ollama
Quantization
Apply
$139k – $284k per year (Estimated) • In office • Full-Time • Palo Alto
Go
TypeScript
AI/ML
Ollama
Apply
Chief of Staff 8 hours ago
$120k – $150k per year • Equity 0.4–0.7% • In office • Full-Time • 3+ years exp • Palo Alto
AI/ML
AI Agents
Apply
$140k – $310k per year (Estimated) • Remote/Hybrid • Bachelor's Degree • Palo Alto
Databases
Apache Kafka
NATS
DevOps
AWS
Azure
CI/CD
Docker
GCP
Grafana
gRPC
Kubernetes
OpenTelemetry
Platform Engineering
Prometheus
Robotics
EtherCAT
IoT
MQTT
OPC UA
Apply
$139k – $294k per year (Estimated) • In office • Palo Alto
Python
AI/ML
Fine-tuning
Hybrid Search
LLM
Prompt Engineering
RAG
Human-in-the-Loop
Knowledge Graph
AI Agents
Function Calling
DevOps
AWS
Apply
$137k – $292k per year (Estimated) • In office • Palo Alto
Python
AI/ML
Hybrid Search
LLM
Prompt Engineering
RAG
Human-in-the-Loop
AI Agents
Function Calling
DevOps
AWS
Apply
$150k – $271k per year (Estimated) • Remote/Hybrid • Full-Time • 5+ years exp • Bachelor's Degree • Palo Alto
C++
Java
Python
Rust
AI/ML
LLM
Apply
See all jobs
This is one of many
368,634 more open roles from verified company boards, updated every day.