368,634open jobs
9,437companies
50,578added this week
Browse all
Salary
$141k – $288k per year (Estimated)
Location
In office (Palo Alto)
Employment
Full-Time
Overview
Company
Impact
Profile match
Ollama is the easiest way to automate your work using open models, while keeping your data safe.

Ollama is the most popular way for developers to access open models. What started as an open-source, local-first runtime is now the largest developer network in the open-model ecosystem: 8.9 million monthly active developers and over 67,000+ community-built integrations. We're backed by Y Combinator, Benchmark, 8VC, and Theory Ventures.

Our team is small and talent dense. We're flat, low-ego, and fast-moving. We like people who are truth-seeking, passionate, design-driven, and who enjoy shipping code.

About the role

You'll work on the heart of Ollama - the local runtime that runs open models on developers' own machines. It loads models, manages memory, drives GPU acceleration across NVIDIA, AMD, Intel, Qualcomm, and Apple Silicon (including our MLX integration), and makes all of it feel instant. You'll work in Go and C/C++ and touch the model formats and inference engines underneath, shipping to macOS, Linux, and Windows across an enormous range of hardware.

What you'll do

  • Make open models run fast and reliably on consumer and enterprise hardware - from a MacBook Pro to server-grade NVIDIA GPUs.

  • Own pieces of the runtime: model loading & scheduling memory management, quantization, GPU hardware backends.

  • Integrate new model architectures and quantization formats so the latest open models work on day one.

  • Improve cold-start, time-to-first-token, and throughput

  • Partner with model labs and hardware vendors on early access and deep integrations.

  • Ship in the open: Ollama is open source, and you'll work with the community

Example projects

  • Add support for a new model family end-to-end - format parsing, weights loading, and the defaults that make it useful out of the box.

  • Cut cold-start for a popular model in half by streaming weights and lazy-loading layers.

  • Land a new quantization format so a 70B model runs on a single consumer GPU.

  • Wire up a new GPU backend and find a 2x throughput win with kernel selection and memory tuning.

  • Improve the "Auto" experience - picking the right model and settings for a machine's hardware without the user thinking about it.

You may be a fit if

  • You have strong systems fundamentals and are comfortable in Go, C, or C++

  • You've worked close to the metal - GPU compute, inference, game engines, databases, OS, or networking.

  • You care about performance and have profiled and optimized real workloads.

  • You're comfortable shipping to millions of users and handling the long tail of hardware and OS combinations.

  • Bonus: experience with model quantization, GPU programming (CUDA/Metal/SYCL), Apple MLX

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
368,634 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
Palo Alto
$39k – $83k per year (Estimated) • In office • Full-Time • 12+ years exp • Pune
C++
Go
Java
Java
Spring Boot
Databases
Apache Kafka
NATS
DevOps
CI/CD
Docker
Jenkins
Kubernetes
Rest API
Apply
$150k per year • In office • Full-Time • New York
C++
Python
Apply
In office • Full-Time • Bachelor's Degree • Gurgaon
C++
COBOL
Java
SQL
Apply
$18k – $46k per year (Estimated) • Remote/Hybrid • Saint Petersburg
C++
DevOps
CI/CD
QEMU
Apply
$17k – $43k per year (Estimated) • Remote • Moscow
C++
Java
Python
Java
Maven
DevOps
Ansible
CI/CD
Docker
Git
Graylog
HAProxy
Jenkins
Nginx
Prometheus
Zabbix
Apply
$142k – $290k per year (Estimated) • In office • Full-Time • Palo Alto
Go
AI/ML
Ollama
DevOps
GitHub
Apply
GTM Engineer 1 month ago
$131k – $235k per year (Estimated) • In office • Full-Time • Palo Alto
AI/ML
Ollama
Apply
$99k – $184k per year (Estimated) • In office • Full-Time • 3+ years exp • Palo Alto
AI/ML
Ollama
Apply
$139k – $284k per year (Estimated) • In office • Full-Time • Palo Alto
AI/ML
Ollama
DevOps
Kubernetes
Apply
$139k – $284k per year (Estimated) • In office • Full-Time • Palo Alto
Go
TypeScript
AI/ML
Ollama
Apply
Chief of Staff 8 hours ago
$120k – $150k per year • Equity 0.4–0.7% • In office • Full-Time • 3+ years exp • Palo Alto
AI/ML
AI Agents
Apply
$140k – $310k per year (Estimated) • Remote/Hybrid • Bachelor's Degree • Palo Alto
Databases
Apache Kafka
NATS
DevOps
AWS
Azure
CI/CD
Docker
GCP
Grafana
gRPC
Kubernetes
OpenTelemetry
Platform Engineering
Prometheus
Robotics
EtherCAT
IoT
MQTT
OPC UA
Apply
$139k – $294k per year (Estimated) • In office • Palo Alto
Python
AI/ML
Fine-tuning
Hybrid Search
LLM
Prompt Engineering
RAG
Human-in-the-Loop
Knowledge Graph
AI Agents
Function Calling
DevOps
AWS
Apply
$137k – $292k per year (Estimated) • In office • Palo Alto
Python
AI/ML
Hybrid Search
LLM
Prompt Engineering
RAG
Human-in-the-Loop
AI Agents
Function Calling
DevOps
AWS
Apply
$150k – $271k per year (Estimated) • Remote/Hybrid • Full-Time • 5+ years exp • Bachelor's Degree • Palo Alto
C++
Java
Python
Rust
AI/ML
LLM
Apply
See all jobs
This is one of many
368,634 more open roles from verified company boards, updated every day.