573,000open jobs
23,970companies
78,613added this week
Browse all
Salary
$138k – $304k per year (Estimated)
Location
In office (Palo Alto)
Seniority
Senior · 5+ years exp
Employment
Full-Time
Overview
Company
Impact
Profile match

Palo Alto, CA | Full-Time | On-site

About Nace AI:

Nace AI is an enterprise AI product and research company in Palo Alto (backed by General Catalyst, Walden Catalyst, and Intel). We build long-running AI agents powered by our own specialized SLMs - we started with financial audit and accounting workflows and are expanding from there. Real enterprise deployments, not demos.

Role Overview:

As a Senior MLOps Engineer, you will own the infrastructure that takes Nace.AI 's models from research to reliable, production-grade systems. Our infrastructure generates task-specific Small Language Models (SLMs) in real time - which means our training, serving, and evaluation infrastructure isn't an afterthought; it is the product. You will design and operate the pipelines, orchestration, and serving layers that allow us to train, deploy, monitor, and continuously improve many specialized models at once, with the reliability that high-stakes audit, compliance, and finance workflows demand. This role sits at the intersection of ML engineering, LLM inference infrastructure, and platform reliability, and requires both strong systems instincts and hands-on execution.

Key Responsibilities:

  • Design, build, and operate end-to-end ML infrastructure: training orchestration, experiment tracking, model registries, CI/CD for models, and automated evaluation pipelines.

  • Own LLM/SLM serving infrastructure - scale low-latency, high-throughput inference using frameworks like vLLM, including batching, caching, and autoscaling strategies.

  • Build and manage multi-GPU training and inference clusters (scheduling, utilization, cost optimization) across cloud and on-prem environments.

  • Implement observability for models in production: latency, throughput, drift, regression, and quality monitoring with actionable alerting.

  • Apply inference-time optimizations - quantization (AWQ, GPTQ, FP8/GGUF), distillation support, KV-cache management, and deployment tuning - in partnership with our ML and Research Engineers.

  • Harden our stack for enterprise deployment: reproducibility, versioning, access controls, and audit-ready traceability of model behavior.

  • Set MLOps best practices and tooling standards as an early, senior member of the infrastructure team.

Qualifications:

  • 5+ years of experience in MLOps, ML infrastructure, or platform engineering, with substantial production ownership.

  • Proven experience deploying and scaling LLM, inference infrastructure in production, including model serving frameworks such as TRT, vLLM, SGLang or TGI.

  • Strong proficiency with Kubernetes, containerization (Docker), and infrastructure-as-code (Terraform or similar).

  • Hands-on experience with GPU cluster management and distributed training/serving environments.

  • Proficient in Python with a strong track record of building substantial, maintainable systems.

  • Experience with ML pipeline and orchestration tooling (e.g., Airflow, Kubeflow, Ray, MLflow, Weights & Biases).

  • Solid foundation in computer science fundamentals and cloud architecture (AWS, GCP, or Azure).

  • BS degree in CS or related technical field.

  • Self-starter comfortable working in a fast-paced, dynamic environment.

Preferred Qualifications:

  • MS in CS or related technical field.

  • Experience operating multi-node GPU training infrastructure.

  • Hands-on experience with quantization techniques (AWQ, GPTQ, FP8/GGUF) and other inference-time optimizations.

  • Familiarity with data processing stacks such as Spark and Airflow.

  • Experience supporting fine-tuning workflows for LLMs/VLMs (instruction tuning, RLHF/DPO pipelines).

  • Experience in regulated or enterprise environments where reliability, security, and auditability are first-class requirements.

  • Contributor to open-source ML infrastructure projects.

Why Nace AI?

  • Pedigree: Work with a team from top-tier institutions and companies, backed by the best VCs in the world.

  • Impact: You are joining early enough to shape the infrastructure foundations of a company aiming to be the "OS" for professional knowledge.

  • Competitive Package: Silicon Valley-standard salary, significant equity, and premium benefits.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
573,000 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
Palo Alto
$52k – $77k per year • In office • Full-Time • Bachelor's Degree • Galway
Python
Bash
DevOps
Terraform
Ansible
Red Hat
Dynatrace
CI/CD
AWS
SRE
Configuration Management
Amazon EC2
Amazon S3
Amazon CloudWatch
Apply
$150k – $250k per year • In office • Full-Time • 2+ years exp • Singapore
Python
AI/ML
Reinforcement Learning
AI Agents
DevOps
Docker
Apply
$150k – $250k per year • In office • Full-Time • 2+ years exp • Singapore
Python
AI/ML
Reinforcement Learning
LLM
Post-training
DevOps
Docker
Apply
$46k – $117k per year (Estimated) • In office • Internship • Bachelor's Degree • Córdoba
Python
Verilog
C++
MATLAB
AI/ML
Copilot
ChatGPT
Apply
$74k – $191k per year (Estimated) • In office • Internship • Bachelor's Degree • Plantation
Python
C++
Perl
MATLAB
DevOps
Git
Management
Agile
Scrum
Apply
VP of Engineering 25 days ago
$166k – $370k per year (Estimated) • In office • Full-Time • 10+ years exp • Palo Alto
AI/ML
AI Agents
LLM
Cybersecurity
SOC 2
Apply
$111k – $250k per year (Estimated) • In office • Full-Time • 4+ years exp • Bachelor's Degree • Palo Alto
AI/ML
AI Agents
LLM
Management
Linear
Notion
Jira
Apply
$120k – $255k per year (Estimated) • In office • Full-Time • 5+ years exp • Palo Alto
JavaScript
AI/ML
Cursor
Claude
OpenAI Codex
Frontend
React.js
Apply
$91k – $195k per year (Estimated) • In office • Full-Time • 5+ years exp • Palo Alto
AI/ML
AI Agents
Apply
$120k – $140k per year • In office • Full-Time • Palo Alto
AI/ML
AI Agents
Apply
$131k – $205k per year • Remote/Hybrid • Full-Time • 10+ years exp • Bachelor's Degree • Palo Alto
Apply
$88k – $210k per year (Estimated) • In office • 6+ years exp • Palo Alto
DevOps
Terraform
GCP
AWS
Akamai
Management
Agile
Scrum
Apply
$159k per year • In office • Full-Time • 4+ years exp • Palo Alto
Apply
$131k – $164k per year • Equity • In office • Full-Time • 5+ years exp • Palo Alto
AI/ML
Multimodal AI
Apply
$150k – $200k per year • Equity • Remote/Hybrid • 3+ years exp • Bachelor's Degree • Palo Alto
Go
JavaScript
C++
AI/ML
Edge AI
Frontend
npm
DevOps
AWS
Docker
Kubernetes
Bazel
AWS Lambda
API Gateway
Apply
See all jobs
This is one of many
573,000 more open roles from verified company boards, updated every day.