793,143open jobs
50,545companies
124,108added this week
Browse all
Salary
$180k – $250k per year
Location
In office (Sunnyvale)
Seniority
Senior · 5+ years exp
Employment
Full-Time

Confirmed on the employer's own hiring board on Sep 25, 2026. First seen by Alion on Jan 16, 2026.

Overview
Company
Impact
Profile match
RealmGuard Technology Resources AI Failure Index Blog News Research Whitepaper Book Demo About Realm Labs was founded by AI researchers from three traditions that converged on the same problem: adversarial machine learning, AI safety and explainability, and large-scale systems and supply-chain security.

Role Overview

We are hiring a Founding ML Infrastructure Engineer to own the end-to-end deployment, optimization, and operation of our suits of models in production.

This is a core founding role focused on building and operating production-grade LLM systems. You will apply deep knowledge of model internals to deploy, optimize, and run modern LLMs at scale, owning performance end-to-end across latency, throughput, and reliability.

You will design and operate the full ML serving stack from model artifacts to GPU execution, and work closely with Product and ML teams to ensure our models can support high QPS, strict SLAs, and production correctness.

This role is ideal for someone who deeply understands how LLMs work internally, but chooses to specialize in making them fast, stable, and production-ready.

About Realm Labs

Realm Labs is an AI trust and security startup. We help enterprises detect, debug, and prevent AI’s misbehaviors in production. We are backed by top VCs and serve some of the most iconic global enterprises.

Key Responsibilities

  • Own the end-to-end LLM inference stack, including:
    • Model loading and execution
    • GPU utilization and memory efficiency
    • Runtime performance tuning
    • Production deployment and scaling
  • Design and operate high-performance LLM serving systems using technologies such as:
    • vLLM, TensorRT / TensorRT-LLM, Triton Inference Server, SGLang
  • Optimize inference across:
    • Latency
    • Throughput (QPS)
    • GPU memory footprint
    • Cost efficiency
  • Work hands-on with PyTorch and TensorFlow models, including:
    • Model graph understanding
    • Attention mechanisms, KV cache behavior, batching strategies
    • Precision tradeoffs (FP16, BF16, INT8, etc.)
  • Build and maintain production-grade GPU services:
    • Multi-model serving
    • Autoscaling strategies
    • Fault isolation and graceful degradation
  • Collaborate with application and platform teams to:
    • Define serving APIs
    • Ensure correctness and safety of outputs
    • Debug production issues end-to-end
  • Build a reproducible model training and versioning system for customer deployments
  • Establish best practices for:
    • Model versioning
    • Rollouts and rollbacks
    • Performance benchmarking
    • Production validation

Expected Qualifications

  • 5+ years of professional experience in ML infrastructure, systems engineering, or production ML roles.
  • Strong software engineering fundamentals; ability to write robust, maintainable production code.
  • Deep hands-on experience with LLM inference infrastructure, including:
    • PyTorch (required)
    • TensorFlow (working knowledge)
  • Proven experience with GPU inference optimization, including:
    • TensorRT / TensorRT-LLM
    • vLLM
    • Triton Inference Server
    • SGLang or similar serving runtimes
  • Strong understanding of LLM internals, such as:
    • Transformer architectures
    • Attention and KV caching
    • Batching, streaming, and token-level generation
  • Experience running ML systems in production with high traffic and SLAs
  • Comfortable working in Linux-based, cloud production environments

Preferred Qualifications

  • Experience deploying LLMs on Kubernetes and GPU clusters.
  • Familiarity with CUDA, NCCL, or low-level GPU performance concepts.
  • Experience with:
    • Model sharding and parallelism strategies
    • Multi-GPU inference
    • Streaming inference systems
  • Knowledge of observability for ML systems (metrics, latency breakdowns, GPU monitoring).
  • Experience working at startups or owning systems with minimal abstraction layers.

Additional Information

  • This is a founding, high-ownership role with direct impact on core product capabilities.
  • You will be expected to build, run, and own systems end-to-end.
  • The role may include limited on-call responsibilities aligned with production ownership.

Compensation & Benefits

  • Market aligned compensation and benefits
  • Founding engineer equity (Equity is a significant component of this role and will be discussed)
  • Medical, Dental, Vision, Life insurance, 401-K, In-office lunch etc.

Visa sponsorship: We do sponsor visas! However, we aren't able to successfully sponsor visas for every role and candidate. But if we make you an offer, we will make all reasonable effort to get you a visa, and we retain an immigration lawyer to help with this.

Compensation

The base pay range for this role is $180,000 - $250,000 per year.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
793,143 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

AI/ML
Similar stack
Same company
Sunnyvale
$143k – $193k per year • Equity • In office • Full-Time • 4+ years exp • Master's Degree • Seattle
Python
Java
C++
AI/ML
AI Agents
Machine Learning
DevOps
AWS
Linux
Unix
Apply
$143k – $193k per year • Equity • In office • Full-Time • 4+ years exp • Master's Degree • Bellevue
Python
Java
C++
AI/ML
RLHF
Reinforcement Learning
Function Calling
AI Agents
LLM
RAG
DPO
SFT
PPO
GRPO
Post-training
Multi-Agent Systems
Tool Use
RLAIF
Reward Modeling
DevOps
Linux
Unix
Robotics
Reinforcement Learning
Apply
$143k – $193k per year • Equity • In office • Full-Time • 4+ years exp • Master's Degree • Bellevue
Python
Java
C++
AI/ML
RLHF
Reinforcement Learning
Function Calling
AI Agents
LLM
RAG
DPO
SFT
PPO
GRPO
Post-training
Multi-Agent Systems
Tool Use
RLAIF
Reward Modeling
DevOps
Linux
Unix
Robotics
Reinforcement Learning
Apply
$165k – $224k per year • Equity • In office • Full-Time • 3+ years exp • Bachelor's Degree • Cupertino
Python
Java
C#
C++
Perl
C++
PyTorch C++
AI/ML
vLLM
CUDA Toolkit
SGLang
TensorRT
PyTorch
LLM
CUDA
Triton
AWS Trainium
CUTLASS
Machine Learning
DevOps
CI/CD
AWS
GitHub
Management
Agile
Apply
$180k – $250k per year • Equity • In office • Full-Time • United States
Python
TypeScript
AI/ML
LLM
Apply
AI Engineer 9 hours ago
In office
Python
C++
Swift
C++
TensorFlow C++
PyTorch C++
AI/ML
OpenCV
Triton Inference Server
Quantization
Scikit-learn
Multimodal AI
Computer Vision
TensorRT
TensorFlow
PyTorch
LLM
Point Cloud Library
Synthetic Data
Triton
Amazon SageMaker
ONNX Runtime
Machine Learning
Mobile
Core ML
ARKit
Metal
DevOps
CI/CD
AWS
Wi-Fi
Robotics
Open3D
Sensor Fusion
Apply
$150k – $220k per year • Equity • Remote (United States) • Full-Time • 5+ years exp
Python
AI/ML
vLLM
CUDA Toolkit
Quantization
SGLang
LLM
RunPod
CUDA
Triton
Speculative Decoding
Apply
In office
Verilog
C++
SystemVerilog
AI/ML
Flash Attention
vLLM
CUDA Toolkit
Quantization
TensorRT
LLM
Mixture of Experts
CUDA
Triton
Apply
In office • Bachelor's Degree
Python
Bash
MATLAB
MATLAB
Simulink
AI/ML
vLLM
Triton Inference Server
AI Agents
LocalAI
Ollama
DevOps
Terraform
Ansible
CI/CD
Git
AWS
Docker
Kubernetes
Platform Engineering
Amazon EC2
Amazon S3
IAM
Amazon ECS
HPC
Linux
Apply
≈ $104k – $225k per year (Estimated) • In office • Contractor • PhD • Lausanne
Python
Rust
C++
Assembly
OCaml
C++
PyTorch C++
Assembly
Keystone Engine
AI/ML
vLLM
CUDA Toolkit
AI Agents
SGLang
PyTorch
LLM
CUDA
Triton
Agentic Workflows
DevOps
CI/CD
Apply
$200k – $260k per year • In office • Full-Time • PhD • Sunnyvale
Python
AI/ML
LoRA
vLLM
Triton Inference Server
Fine-tuning
Multimodal AI
NLP
PEFT
Transformers
PyTorch
LLM
Anthropic
Hugging Face
LLM Guardrails
Interpretability
DevOps
GCP
Git
AWS
Docker
Unix
Apply
≈ $51k – $87k per year (Estimated) • In office • Internship • Sunnyvale
Python
AI/ML
Fine-tuning
Multimodal AI
AI Agents
NLP
PyTorch
LLM
Hugging Face
Red Teaming
LLM Guardrails
Interpretability
Multi-Agent Systems
Tool Use
Machine Learning
DevOps
GCP
AWS
Unix
Apply
Software Engineer 1 month ago
≈ $101k – $204k per year (Estimated) • In office • Full-Time • 2+ years exp • Bachelor's Degree • Sunnyvale
Python
Go
Java
C++
Python
FastAPI
Databases
MySQL
PostgreSQL
Redis
RabbitMQ
Apache Kafka
AI/ML
LLM
DevOps
Rest API
gRPC
Terraform
GCP
GitHub Actions
OpenTelemetry
Prometheus
Azure
AWS
Docker
Kubernetes
Grafana
Apply
Marketing Lead 3 months ago
$140k per year • Hybrid • Full-Time • 5+ years exp • Sunnyvale
AI/ML
Claude
Claude Code
DevOps
Vercel
Management
Google Workspace
Apply
$192k – $260k per year • Equity • In office • Full-Time • 6+ years exp • Master's Degree • Sunnyvale
Python
Java
C++
AI/ML
RLHF
Reinforcement Learning
Multimodal AI
Speech Recognition
SFT
Post-training
Pre-training
Text-to-Speech
RLAIF
Reward Modeling
Machine Learning
Apply
Account Executive 1 day ago
In office • Sunnyvale
Apply
≈ $149k – $299k per year (Estimated) • Equity • In office • 2+ years exp • PhD • Sunnyvale
Apply
≈ $110k – $209k per year (Estimated) • In office • Full-Time • Bachelor's Degree • Sunnyvale
AI/ML
Multimodal AI
Computer Vision
Design
SolidWorks
Fusion 360
Apply
$157k – $213k per year • Equity • In office • Full-Time • 3+ years exp • Bachelor's Degree • Sunnyvale
Python
SQL
Perl
MATLAB
SAS
AI/ML
Machine Learning
Apply
See all jobs
This is one of many
793,143 more open roles from verified company boards, updated every day.