1,008,666open jobs
59,992companies
167,270added this week
Browse all
Salary
≈ $175k – $348k per year (Estimated)
Location
Hybrid (San Francisco, United States)
Seniority
Principal · 10+ years exp
Employment
Full-Time

Confirmed on the employer's own hiring board on Oct 1, 2026. First seen by Alion on Apr 24, 2026. SPREEAI scores B on the Alion truth index.

Overview
Company
Impact
Profile match
SPREEAI is the AI fashion-commerce platform connecting photorealistic virtual try-on, fit guidance and outfit intelligence.

SPREEAI is a fast-growing, innovative AI company at the forefront of fashion and e-commerce,revolutionizing how consumers engage with fashion through lifelike photorealistic try-on technology and hyper-personalized shopping experiences.Our mission is to redefine the retail landscape with cutting-edge AI solutions that blend high fashion and technology. We thrive in a dynamic, fast-paced environment where creativity meets technology to drive real impact. If you are passionate about innovation and shaping the future of fashion, SPREEAI offers a platform to make your mark.

About the Role

SPREEAI is building the future of AI-powered commerce through photorealistic virtual try-on and multimodal intelligence. We bring together cutting-edge AI and real-world retail to deliver production systems that redefine how people shop online.

We are looking for a Principal Engineer to build the infrastructure, deployment pipelines, and observability systems that enable multimodal AI models to move from research prototypes to reliable, production-grade deployments powering real-time virtual try-on experiences for global retail partners.

This role spans ML platform engineering, deployment systems, GPU infrastructure, and observability. You will partner closely with Applied Science, AI Platform, Product, and Partner Engineering to enable rapid research iteration and reliable model delivery at scale.

What You'll Own

ML Platform & Training Enablement

  • Build and operate SPREEAI’s end-to-end ML platform spanning training, evaluation, deployment, and monitoring.
  • Enable scalable and reliable training workflows through orchestration, infrastructure, and resource management systems.
  • Define platform standards for model packaging, model registry, dataset lineage, experiment tracking, checkpointing, and deployment automation.

Deployment, Inference & Observability

  • Enable reliable and scalable inference deployments through standardized serving, orchestration, and monitoring frameworks.
  • Build and operate model deployment pipelines with versioning, reproducibility, rollback, approval gates, evaluation gates, and production observability.
  • Establish production SLOs for latency, availability, error rate, GPU saturation, cold-start time, cost per inference, and model quality drift.
  • Standardize and support serving infrastructure using modern inference runtimes such as vLLM, NVIDIA Triton, TensorRT-LLM, Ray Serve, TorchServe, ONNX Runtime, or equivalent systems.

GPU Infrastructure & System Efficiency

  • Design and manage GPU allocation, scheduling, and resource utilization across training and inference workloads.
  • Improve GPU utilization, throughput, latency, reliability, and cost efficiency across model lifecycle systems.
  • Design and operate model evaluation and benchmarking systems, including automated regression detection and quality gates for production releases.
  • Partner with research teams to productionize new capabilities by providing robust infrastructure, tooling, and deployment pathways.

What We're Looking For

  • 10+ years of software engineering / infrastructure experience, with 5+ years in ML infrastructure, MLOps, distributed systems, or AI platform engineering.
  • Deep experience with Python, PyTorch, Kubernetes, Docker, cloud infrastructure, and GPU-based workloads.
  • Strong understanding of distributed systems and large-scale ML infrastructure design.
  • Experience with ML workflow orchestration systems such as Ray, Kubeflow, Argo, Airflow, Flyte, or Metaflow.
  • Experience deploying and managing production inference systems using platforms like Triton, vLLM, TensorRT-LLM, Ray Serve, KServe, Seldon, BentoML, TorchServe, or custom services.
  • Strong understanding of inference optimization techniques such as batching, quantization, CUDA graphs, and memory-aware scheduling.
  • Experience with model registries, experiment tracking, CI/CD for ML, canary deployments, shadow traffic, rollback strategies, and production monitoring.
  • Strong cloud experience across AWS, GCP, Azure, or GPU-focused providers like CoreWeave, Lambda Labs, or RunPod.
  • Ability to debug performance bottlenecks across distributed systems, containers, networking, GPU memory, and storage layers.

Strong ownership mindset with the ability to define architecture, set platform standards, and drive execution across teams.

Nice to Have

  • Experience with multimodal, vision, or generative AI systems.
  • Experience with large-scale GPU clusters e.g. A100/H100, NCCL, and high-throughput data pipelines.
  • Experience designing evaluation and monitoring systems for generative AI workloads.
  • Familiarity with ML security, privacy, and data governance practices.
  • Experience building internal developer platforms for research teams.

Success Looks Like

  • Within 6 months, you will:
  • Create reliable research-to-production pathways for SPREEAI’s core AI models.
  • Reduce manual model deployment friction through standardized pipelines and tooling.
  • Improve GPU utilization and reduce training and inference costs.
  • Establish robust observability and evaluation gates for production model releases.
  • Accelerate the delivery of new AI capabilities into partner-facing experiences.

Why This Role Matters

This is not a traditional DevOps role. This is the infrastructure backbone that enables SPREEAI to turn frontier AI research into reliable, scalable, production-grade systems. You will define the systems powering real-time AI experiences where latency, cost, and model quality directly impact end-user experience.

Why Join SPREEAI?

  • Build the Core AI Infrastructure, Not Just Features: You will define how multimodal AI systems are reliably deployed, monitored, and scaled-directly shaping the performance, cost efficiency, and reliability of real-world AI products.
  • Own Systems End-to-End: You will own critical infrastructure decisions across deployment, observability, and resource management, with direct impact on production systems serving real partner traffic.
  • Work on Hard, High-Leverage Problems: From GPU efficiency to large-scale deployment systems, you will tackle challenges that sit at the frontier of real-time AI infrastructure.
  • High Velocity, Low Bureaucracy & Direct Impact: We operate with tight feedback loops between research, platform, and product, enabling rapid iteration and meaningful impact without organizational friction.
Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
1,008,666 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Backend
Similar stack
Same company
San Francisco
$95k – $148k per year • Remote (United States) • Full-Time • 8+ years exp • Bachelor's Degree • United States
C#
Databases
MS SQL
Frontend
Redux
Mobile
State Management
Apply
≈ $151k – $273k per year (Estimated) • Hybrid • 15+ years exp • Master's Degree • Richmond
Management
Agile
Apply
$206k – $211k per year • Equity • In office • Full-Time • 7+ years exp • Bachelor's Degree • Bellevue
Java
Kotlin
Java
Gradle
Kotlin
Kotlin Coroutines
Detekt
Ktlint
Mobile
Jetpack Compose
MVVM
Clean Architecture
Espresso
Fastlane
ProGuard
DevOps
CI/CD
Jenkins
Docker
GitLab
Cybersecurity
SonarQube
Apply
Staff Engineer 1 hour ago
$269k – $307k per year • In office • Full-Time • 10+ years exp • Bachelor's Degree • McLean • Plano • Richmond
Python
Go
JavaScript
Rust
TypeScript
C#
Scala
AI/ML
AI Agents
Machine Learning
DevOps
GCP
Azure
AWS
HPC
Apply
$286k – $327k per year • In office • Full-Time • 12+ years exp • Bachelor's Degree • McLean • Richmond • Plano
Python
Go
JavaScript
Rust
TypeScript
C#
Scala
AI/ML
AI Agents
Machine Learning
DevOps
GCP
Azure
AWS
HPC
Apply
≈ $47k – $130k per year (Estimated) • In office • Full-Time • Brazil
Python
Java
SQL
PowerShell
Databases
Oracle
Apply
≈ $48k – $131k per year (Estimated) • In office • Full-Time • Brazil
Python
JavaScript
Java
Ruby
Java
Spring Boot
Frontend
React.js
Mobile
React Native
DevOps
CI/CD
GitLab
QA
Postman
SoapUI
Apply
≈ $50k – $137k per year (Estimated) • In office • Full-Time • Brazil
Python
Java
AI/ML
Machine Learning
Apply
≈ $30k – $75k per year (Estimated) • In office • Full-Time • Brazil
Python
SQL
Databases
MS SQL
DevOps
Azure DevOps
Azure
CI/CD
Jenkins
GitLab
Analytics
SSIS
Azure Data Factory
Apply
≈ $25k – $61k per year (Estimated) • In office • Full-Time • Brazil
Python
SQL
Databases
MS SQL
DevOps
Azure DevOps
Azure
CI/CD
Jenkins
GitLab
Analytics
SSIS
Azure Data Factory
Apply
Full Stack Engineer 22 days ago
≈ $125k – $250k per year (Estimated) • Hybrid • Full-Time • 3+ years exp • San Francisco
Go
JavaScript
Node JS
Databases
PostgreSQL
DynamoDB
AI/ML
Edge AI
Frontend
Zustand
Redux
React.js
React Query
Redux Toolkit
Mobile
State Management
DevOps
Rest API
Design
Figma
Apply
≈ $127k – $253k per year (Estimated) • Hybrid • Full-Time • 3+ years exp • San Francisco
Go
JavaScript
Node JS
Databases
PostgreSQL
DynamoDB
AI/ML
Edge AI
DevOps
Rest API
Docker
Kubernetes
Apply
≈ $151k – $314k per year (Estimated) • Hybrid • Full-Time • Bachelor's Degree • San Francisco
Python
Go
Java
C++
C++
PyTorch C++
AI/ML
Ray Serve
vLLM
Multimodal AI
PyTorch
Ray
Triton
Edge AI
Machine Learning
DevOps
Docker
Kubernetes
Apply
MLOps Engineer 22 days ago
≈ $157k – $343k per year (Estimated) • Hybrid • Full-Time • San Francisco
Python
Databases
Delta Lake
AI/ML
Weights & Biases
MLFlow
Kubeflow
Ray
Edge AI
DevOps
Helm
CI/CD
Docker
Kubernetes
Analytics
A/B Testing
Apply
Frontend Engineer 2 months ago
≈ $133k – $273k per year (Estimated) • Hybrid • Full-Time • 3+ years exp • San Francisco
Go
JavaScript
AI/ML
Edge AI
Frontend
Zustand
Redux
React.js
React Query
Redux Toolkit
Mobile
State Management
Apply
$195k – $250k per year • Equity • Hybrid • Full-Time • Bachelor's Degree • San Francisco
Python
Go
JavaScript
Java
Ruby
C
C++
Node JS
C
FFmpeg
AI/ML
Machine Learning
DevOps
Azure
CI/CD
AWS
Docker
Kubernetes
Management
Agile
Scrum
Apply
$185k – $200k per year • In office • 7+ years exp • San Francisco
DevOps
CI/CD
Apply
$50k per year • In office • 5+ years exp • San Francisco
Python
Go
JavaScript
Java
Kotlin
TypeScript
Node JS
Node JS
Nest.JS
Databases
Redis
Frontend
Next.js
React.js
DevOps
GitHub Actions
CI/CD
Docker
Kubernetes
Apply
≈ $142k – $296k per year (Estimated) • In office • San Francisco
Apply
Founding Engineer 11 hours ago
≈ $142k – $296k per year (Estimated) • In office • San Francisco
Apply
See all jobs
This is one of many
1,008,666 more open roles from verified company boards, updated every day.