Overview
Company
Profile match
Impact
Conditions
Benefits
Hiring process
Similar jobs

Rockstar

Human-led recruiting. Reinvented with AI.For 10x Less.. Personalized, multi-step sequences sent via Rockstar or "on behalf of" your team.

Rockstar is recruiting for a fast-growing startup that is building the AI backbone for the next generation of intelligent products. They help fast-growing AI startups design, fine-tune, evaluate, deploy, and maintain specialized models across text, vision, and embeddings. Think of them as “AWS for AI models”-not data or raw compute, but a full-stack backend for fine-tuning, reinforcement learning, inference, and long-term model maintenance. Their customers are Series A-C AI companies building enterprise-grade products. Their promise is simple: they make your AI system better.

They are hiring a Backend Software Engineer (ML Infrastructure) to help design, build, and scale the core systems that power large-scale model training and deployment.

The candidate will work on distributed training pipelines, cloud-native infrastructure, and internal developer platforms that support fine-tuning, reinforcement learning, and inference at scale. This role sits at the intersection of backend engineering and ML systems-the candidate will collaborate closely with ML engineers while owning production-grade infrastructure.

This is an ideal role for an early-career engineer who wants to work on real distributed systems, GPU workloads, and modern ML infrastructure-not dashboards or CRUD apps.

What You’ll Do

Build & Scale Core Infrastructure

- Design and implement backend systems that support large-scale ML workloads, including fine-tuning and reinforcement learning.

- Build distributed training and inference pipelines that are efficient, fault-tolerant, and observable.

- Develop internal developer tools and platforms that make it easier for ML engineers to train, evaluate, and deploy models.

Cloud & Systems Engineering

- Work on cloud-native systems using containers and orchestration (e.g., Kubernetes).

- Optimize systems for performance, reliability, and cost efficiency, especially for GPU-heavy workloads.

- Implement monitoring, logging, and observability for long-running training jobs and production services.

Collaborate with ML Engineers

- Partner closely with ML engineers to support evolving model architectures, training workflows, and evaluation needs.

- Translate ML requirements into scalable backend and infrastructure solutions.

Who You Are

Required

- 1-3 years of backend engineering experience, ideally working on production systems.

- Strong fundamentals in distributed systems, networking, and backend architecture.

- Experience building systems that scale under real load.

- Comfortable working in Python and/or Go (or similar backend languages).

- Excited to work on-site in San Francisco with a fast-moving early-stage team.

Strongly Preferred

- Experience with or exposure to ML infrastructure or ML platforms.

- Familiarity with GPU workloads, training pipelines, or inference systems.

- Experience with containerization and orchestration (Docker, Kubernetes).

- Contributions to or deep familiarity with ML infrastructure libraries such as:

- Ray

- vLLM

- SGLang

- or similar distributed ML systems

Bonus

- Computer science background from a top-tier program or equivalent demonstrated excellence.

- Open-source contributions, research projects, or side projects in systems or ML infrastructure.

- A track record of high ownership and technical curiosity.

Recommended for you based on this role

Similar stack
Same company
In your city
Product Manager 2 days ago
In office • San Francisco
AI/ML
AI Agents
Claude
Claude Code
Cursor
Apply
Sr. Product Manager 3 days ago
$145k – $190k per year
DevOps
Datadog
Kubernetes
Apply
JavaScript
Python
SQL
Databases
MySQL
PostGIS
PostgreSQL
AI/ML
ChatGPT
Claude
DevOps
SLI/SLO/SLA
Management
Jira
Marketing
HubSpot
Zendesk
SpaceTech
QGIS
Apply
Remote • San Francisco
Apply
Remote • Full-Time • 8+ year exp • Bachelor's Degree • United States
Go
Python
Databases
Dgraph
DevOps
Platform Engineering
Apply
Software Engineer 28 days ago
$150k – $225k per year • Seattle
TypeScript
JavaScript
Frontend
GraphQL
Next.js
React.js
Apply
$175k – $200k per year • Remote
Python
Python
Django
Databases
PostgreSQL
DevOps
AWS
Incident Management
Apply
$150k – $225k per year • Full-Time • Los Angeles
TypeScript
Databases
Databricks
PostgreSQL
AI/ML
AI Agents
DevOps
Docker
Apply
Senior AI Engineer 3 months ago
Full-Time
Python
AI/ML
AutoGen
BentoML
CrewAI
Embeddings
Fine-tuning
LangGraph
LLM
LoRA
Multimodal AI
NLP
Prompt Engineering
PyTorch
QLoRA
Quantization
RAG
Ray
Ray Serve
Reranking
Scikit-learn
Semantic Search
TensorFlow
Transformers
vLLM
LangChain
PEFT
DevOps
CI/CD
Docker
Kubernetes
Vector
Apply
VP of Engineering 4 months ago
$250k – $325k per year • Full-Time • New York
JavaScript
Python
Apply
Career impact
Discover how this job can transform your career
Get a personal career forecast for this job - salary uplift, next-level role, skill boost and a 3-year financial impact, all calculated from your profile.
Personal salary uplift vs. your current pay
Your 3-year career trajectory
Skills you will level up in this role
3-year financial impact in dollars
Create free account
Free forever • Less than a minute • No credit card

Work setup

Location
San Francisco
Remote work
In office