368,530open jobs
9,432companies
50,439added this week
Browse all
Salary
$87k – $226k per year (Estimated)
Location
Remote (San Francisco, United States)
Employment
Full-Time
Overview
Company
Impact
Profile match
Andromeda connects AI teams with high-performance compute fast, at scale, and on terms that work. Buy, sell, and operate gpu clusters without complexity.

Forward Deployed Engineer - SRE

Location: North America Remote/SF-Hybrid · Full-Time

About Andromeda

Andromeda gives AI companies access to the kind of scaled compute once reserved for hyperscalers. Our platform connects 100+ AI customers to 50+ global providers, with billions of GPU-hours supported, and those numbers are all rapidly growing. We combine enterprise-grade reliability with the speed and economics of an open market, serving teams running everything from large-scale training to production inference.

Nat Friedman (former CEO of GitHub) and Daniel Gross (former head of AI at Apple, YC partner) started Andromeda in 2023 with a single GPU cluster. It filled almost immediately. Three years later, we're a $1.5B company, profitable since day one, with a Series A from Paradigm to scale the platform globally. The global flow of compute is already a multi-trillion dollar market, and our team is building the infrastructure that enables it to continue to scale.

The problem is deceptively hard. Not all compute is equal: interconnect, networking, OEM, firmware, and cluster age all vary across providers, and the differences matter at scale. Our platform benchmarks and validates capacity, takes positions, structures contracts, and operates clusters globally, delivering a consistent product regardless of where it runs. No one else has built this layer, and the AI industry can't scale without it.

The Role

This is not a generalist SRE role, and it is not a support role. You will embed directly with the teams running large-scale training and inference on our clusters. You are responsible for onboarding them, tuning their jobs, and debugging their failures alongside them, while owning the infrastructure and automation that makes those clusters reliable in the first place.

Forward deployed means you spend real time inside customer environments: reading their training code, sitting in their Slack channels, watching their runs, and shipping fixes that land in our platform. When a multi-hundred-GPU run stalls, you are the person who figures out whether it's the fabric, the driver, the scheduler, or their dataloader, then you make sure it can't happen the same way twice.

We're looking for engineers who have personally run GPU clusters in production, understand the failure modes of distributed training, and can reason about performance from network fabric → kernel → framework. Equally important: you can explain what you found to someone else's engineering team without condescension, and turn that conversation into a product improvement.

What You’ll Do

  • Serve as the primary technical point of contact for teams running large-scale training and inference workloads. Own onboarding end to end; environment setup, orchestration choice (Slurm, Kubernetes, or direct SSH), storage layout, first successful run at scale. You will continue to stay engaged as their workloads grow.

  • Work inside customer environments to diagnose real failures: NCCL timeouts, stragglers, checkpoint I/O stalls, degraded links, OOM patterns, container and driver mismatches. Read their code when you need to. Reproduce, isolate, fix, and write it down.

  • Profile and improve distributed training performance on live workloads. Improving MFU, cutting idle GPU time, and reducing time-to-first-successful-run for new deployments.

  • Own reliability outcomes for the accounts you're deployed on.

  • Ensure the health and performance of high-speed interconnects (InfiniBand, RoCE, NVLink) that underpin distributed training. Diagnose and resolve fabric-level issues that degrade collective operations.

  • Build deep visibility into GPU utilization, memory pressure, interconnect throughput, job performance, and hardware health.

  • Turn every repeated deployment problem into automation: cluster provisioning, GPU health checks and burn-in, preflight validation, self-healing, firmware/driver lifecycle management, and reusable reference configurations for common training and serving stacks.

  • Lead incident response for complex, multi-layer failures spanning hardware, networking, orchestration, and ML frameworks. Own the customer-facing communication during the incident and the blameless postmortem and systemic fix after it.

  • You will see our rough edges before anyone else does. Bring that signal back to influence the roadmap, file the hard bugs, and build the missing pieces yourself when that's the fastest path.

What We’re Looking For

  • Hands-on experience operating GPU clusters in production (NVIDIA A100/H100/H200/B200 or equivalent). You understand GPU memory hierarchies, ECC behavior, thermal throttling, and hardware failure modes from direct experience.

  • Production experience with InfiniBand, RoCE, or NVLink fabrics in the context of distributed training. You can diagnose why an all-reduce is slow, identify a degraded link in a fat-tree topology, and reason about congestion control at scale.

  • Working knowledge of how large training and inference jobs actually run. You don't need to design models, but you need to understand what's happening at the systems level when a large run stalls.

  • Expert-level Linux experience: kernel tuning, driver management (NVIDIA drivers, CUDA toolkit), cgroup/namespace internals, container runtimes, and performance profiling at the syscall and hardware level.

  • Strong experience running Kubernetes in production with GPU workloads. Experience with device plugins, topology-aware scheduling, multi-cluster, custom operators. Experience with Slurm or other HPC schedulers is equally valued.

  • Strong engineering skills in Python, Go, or Bash. You build production-grade tools and services, not just scripts.

  • Infrastructure-as-Code proficiency (Terraform, Helm, Ansible, or equivalent).

  • Hands-on experience building monitoring and alerting for GPU-specific telemetry (DCGM, nvidia-smi, fabric manager metrics) integrated into actionable dashboards.

  • You can go deep on architecture with a customer's infra team and clearly articulate tradeoffs to their leadership. You're comfortable being the only one in the room who knows the answer, and equally comfortable saying you don't yet.

  • Proven track record leading incident response for complex distributed systems.

Strong Candidates May Have

  • Experience with high-performance parallel file systems (VAST, WEKA, Lustre, GPFS) and the checkpoint I/O and data-loading bottlenecks that come with large training runs.

  • Time spent embedded with external engineering teams, i.e solutions architecture, professional services, deployed SRE, or technical account ownership at an infrastructure company.

  • Experience operating production inference. Hands on experience with autoscaling, batching, KV cache behavior, cold starts, and multi-tenant GPU sharing.

  • Contributions to relevant OSS projects, or benchmarks, postmortems, and deep-dives you've published.

  • Experience working across heterogeneous providers and regions rather than a single hyperscaler.

What Success Looks Like

By the end of your first year:

  • You know each of your accounts' actual technical goals including what they're training, what their scaling roadmap looks like over the next two quarters, what their real constraints are (budget, deadline, headcount, data), and you've written that down somewhere the rest of us can read it.

  • You are the person your accounts' engineers message first, before they file a ticket, because you've earned it.

  • Their reliability and throughput numbers are visibly better than at onboarding, and you can point to the specific changes that did it.

  • Recurring problems you found in the field exist as automation, preflight checks, or documentation, not as tribal knowledge in your head.

  • You've advocated internally for at least one roadmap change on behalf of a strategic customer, and it shipped.

Why You’ll Love It Here

  • High-growth environment: Get in early at a company at the center of the AI infrastructure boom

  • Ownership: First FDE for the solutions engineering team, you’ll get to build this function from the ground up

  • Competitive compensation: + meaningful equity

  • Comprehensive benefits: for you and your dependents, including healthcare, dental, and vision coverage, 401(k), and unlimited PTO

Andromeda Cluster is an equal opportunity employer. We celebrate diversity and are committed to creating an inclusive environment for all employees. We do not discriminate on the basis of race, religion, color, national origin, gender, sexual orientation, age, marital status, veteran status, or disability status.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
368,530 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
San Francisco
$28k – $57k per year (Estimated) • Remote • Moscow
DevOps
Ansible
ArgoCD
CI/CD
GitLab CI
GitOps
Grafana
Helm
Jaeger
Kubernetes
Kustomize
Prometheus
Service Mesh
Terraform
GitLab
Cybersecurity
Kyverno
Apply
$91k – $216k per year (Estimated) • In office • Full-Time • 10+ years exp • Amsterdam
Java
JavaScript
Java
Gradle
Frontend
npm
pnpm
DevOps
ArgoCD
AWS
Bazel
CI/CD
Datadog
GitHub Actions
GitOps
Helm
JFrog Artifactory
Kubernetes
SRE
Terraform
GitHub
IAM
Cybersecurity
GDPR
Apply
$27k – $66k per year (Estimated) • Remote/Hybrid • Full-Time • 3+ years exp • Bachelor's Degree • Moscow
Databases
MinIO
OpenSearch
DevOps
Ansible
CI/CD
Docker
Git
Grafana
Kubernetes
Prometheus
Terraform
Amazon S3
GitLab
Management
Confluence
Jira
Apply
$47k – $102k per year (Estimated) • In office • Full-Time • 8+ years exp • Bachelor's Degree • Bengaluru
C++
Java
Python
SQL
C#
TypeScript
JavaScript
Java
Maven
C#
.NET
Databases
Apache Kafka
MySQL
AI/ML
ChatGPT
Copilot
Frontend
Angular
DevOps
CI/CD
Docker
Jenkins
Kubernetes
Prometheus
Apply
Remote • Full-Time • 3+ years exp • Prague
Bash
PowerShell
Python
DevOps
Splunk
Cybersecurity
Crowdstrike
IBM QRadar
Apply
Solutions Architect 1 month ago
$132k – $280k per year (Estimated) • Remote/Hybrid • Full-Time • 2+ years exp • San Francisco
Python
AI/ML
InfiniBand
DevOps
Kubernetes
SLURM
GitHub
HPC
Apply
$115k – $288k per year (Estimated) • Remote/Hybrid • Full-Time • San Francisco
Node JS
JavaScript
Node JS
Commander.js
AI/ML
InfiniBand
DevOps
Incident Management
Kubernetes
SLURM
Apply
$81k – $170k per year (Estimated) • Remote • Full-Time • San Francisco
Bash
Python
AI/ML
CUDA Toolkit
DeepSpeed
PyTorch
CUDA
FSDP
InfiniBand
Megatron-LM
NCCL
NVLink
DevOps
Ansible
Grafana
Helm
Incident Management
Kubernetes
Prometheus
Self-Healing
SLURM
Terraform
HPC
Apply
$104k – $203k per year (Estimated) • Remote • Full-Time • 5+ years exp • San Francisco
Go
Python
DevOps
Ansible
Helm
Kubernetes
Terraform
Apply
$223k – $424k per year (Estimated) • In office • Bachelor's Degree • San Francisco
AI/ML
AI Agents
LLM
Recommender Systems
Apply
$160k – $283k per year • Equity • In office • 5+ years exp • San Francisco
AI/ML
AI Agents
Apply
$83k – $188k per year (Estimated) • In office • 2+ years exp • San Francisco
Python
AI/ML
AI Agents
LLM Guardrails
Model Context Protocol
DevOps
Terraform
Cybersecurity
Crowdstrike
GDPR
Least Privilege
Okta
SentinelOne
Management
Google Workspace
Slack
Apply
$171k – $273k per year • In office • Full-Time • 8+ years exp • PhD • San Francisco • Washington
AI/ML
A2A
Agentforce
AI Agents
Model Context Protocol
DevOps
AWS
GCP
Marketing
Salesforce
Apply
Security GRC Analyst 3 hours ago
$119k – $268k per year (Estimated) • Remote/Hybrid • 4+ years exp • Bachelor's Degree • San Francisco
AI/ML
Ignite
PyTorch
Cybersecurity
ISO 27001
NIST CSF
SOC 2
Apply
See all jobs
This is one of many
368,530 more open roles from verified company boards, updated every day.