372,576open jobs
9,648companies
49,982added this week
Browse all
Salary
$120k – $160k per year
Location
Remote (United States)
Seniority
Middle · 3+ years exp
Employment
Full-Time
Overview
Company
Impact
Profile match
RunPod is a cloud computing company headquartered in Mount Laurel, New Jersey, and founded in 2022. The company provides a specialized AI developer cloud featuring on-demand GPU instances, serverless GPU endpoints, and multi-node GPU clusters for training and inference. It operates a global infrastructure across more than 30 regions, serving over one million developers with scalable compute resources for machine learning and generative AI workloads.

Runpod is the AI Developer Cloud. More than one million developers, from indie researchers to teams running frontier models in production, use Runpod to experiment, train, fine-tune, deploy, and scale AI on one platform. The platform has processed more than 20 billion inference requests. We closed a $100M Series A in June 2026. We're at an inflection point for AI infrastructure, and we're building the platform the next generation of developers will depend on.

We're a small, remote-first team. We take ownership seriously, move fast, and ship work that more than a million developers rely on every day. We're looking for people who care deeply, build with urgency, and want to matter at scale.

Learn more in our CEO's funding announcement: https://www.runpod.io/blog/one-million-developers.

We are looking for a Datacenter Infrastructure Specialist to be the operational linchpin of our global fleet. Reporting to the Manager of Infrastructure Capacity & Management, you will serve as the technical authority bridging our hardware partners and internal engineering teams.

As our Datacenter Infrastructure Specialist, you will own the technical lifecycle and operational health of Runpod’s rapidly expanding, high-density GPU fleet. You will be part of a team that acts as the technical anchor for our hardware partners-serving as their infrastructure advisor, technical translator, adopter, and incident commander. This role blends deep HPC systems engineering, advanced network troubleshooting, and process automation to ensure rock-solid uptime for the world's most demanding AI workloads.

This is a high-visibility, high-impact position where you will move beyond traditional ticket-closing. You will work directly with cutting-edge GPU cloud technologies, advanced RDMA fabrics, and modern observability stacks to solve complex hardware challenges at scale. If you want the autonomy to build automated infrastructure tooling and directly contribute to the resilience, scalability, and revenue velocity of Runpod’s global physical backbone, this is where you do it.

Responsibilities

  • Hardware Validation & Benchmarking: Assist in validating new hardware, ensuring partner deployments meet Runpod’s specifications for distributed AI/ML workloads.

  • Uptime & SLA Enforcement: Monitor fleet health to identify performance degradation. You will help audit downtime and provide the technical data needed to protect customer SLAs.

  • AI-Driven Operations: We operate with an AI-first mindset, powering our operations with the technology we host. You will work with LLMs and AI agents to help automate network triage and generate dynamic runbooks for our fleet.

  • Incident Support: Coordinate technical incident communications with clear updates, acting as a steady hand that translates outages into actionable resolutions.

  • Partner Technical Support: Support the growth of our infrastructure partners

Requirements

  • Professional Background: 3-5 years of experience in infrastructure operations, systems reliability, or datacenter engineering.

  • Datacenter Networking: Strong proficiency in standard datacenter networking and performance troubleshooting. Exposure to RDMA, InfiniBand, or RoCE is highly preferred.

  • GPU & AI Stack: Hands-on experience with the NVIDIA Software Stack (driver installation, performance utilities) and an understanding of multi-node performance tuning.

  • Systems & Diagnostics: Solid Linux system administration skills and experience with containerization (Docker). You are comfortable performing system-level troubleshooting and performance tuning at the kernel and hardware interface layers.

  • Effective Communication: Clear written and verbal communication skills. You can explain hardware or networking issues to both technical partners and internal leadership.

  • Operational Flexibility: As our global fleet scales, this role may require participating in an on-call rotation in the future.

  • Strategic Problem-Solver: You are detail-oriented and proactive when it comes to identifying potential failures before they impact customers.

Preferred

  • Startup Experience: Experience working in a fast-paced environment where you have contributed to building operational workflows.

  • HPC Exposure: Experience managing or optimizing bare-metal High-Performance Computing environments at massive scale.

  • Observability Tools: Experience with Grafana, Prometheus, or Datadog to monitor system health.

  • Automation: Proficiency in Python, Go (Golang), or Bash to automate repetitive infrastructure tasks and interface with internal APIs.

What You’ll Receive:

  • The competitive base pay for this position ranges from $120,000.00 - $160,000.00. This salary range may be inclusive of several career levels at Runpod and will be narrowed during the interview process based on a number of factors, including the candidate’s experience, qualifications, and location.

  • Meaningful equity in a fast-growing AI infra company - everyone on the team receives stock options - your impact drives our growth, and you share in the upside.

  • Generous medical, dental & vision plans - we cover 100% for all employees and partial for dependents.

  • Flexible PTO - take the time you need to recharge.

  • Most roles are remote work first with inclusive, collaborative teams utilizing Slack as the main form of internal communication.

  • Join a passionate team on the cutting edge of AI infrastructure - where culture, learning, and ownership are at the heart of how we scale.

  • $1,200 Home Office & Equipment Stipend - We set you up for success from day one with gear and support to create your ideal workspace.

Runpod is committed to maintaining a workplace free from discrimination and upholding the principles of equality and respect for all individuals. We believe that diversity in all its forms enhances our team. As an equal opportunity employer, Runpod is committed to creating an inclusive workforce at every level. We evaluate qualified applicants without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, age, marital status, protected veteran status, disability status, or any other characteristic protected by law. We welcome every qualified candidate eligible to work in the United States; however, we are currently unable to sponsor employment visas.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
372,576 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
In your city
$131k – $237k per year • In office • Full-Time • 4+ years exp • Bachelor's Degree • Chantilly • Columbia
Bash
Python
TypeScript
JavaScript
Databases
PostgreSQL
Frontend
Angular
DevOps
Ansible
ArgoCD
AWS
Azure
buildah
CI/CD
Docker
Docker Swarm
FluxCD
GitLab
GitLab CI
Grafana
Helm
Kubernetes
Loki
Prometheus
Terraform
Twelve-Factor App
Podman
Apply
$108k – $195k per year • In office • Full-Time • 8+ years exp • Orlando
Bash
Python
DevOps
Ansible
AWS
Azure
CI/CD
Configuration Management
Docker
Git
GitLab
Kubernetes
OpenShift
Red Hat
Apply
$108k – $195k per year • In office • Full-Time • 8+ years exp • Bachelor's Degree • United States
Python
AI/ML
AI Agents
Amazon SageMaker
AWS Bedrock
AWS Bedrock AgentCore
LLM
DevOps
AWS
CI/CD
CloudFormation
Docker
IAM
Kubernetes
Platform Engineering
Terraform
Cybersecurity
FedRAMP
Apply
$215k – $240k per year • In office • Full-Time • 5+ years exp • Seattle
Python
Rust
AI/ML
Flyte
OpenAI
DevOps
ArgoCD
AWS
Azure
Buildkite
CI/CD
Docker
GCP
Helm
Kubernetes
Terraform
Apply
$70k – $126k per year • In office • Full-Time • 3+ years exp • United States
Python
DevOps
Azure
Azure DevOps
Bitbucket
CI/CD
Docker
GitHub
GitLab
GitLab CI
Jenkins
Kubernetes
Cybersecurity
SonarQube
QA
Postman
Robot Framework
Selenium
Apply
$225k – $325k per year • Equity • Remote • Full-Time • 7+ years exp
AI/ML
LLM
RunPod
DevOps
Kubernetes
HPC
Apply
$140k – $165k per year • Equity • Remote • Full-Time • 4+ years exp • Bachelor's Degree
AI/ML
RunPod
Apply
$100k – $180k per year • Equity • Remote • Full-Time • 3+ years exp • Bachelor's Degree
Go
JavaScript
Node JS
Python
SQL
Python
Django
Flask
Databases
MySQL
PostgreSQL
AI/ML
Fine-tuning
LLM
PyTorch
RunPod
AI Agents
Frontend
React.js
DevOps
Docker
Ubuntu
Apply
Senior Data Engineer 19 days ago
$175k – $220k per year • Equity • Remote • Full-Time • 5+ years exp • Bachelor's Degree • San Francisco
Databases
Amazon Redshift
Databricks
Snowflake
AI/ML
Dagster
dbt
RunPod
Cybersecurity
SOC 2
Analytics
ETL/ELT
Apply
$130k – $200k per year • Equity • Remote • Full-Time
Go
JavaScript
Python
TypeScript
Python
FastAPI
Pydantic
AI/ML
RunPod
Frontend
React.js
DevOps
CI/CD
Docker
Git
Apply
See all jobs
This is one of many
372,576 more open roles from verified company boards, updated every day.