699,301open jobs
41,468companies
98,642added this week
Browse all
Salary
$90k – $168k per year (Estimated)
Location
Remote (Germany, France, United Kingdom, Italy, Spain, Sweden)
Seniority
Senior
Employment
Full-Time
Overview
Company
Impact
Profile match
Build with generative AI on a unified inference API. Image generation, video generation, audio, 3D, and large language models. 400K+ models, managed infrastructure, usage-based pricing.

Runware is building high-performance infrastructure and products to power the worlds intelligence. Our platform enables developers and businesses to run fast, scalable inference across image, video and emerging modalities, while our Serverless platform allows customers to deploy and scale their own AI models on production-grade GPU infrastructure.

As a Site Reliability Engineer at Runware, you will help ensure these systems remain reliable, performant and resilient as we scale. This is a highly technical, hands-on role working across software, infrastructure and production operations to improve observability, reduce incidents, eliminate operational toil and build lasting improvements across complex distributed systems.

What you’ll do

  • Own and improve the reliability, availability and performance of critical production services across the Runware platform
  • Define and evolve our reliability practices, including SLIs, SLOs, alerting, observability and production-readiness standards
  • Investigate complex production issues across distributed systems, APIs, networking, queues, databases and GPU-backed workloads, participating in our engineering on-call rotation
  • Lead and contribute to incident reviews and RCAs, turning recurring failure modes into lasting engineering improvements
  • Reduce operational toil through automation, automated remediation and improvements to deployment safety, recovery and system resilience
  • Work closely with Engineering and DevOps teams on capacity planning, performance, scaling and architectural improvements as the platform grows

Requirements

  • Have strong experience operating and troubleshooting production systems at scale in an SRE, Production Engineering, Platform Engineering or similar role
  • Have a strong understanding of distributed systems and are comfortable debugging across applications, databases, queues, containers, networking and infrastructure
  • Have experience designing and operating observability systems using metrics, logs and distributed tracing
  • Understand SRE principles including SLIs, SLOs, error budgets, capacity planning, incident management and reducing operational toil
  • Have experience with Kubernetes, containers, IaC and automated deployment practices, alongside the ability to write software and automation using languages such as Python, Go or PHP
  • Take strong ownership of production problems and are comfortable participating in an engineering on-call rotation, taking issues from initial investigation through to long-term remediation

Bonus

  • Experience operating high-throughput or low-latency APIs and distributed systems
  • Experience with bare-metal infrastructure, GPU environments or AI and ML workloads
  • Experience with RabbitMQ or other distributed messaging and queueing systems
  • Experience operating MySQL, Redis, ClickHouse or similar production data systems
  • Experience with global traffic management, load balancing, CDN platforms and hybrid infrastructure environments
  • Experience building automated scaling, capacity management or self-healing systems

Benefits

We’re a remote-first collective, meeting in person twice a year to plan, brainstorm, celebrate wins, and enjoy some face-to-face time. We have core hours for cooperative working and calls, but outside of that your calendar is yours. Work the hours that let you perform at your peak while also building a healthy life.

Our release cycles are fast and intense, but they’re followed by real downtime. After big pushes we expect the team to unplug, recharge, and come back ready & stronger than ever for the next leap.

  • Generous paid time off - vacation, sick days, public holidays
  • Meaningful stock options - share in the upside you create
  • Remote-first setup - work from home anywhere we can employ you
  • Flexible hours - own your schedule outside core collaboration blocks
  • Family leave - paid maternity, paternity, and caregiver time
  • Company retreats - twice-yearly gatherings in inspiring locations
Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
699,301 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
In your city
SRE-инженер 2 days ago
$15k – $29k per year (Estimated) • In office • 4+ years exp • Saint Petersburg
Python
JavaScript
SQL
Node JS
Python
Django
Celery
Gunicorn
Uvicorn
Node JS
Axios
Databases
PostgreSQL
Redis
ClickHouse
MinIO
Apache Kafka
Frontend
Webpack
React.js
Material UI
React Router
DevOps
Terraform
Ansible
Docker Compose
Helm
Azure DevOps
Prometheus
Azure
CI/CD
GitOps
Jenkins
Git
Docker
Kubernetes
Nginx
Grafana
Gitflow
Trunk-Based Development
Error Budget
SLI/SLO/SLA
Astra Linux
Cybersecurity
Keycloak
HashiCorp Vault
Cryptography
Vault
QA
Selenium
Jest
Pytest
Apply
$20k – $44k per year (Estimated) • Remote/Hybrid • Full-Time • Saint Petersburg
PHP
1C
PHP
Bitrix
Databases
MySQL
DevOps
Rest API
Apply
$106k – $143k per year • In office
Python
SQL
Databases
Google BigQuery
BigQuery
AI/ML
Copilot
Claude
Model Context Protocol
Vertex AI
AI Agents
AWS Bedrock
LLM
DevOps
Terraform
GCP
Azure
CI/CD
AWS
Docker
Kubernetes
FinOps
Cybersecurity
SOC 2
HIPAA
Management
Service Desk
Apply
SDE 2 / 3 (Backend) 2 days ago
$22k – $55k per year (Estimated) • In office • 4+ years exp • Mumbai
Go
Java
C++
Databases
Redis
Aerospike
RabbitMQ
Apache Kafka
Apply
$19k per year (gross) • In office • Full-Time • Saint Petersburg
PHP
PHP
Bitrix
Design
Adobe Photoshop
Management
Telegram
Marketing
Mailchimp
Unisender
AmoCRM
Apply
$10k – $29k per year (Estimated) • In office • Full-Time • Bucharest
Apply
$103k – $207k per year (Estimated) • Equity • Remote • Full-Time • Bachelor's Degree
Python
C++
C++
PyTorch C++
AI/ML
LoRA
vLLM
CUDA Toolkit
Fine-tuning
Multimodal AI
Diffusion Models
PEFT
Transformers
PyTorch
CUDA
Triton
Edge AI
Machine Learning
DevOps
Docker
Kubernetes
GitHub
Apply
$99k – $199k per year (Estimated) • Equity • Remote • Full-Time
PHP
PHP
Symfony
Doctrine
DevOps
Rest API
WebSockets
Management
Stripe
Apply
$137k – $226k per year (Estimated) • Equity • Remote • Full-Time
Go
PHP
Apply
Engineering Manager 1 month ago
$96k – $195k per year (Estimated) • Equity • Remote • Full-Time • Master's Degree • London
Python
Go
PHP
Rust
DevOps
CI/CD
Apply
See all jobs
This is one of many
699,301 more open roles from verified company boards, updated every day.