699,301open jobs
41,468companies
98,642added this week
Browse all
Salary
$53k – $140k per year (Estimated)
Location
Remote (United Kingdom)
Employment
Full-Time
Overview
Company
Impact
Profile match
Build with generative AI on a unified inference API. Image generation, video generation, audio, 3D, and large language models. 400K+ models, managed infrastructure, usage-based pricing.

About Runware

Runware is building a high-performance, full-stack AI media-creation platform - empowering developers and companies to generate any type of media instantly. As we scale fast and integrate increasingly complex models, we need stronger visibility, analytics, and monitoring across the whole platform stack.

We’re looking for a Data Expert (Analytics + Monitoring + Observability) to help us better understand, measure, and optimize how the Runware platform performs at scale - internally and for our clients.

Mission

Your main goal is to give Runware full visibility over:

  • End-to-end inference performance
  • Integration usage and model activity
  • Errors, delays, bottlenecks, regressions
  • Internal and client-facing analytics dashboards
  • Health and performance of production pipelines

You will provide the data insights that allow engineering, ML, backend, DevOps, and leadership to make informed decisions - and to continuously improve performance and reliability.

What You Will Do

Performance Monitoring & Benchmarking

  • Build and maintain E2E inference time tracking (global and per-model).
  • Monitor how implementation changes impact total request latency.
  • Detect regressions introduced by suboptimal code paths.
  • Provide automated alerts & historical trends.

Usage & Analytics Reporting

  • Build dashboards for internal use (engineering, product, leadership).
  • Provide client-facing usage dashboards (requests, errors, success rate, performance).
  • Support clients who need visibility to debug their integrations.
  • Track model-level usage, API endpoints usage, adoption metrics, etc.

Platform Observability

  • Implement metrics, logs, and traces that help the entire platform scale smoothly.
  • Work closely with DevOps & backend teams to improve system observability.
  • Provide insights that guide infra decisions (GPU allocation, autoscaling, caching, batching, etc.).

Data Infrastructure Ownership

  • Select and maintain tooling (e.g., Prometheus/Grafana, Datadog, OpenTelemetry, ELK, BigQuery, etc.).
  • Ensure data pipelines are reliable, accessible, and always up-to-date.
  • Build simple, easy-to-read dashboards for both technical and non-technical teams.

Requirements

Must-Have

  • Strong experience with data analytics, observability, or monitoring
  • Hands-on with metrics/logging/tracing frameworks (Prometheus, Grafana, Datadog, New Relic, etc.)
  • Good understanding of backend systems and distributed architectures
  • Ability to turn raw metrics into actionable insights
  • Experience building dashboards for internal and external stakeholders
  • Familiarity with AI model monitoring (latency, throughput, error codes, GPU utilization)

Nice-to-Have

  • Experience with AI/ML infrastructure, inference pipelines, GPUs
  • Understanding of Python APIs, FastAPI, or Node environments
  • Experience working with high-throughput real-time systems
  • Startup or scale-up experience

What You Bring

  • A problem-solver mindset
  • Proactivity - you like digging into the data and flagging problems before anyone else sees them
  • Ability to work with ML, backend, DevOps, and product teams
  • Comfort with autonomous ownership

You help Runware go from “it works” to “we know exactly how well it works - and how to make it better.”

Benefits

We’re a remote-first collective, meeting in person twice a year to plan, brainstorm, celebrate wins, and enjoy some face-to-face time. We have core hours for cooperative working and calls, but outside of that your calendar is yours. Work the hours that let you perform at your peak while also building a healthy life.

Our release cycles are fast and intense, but they’re followed by real downtime. After big pushes we expect the team to unplug, recharge, and come back ready & stronger than ever for the next leap.

  • Generous paid time off - vacation, sick days, public holidays
  • Meaningful stock options - share in the upside you create
  • Remote-first setup - work from home anywhere we can employ you
  • Flexible hours - own your schedule outside core collaboration blocks
  • Family leave - paid maternity, paternity, and caregiver time
  • Company retreats - twice-yearly gatherings in inspiring locations
Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
699,301 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
In your city
In office • 2+ years exp • Bachelor's Degree • Bengaluru
Python
JavaScript
Node JS
Node JS
Commander.js
Databases
PostgreSQL
AI/ML
AI Agents
LLM
DevOps
Terraform
GitHub Actions
OpenTelemetry
PagerDuty
Prometheus
CI/CD
Git
AWS
Kubernetes
Grafana
Platform Engineering
Self-Healing
Amazon EKS
AWS Fargate
AWS Lambda
Amazon EC2
Incident Management
SLI/SLO/SLA
Amazon S3
IAM
Amazon ECS
Amazon CloudWatch
DNS
Cybersecurity
ISO 27001
SOC 2
Least Privilege
Apply
$33k – $73k per year (Estimated) • In office • Full-Time • Bengaluru
Python
DevOps
Terraform
OpenShift
Helm
OpenTelemetry
Prometheus
AWS
Docker
Kubernetes
Grafana
Thanos
Amazon EKS
Incident Management
TCP/IP
DNS
Cybersecurity
ISO 27001
PCI DSS
SOC 2
HIPAA
Apply
$21k – $45k per year (Estimated) • Remote/Hybrid • Full-Time • 5+ years exp • Penza
Bash
Databases
PostgreSQL
DevOps
Ansible
Helm
Prometheus
GitLab CI
CI/CD
Git
Docker
Kubernetes
Ubuntu
Nginx
Grafana
Linux
Apply
$14k – $27k per year (Estimated) • Remote/Hybrid • Full-Time • Moscow
DevOps
Zabbix
Prometheus
Grafana
Proxmox VE
TCP/IP
Cybersecurity
Wireshark
Tcpdump
ViPNet
Apply
Head of Engineering 2 days ago
$167k – $241k per year • In office • Full-Time • Berlin
Python
JavaScript
TypeScript
Node JS
Node JS
Fastify
Databases
PostgreSQL
AI/ML
OpenAI
Semantic Search
DevOps
GCP
Datadog
Azure
AWS
Apply
$10k – $29k per year (Estimated) • In office • Full-Time • Bucharest
Apply
$103k – $207k per year (Estimated) • Equity • Remote • Full-Time • Bachelor's Degree
Python
C++
C++
PyTorch C++
AI/ML
LoRA
vLLM
CUDA Toolkit
Fine-tuning
Multimodal AI
Diffusion Models
PEFT
Transformers
PyTorch
CUDA
Triton
Edge AI
Machine Learning
DevOps
Docker
Kubernetes
GitHub
Apply
$99k – $199k per year (Estimated) • Equity • Remote • Full-Time
PHP
PHP
Symfony
Doctrine
DevOps
Rest API
WebSockets
Management
Stripe
Apply
$137k – $226k per year (Estimated) • Equity • Remote • Full-Time
Go
PHP
Apply
Engineering Manager 1 month ago
$96k – $195k per year (Estimated) • Equity • Remote • Full-Time • Master's Degree • London
Python
Go
PHP
Rust
DevOps
CI/CD
Apply
See all jobs
This is one of many
699,301 more open roles from verified company boards, updated every day.