376,406open jobs
9,791companies
48,125added this week
Browse all
Salary
$87k – $181k per year (Estimated)
Location
In office (Reading)
Seniority
Senior · 5+ years exp
Employment
Full-Time
Overview
Company
Impact
Profile match
CloudFactory is a company founded in 2010 that provides managed teams for data annotation, document processing and model evaluation. It built its delivery model around training workers in Nepal, Kenya and other emerging markets, pairing them with tooling and quality management for machine learning customers. The company serves autonomous vehicle, medical imaging and geospatial programmes that need consistent human-in-the-loop work at scale.

At CloudFactory, we are a mission-driven team passionate about unlocking the potential of AI to transform the world. By combining advanced technology with a global network of talented people, we make unusable data usable, driving real-world impact at scale.

More than just a workplace, we’re a global community founded on strong relationships and the belief that meaningful work transforms lives. Our commitment to earning, learning, and serving fuels everything we do as we strive to connect one million people to meaningful work and build leaders worth following.

Our Culture

At CloudFactory, we believe in building a workplace where everyone feels empowered, valued, and inspired to bring their authentic selves to work. We are:

  • Mission-Driven: We focus on creating economic and social impact.
  • People-Centric: We care deeply about our team’s growth, well-being, and sense of belonging.
  • Innovative: We embrace change and find better ways to do things together.
  • Globally Connected: We foster collaboration between diverse cultures and perspectives.

If you’re passionate about innovation, collaboration, and making a real impact, we’d love to have you on board!

Role Summary

As a Site Reliability Engineer, you will play a key role in keeping all production systems running smoothly. You will work closely with other engineers and operators to fuse engineering principles, operational knowledge, security, and automation to work towards platform/service production excellence from an angle of infrastructure, reliability, and security.

The SRE team owns the foundation of AI Platform’s Core platform - the services and infrastructure that let us deploy to a multitude of public cloud providers and that powers many ML and LLM powered features. We give every other engineering team a reliable base to build on, and we own the software delivery lifecycle end to end: the tooling, patterns, and automation that reduce friction for the whole org.

This is an exciting opportunity to grow professionally while contributing to a mission-driven organization.

Responsibilities:

What you’ll own

  • Reliability of platform(includes ML and LLM workloads) - model serving and inference infrastructure (GPU-backed endpoints, autoscaling, latency and cost tradeoffs), with SLOs, on-call, and incident response that cover models, not just services
  • Observability(includes ML models) - drift and performance monitoring for ML, plus LLM-specific tracing, evals, and guardrails, wired into the same metrics and logging stacks we run everywhere else
  • Company-wide technical direction: shaping the roadmap and building golden paths that raise the baseline for every team
  • Developer tooling and automation that compounds - reusable GitHub Actions, GitOps workflows, Terraform modules - so every engineer ships faster
  • Reusable components packaging common open-source tools (Grafana, Istio, CloudNative stack, and ML tooling such as model registries and feature stores) for teams to deploy in any environment
  • Secure-by-default infrastructure - baking security, compliance audits, cost governance, and audit trails into the platform in close partnership with our lead/backend/staff engineers.

Requirements

Who you are (must-haves)

  • 5+ years in infrastructure engineering, DevOps, or SRE, operating large-scale, high-availability production systems using Kubernetes
  • Production Operational experience - a live cluster under real load, not a lab. Fluent with Helm, and Terraform or Cloudformation, on at least one major cloud (AWS preferred).
  • Good proficiency in Python or Go or general scripting for automation and tooling(automation with higher language preferred)
  • AI is already in your daily loop - Agentic tooling (Claude Code, Codex, Droid, internal skills) is part of how you ship and not what you are experimenting with. We believe AI tools can be great with human judgement and we want the SRE team to bring the next wave day to day operations.
  • First-principles reasoning - Reasoning from constraints and failure modes naming the tradeoff in business terms (reliability vs. velocity, cost vs. blast radius, standardisation vs. one-off)
  • At least one infrastructure build you owned end to end - with the outcome metric attached (deploy time, MTTR, cost, adoption, availability).
  • Cross-functional strength. Track record working with product, backend/frontend teams to pull through collective initiative.

ML & AI platform (strongly preferred)

  • Running ML workloads on Kubernetes - GPU scheduling, capacity, and cost management
  • Model serving and inference at production scale (eg KServe, RayServe, Triton, vLLM, or similar) with real latency and cost constraints(preferred RayServe)
  • MLOps pipeline tooling - training pipelines, model registries, feature stores, and lineage (Kubeflow, MLflow, Feast, Weights & Biases, or equivalents)
  • LLMOps in production - inference serving, prompt/version management, and LLM observability (tracing, evals, drift, guardrails, cost per request)
  • Governing ML/LLM workloads as platform capabilities: data-residency and PII controls, and audit trails

Any other General requirements

  • Global Collaboration: Ability to work across global teams and different cultures across various time zones with strong communication skills.
  • Problem Solving: Ability to break down complex problems into simple, actionable solutions.
  • Ownership & Drive: Tendency to go above and beyond to meet deadlines, manage own deliverables, and assist team members.
  • Availability: Willingness to support processes for 24x7 operational support.

Benefits

At CloudFactory, we believe that work should be more than just a job-it should be a platform for growth, impact, and community. Here, you’ll earn with purpose, learn every day, and serve a mission that truly matters. If you're looking for a career where you can develop professionally, contribute meaningfully, and be part of a global movement, we’d love to have you on this journey!

Join us today and be part of our mission to connect people and technology for a better world! Apply now and bring your whole, authentic self to work-we can’t wait to meet you!

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
376,406 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
Reading
$200k – $270k per year • In office • Internship • 8+ years exp • PhD • San Francisco
Python
AI/ML
AI Agents
Fine-tuning
LLM
LLM Guardrails
Prompt Engineering
RAG
DevOps
AWS
CI/CD
GCP
Vector
Cybersecurity
GDPR
HIPAA
Apply
$130k – $160k per year • In office • Full-Time • 6+ years exp • Chandler
Python
Python
FastAPI
Databases
PostgreSQL
DevOps
Amazon CloudWatch
Amazon EC2
Amazon ECS
AWS
AWS Fargate
SLI/SLO/SLA
Terraform
Apply
$25k – $86k per year (Estimated) • In office • Full-Time • 3+ years exp • Tokyo
Go
JavaScript
TypeScript
AI/ML
Human-in-the-Loop
Frontend
D3.js
Deck.gl
React.js
Vue.js
Mobile
Clean Architecture
DevOps
AWS
GCP
gRPC
Apply
Senior .NET Developer 10 hours ago
$132k – $219k per year (Estimated) • Remote • Public Trust • Full-Time • 8+ years exp • Bachelor's Degree • United States
C#
JavaScript
Python
SQL
C#
.NET
Entity Framework Core
Databases
DynamoDB
DevOps
Amazon CloudWatch
Amazon Kinesis
Amazon S3
AWS
AWS Lambda
AWS Step Functions
Azure
CI/CD
Configuration Management
Docker
Git
GitHub
GitHub Actions
GitLab
GitLab CI
Kubernetes
Management
Jira
Apply
$90k – $177k per year (Estimated) • In office • Full-Time • Master's Degree • Charlotte
C#
Go
JavaScript
Kotlin
Node JS
Python
SQL
TypeScript
AI/ML
AI Agents
Embeddings
Human-in-the-Loop
LLM
LLM Guardrails
Prompt Engineering
RAG
Frontend
React.js
DevOps
CI/CD
Rest API
Apply
$82k – $184k per year (Estimated) • In office • Full-Time • 2+ years exp • Reading
Python
AI/ML
AI Agents
Claude
Fine-tuning
Human-in-the-Loop
DevOps
AWS
Azure
GCP
Apply
In office • Full-Time • 5+ years exp • Kathmandu
JavaScript
TypeScript
AI/ML
Prompt Engineering
Frontend
esbuild
React.js
Redux
Redux Toolkit
Tailwind CSS
Vite
Webpack
Zustand
Mobile
State Management
DevOps
CI/CD
Git
Rest API
QA
Cypress
Playwright
Postman
Apply
Remote/Hybrid • Full-Time • 3+ years exp • Bachelor's Degree
SQL
AI/ML
Spark
Apply
Software Engineer 17 days ago
Remote/Hybrid • Full-Time • 2+ years exp • Bachelor's Degree
Go
JavaScript
Python
Ruby
TypeScript
Ruby
Ruby on Rails
Databases
Amazon DocumentDB
DynamoDB
PostgreSQL
AI/ML
Prompt Engineering
Frontend
React.js
DevOps
AWS
Azure
CI/CD
Docker
GCP
GitHub Actions
Grafana
Kubernetes
New Relic
GitHub
Apply
Remote/Hybrid • Full-Time • 5+ years exp • Bachelor's Degree
Go
JavaScript
Python
Ruby
TypeScript
Ruby
Ruby on Rails
Databases
Amazon DocumentDB
DynamoDB
PostgreSQL
AI/ML
Prompt Engineering
Frontend
React.js
DevOps
AWS
Azure
CI/CD
CloudFormation
GCP
GitHub Actions
Grafana
Kubernetes
New Relic
Prometheus
Terraform
GitHub
Apply
$72k – $135k per year (Estimated) • Remote/Hybrid • Full-Time • Reading
Apply
$55k – $149k per year (Estimated) • Remote/Hybrid • Full-Time • Reading
Analytics
Power BI
Tableau
Apply
$40k – $54k per year • Remote/Hybrid • Full-Time • Reading
Apply
IMS Core Engineer-3 6 days ago
$31k – $97k per year (Estimated) • Remote/Hybrid • Full-Time • Reading
DevOps
VMWare
Cybersecurity
Wireshark
Management
Jira
Apply
$52k – $135k per year (Estimated) • Remote/Hybrid • Full-Time • Reading
DevOps
CI/CD
Apply
See all jobs
This is one of many
376,406 more open roles from verified company boards, updated every day.