368,657open jobs
9,442companies
50,883added this week
Browse all
Salary
$96k – $227k per year (Estimated)
Location
Remote/Hybrid (EU, Europe)
Seniority
Staff
Employment
Full-Time
Overview
Company
Impact
Profile match
Wand is the world’s first agentic labor infrastructure provider built for governments and global enterprises. Wand builds systems that allow AI agents to collaborate alongside humans — and is already operating at scale inside some of the world’s largest organizations.

Build the Future Workforce

Wand turns AI into labor. It enables humans and AI agents to operate together as a unified, hybrid workforce, with comprehensive management and oversight. And it’s already operating at scale inside some of the world’s largest organizations.

Wand built the world’s first Agentic Labor Infrastructure enabling governments and global enterprises to create, manage, and scale digital workforces.

Our mission is to integrate agent ecosystems into the core of work and business, unlocking a generational leap in the global economy. We’re building the infrastructure that lets humans and AI agents operate together safely, transparently, and at scale.

Join Wand in leading the Agentic Shift

Wand is building a high-performing global team who take full ownership of what they build. We lead by example, move fast, make data-aware decisions, and continuously push for more- always with a focus on delivering real value to customers.

You would be joining a world-class team that combines deep research expertise and real-world product execution, with experience spanning Deepmind, Google, Amazon, Miro, Elise AI, IBM and Accern.

Position Summary

We are hiring for a highly experienced Senior Staff SRE Engineer to act as a senior technical authority within our reliability function.

This is a deeply hands-on individual contributor role, to build and operate SRE practices at scale. You will design and evolve resilient infrastructure, drive reliability across multiple engineering streams, and ensure our AI-driven products operate with high availability, performance, and security.

You will work across platform, product, data, and ML teams, helping us productionise models, absorb and standardise customer environments, strengthen Kubernetes-based architecture, and mature our CI/CD pipelines end-to-end.

You will also collaborate with other Staff engineers and Architects to shape the global product architect and technology vision.

Responsibilities

  • Architect, deploy, and operate scalable, secure production environments (AWS preferred).

  • Lead reliability improvements across multiple engineering streams.

  • Design and evolve Kubernetes-based infrastructure, including migration and optimisation initiatives.

  • Build and enforce strong Infrastructure-as-Code standards.

  • Define and operationalise SLIs, SLOs, and error budgets.

  • Strengthen observability across applications, infrastructure, data pipelines, and ML systems.

  • Work closely with product and data teams to integrate model analytics and product telemetry into reliability insights.

  • Work across and optimise the entire CI/CD pipeline, from build to deploy to rollback.

  • Improve release safety, deployment frequency, and predictability of SLAs.

  • Lead incident response for complex cross-system failures and drive postmortems.

  • Reduce operational toil through automation and platform engineering improvements.

  • Design processes and tooling to absorb, standardise, and troubleshoot customer environments.

  • Support and productionise ML workloads (MLOps practices including model deployment, monitoring, retraining workflows).

  • Ensure infrastructure aligns with enterprise-grade security and regulatory requirements.

  • Mentor engineers and raise the overall reliability bar across teams.

Key Requirements

  • Extensive hands-on experience in SRE or Production Engineering roles.

  • Demonstrated experience building or scaling SRE practices in high-growth or complex environments.

  • Deep expertise in AWS or Azure-based cloud infrastructure.

  • Strong experience with Kubernetes (including migration, scaling, and production hardening).

  • Advanced Infrastructure-as-Code experience (Terraform or equivalent).

  • End-to-end CI/CD pipeline design and optimisation experience.

  • Strong experience with observability tooling across distributed systems.

  • Experience troubleshooting complex multi-tenant or customer-hosted environments.

  • Experience supporting production data platforms and ML systems.

  • MLOps experience, including model deployment and monitoring.

  • Strong understanding of distributed systems, scalability, and fault tolerance.

  • Systems thinker who understands interactions across infrastructure, product, data, and ML.

  • Excellent communication skills and ability to work cross-functionally.

Preferred Experience

  • Experience in large-scale global B2B/B2C products.

  • Experience working with AI/ML systems, NLP, or LLM-based products.

  • Experience integrating product analytics and model performance metrics into operational monitoring.

  • Background in enterprise environments with strong security and compliance requirements.

  • Experience implementing regulatory controls within cloud infrastructure.

  • Experience scaling infrastructure during rapid growth phases.

  • Experience evaluating infrastructure tooling and vendors.

  • Experience in collaborating with large scale enterprise customers to deploy and operate environments within their accounts and VPCs.

Personal Characteristics

  • Strong problem solver who anticipates failure modes.

  • High ownership mentality and accountability.

  • Comfortable working across streams and influencing without formal authority.

  • Learning-oriented with a drive for continuous improvement.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
368,657 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
In your city
$84k – $178k per year (Estimated) • In office • Full-Time • 10+ years exp • Wellington
Java
Python
DevOps
Ansible
AWS
Azure
CI/CD
Docker
GCP
Helm
Kubernetes
Platform Engineering
Prometheus
Service Mesh
Terraform
GitLab
IAM
Apply
Platform Engineer 1 day ago
$87k – $140k per year • In office • Full-Time • 3+ years exp • Berlin
Databases
PostgreSQL
Redis
DevOps
AWS
Azure
Bicep
CI/CD
Docker
GCP
GitHub Actions
Kubernetes
OpenShift
Terraform
GitHub
Apply
Founding Engineer 1 day ago
$81k – $116k per year • In office • Full-Time • Bachelor's Degree • Munich
JavaScript
Python
TypeScript
Databases
MySQL
PostgreSQL
Frontend
Next.js
React.js
Tailwind CSS
DevOps
AWS
Azure
CI/CD
Docker
GCP
Grafana
Kubernetes
OpenTelemetry
Prometheus
Apply
$98k – $195k per year (Estimated) • In office • Full-Time • 7+ years exp • Wellington
Java
Python
SQL
Java
Spring Boot
Databases
Apache Kafka
Databricks
Neo4j
AI/ML
Flink
Spark
Frontend
GraphQL
DevOps
Azure
CI/CD
Datadog
Dynatrace
Kibana
Kubernetes
OpenShift
Platform Engineering
Splunk
Amazon ECS
Apply
$123k – $251k per year (Estimated) • In office • Full-Time • 5+ years exp • Bachelor's Degree • Dallas • Denver • Birmingham
Java
SQL
Java
Gradle
Hibernate
Maven
Spring Boot
Spring Framework
Databases
Apache Kafka
MySQL
Redis
DevOps
CI/CD
Dynatrace
Jenkins
Kubernetes
OpenShift
Cybersecurity
SonarQube
Apply
In office • Full-Time
JavaScript
Python
TypeScript
AI/ML
AI Agents
CrewAI
LangChain
LangGraph
LLM
Prompt Engineering
RAG
DevOps
Rest API
Management
Miro
Apply
$137k – $268k per year (Estimated) • Remote • Full-Time • 7+ years exp • Bachelor's Degree
AI/ML
LLM
RAG
AI Agents
Apply
$162k – $329k per year (Estimated) • Remote/Hybrid • Full-Time • PhD
Databases
pgvector
Pinecone
Weaviate
PostgreSQL
AI/ML
AI Agents
CUDA
CUDA Toolkit
Fine-tuning
LangChain
LangGraph
LangSmith
LlamaIndex
LLM
Ragas
TruLens
Anthropic
Context Engineering
LLM Evaluation
OpenAI
Management
Miro
Apply
$117k – $244k per year (Estimated) • Remote/Hybrid • Full-Time
Go
Node JS
JavaScript
AI/ML
AI Agents
Claude
Claude Code
Management
Miro
Apply
$68k – $193k per year (Estimated) • Remote/Hybrid • Full-Time
Go
Node JS
TypeScript
JavaScript
AI/ML
AI Agents
Claude
Claude Code
Management
Miro
Apply
See all jobs
This is one of many
368,657 more open roles from verified company boards, updated every day.