431,271open jobs
14,748companies
59,958added this week
Browse all
Salary
$93k – $178k per year (Estimated)
Location
In office (London)
Seniority
Senior
Overview
Company
Impact
Profile match

About Clear Street:

Clear Street’s mission is to give every sophisticated investor access to every asset, in every market, through a unified platform built for speed, transparency and scale.

We give our clients the technology, tools, and service once reserved for the largest institutions, rebuilt with modern infrastructure. Our single, cloud-native, end-to-end capital markets platform powers investor growth today and is transforming how they can interact with markets tomorrow.

For more information, visit https://clearstreet.io.

The Role

As a Production Engineer, you sit at the intersection of software reliability and operational excellence. You own the health, resilience, and recovery of our production systems-while spending equal energy innovating solutions that eliminate human toil, reduce incident blast radius, and raise the reliability bar across the entire platform.  You will partner closely with engineering, operations, and business teams to understand daily pain points and translate them into lasting automated solutions. Half your time is spent in the trenches-supporting production, responding to incidents, and deeply understanding how our

systems behave under real conditions. The other half is yours to build: automation, tooling, and observability platforms that make tomorrow's on-call shift meaningfully easier than today's.  

You will work on challenges like:

  • Design and build comprehensive monitoring and observability platforms that surface the right signal at the right time-eliminating alert fatigue and accelerating root-cause analysis.
  • Develop intelligent automation and self-healing capabilities that diagnose issues, trigger recovery workflows, and reduce mean time to recovery (MTTR) without manual intervention.
  • Analyze incidents, identify systemic trends, and engineer solutions that prevent entire classes of failures from recurring.
  • Build reusable runbooks, diagnostic tooling, and recovery playbooks that turn tribal knowledge into scalable platform capabilities.
  • Create golden-path operational workflows-making the safest, most reliable path also the easiest one for engineering teams to follow.
  • Partner with Platform Engineering to influence CI/CD pipelines, deployment safety, and infrastructure resilience from a production reliability perspective.
  • Champion Infrastructure as Code, GitOps, and SRE best practices while helping teams adopt modern engineering workflows.
  • Continuously measure production health through SLIs/SLOs/SLAs, and drive engineering priorities based on reliability data.
  • Explore emerging technologies-including AI-assisted diagnostics and developer tooling-that transform how we operate production systems.

The Team

We believe resilient systems are built by engineers who understand them end to end.  Our Production Engineering team is the first and last line of defense for our production platform.  We treat reliability as a product, with uptime and engineer experience as our north stars. We combine the discipline of SRE with a builder's mindset: when we see a recurring problem, we build a solution-not a workaround.

You will work across every engineering and operations team to understand failure modes, quantify reliability gaps, and build platform capabilities that scale with the organization. Whether it's reducing MTTR from hours to minutes, building self-service diagnostic tools, or designing proactive alerting that catches issues before customers notice, your work will have immediate, measurable impact.

If you're passionate about making production systems invisible to end users-and you get energy from both firefighting and building the systems that make fires less likely-you'll thrive here.

What We're Looking For

We're looking for engineers who combine operational instinct with a builder's discipline.

You should have:

  • Strong hands-on Python skills-this is your primary language for automation and tooling.
  • Experience in SRE, Production Engineering, Platform Engineering, or a related discipline with direct production ownership.
  • Proven track record of building automation and diagnostic tooling that improved recovery times or reduced operational toil.
  • Deep familiarity with cloud-native technologies-Kubernetes, containers, distributed systems-and how they fail in production.
  • Experience with observability platforms such as Datadog, and a strong intuition for what "good" monitoring looks like.
  • Exposure to Infrastructure as Code (Terraform) and GitOps-based deployment workflows (ArgoCD, GitHub Actions, or similar).
  • Familiarity with the broader technology stack: Java, Go, Kafka, Redis, Snowflake, and Postgres.
  • Strong analytical and problem-solving skills-you thrive on ambiguous, high-stakes production problems.
  • A product mindset applied to operational tooling: you think about usability, adoption, and documentation when building internal solutions.
  • Excellent communication skills and the ability to work fluidly across engineering, operations, and business stakeholders.
  • Self-starter mentality-you identify opportunities, take initiative, and deliver with minimal supervision.
  • Curiosity and a continuous learning mindset; fintech or financial industry background is a plus.

The Technology You'll Work With

You'll operate and build on a modern cloud-native platform that includes:

  • Kubernetes & AWS
  • Terraform & ArgoCD
  • GitHub Actions
  • Kafka, Redis
  • PostgreSQL & Snowflake
  • Datadog
  • Python, Go, Java
  • gRPC & Protobuf
  • Internal Platform APIs and Developer Tooling

What Success Looks Like

Within your first year, you'll have made a measurable impact on production reliability. Success looks like:

  • Reducing mean time to detection (MTTD) and mean time to recovery (MTTR) across key production systems.
  • Building automation that handles a meaningful percentage of incident scenarios without human intervention.
  • Becoming a trusted subject matter expert for core platform components and their failure modes.
  • Delivering observability and diagnostic tools that other engineers actually use and depend on.
  • Establishing SLO baselines and driving engineering investment based on reliability data.
  • Spending less of your time-and your teammates time-on repetitive manual toil.

Your impact won't be measured by the number of incidents you respond to-it will be measured by how reliably our systems run and how quickly we recover when they don't.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
431,271 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
London
Senior AI Engineer 5 hours ago
$84k – $166k per year (Estimated) • Remote • Contractor • 5+ years exp
Go
JavaScript
Java
Kotlin
TypeScript
SQL
C#
Databases
PostgreSQL
Redis
DynamoDB
Apache Kafka
AI/ML
Cursor
Claude
Claude Code
AI Agents
OpenAI Codex
Frontend
GraphQL
React.js
DevOps
gRPC
Terraform
GCP
Azure
AWS
Docker
Apply
$80k – $159k per year (Estimated) • Remote • Full-Time • 5+ years exp
Python
Java
PHP
TypeScript
Databases
PostgreSQL
DevOps
Rest API
Terraform
Helm
CI/CD
ArgoCD
AWS
Docker
Kubernetes
Apply
$53k – $133k per year (Estimated) • Remote • Full-Time
Python
Go
Ruby
Lua
DevOps
Terraform
New Relic
Datadog
Prometheus
Docker
Kubernetes
Nginx
Grafana
Apply
TDT CMOS PI 5 hours ago
$42k – $92k per year (Estimated) • In office • Full-Time • 5+ years exp • Master's Degree • Taichung
Python
MATLAB
Apply
$41k – $67k per year (Estimated) • In office • Internship • Bachelor's Degree • Singapore
Python
SQL
Analytics
Tableau
Power BI
Microsoft Excel
Apply
$85k – $215k per year (Estimated) • In office • London
Python
Java
TypeScript
Java
Spring Boot
Databases
Apache Kafka
DevOps
gRPC
Azure
AWS
Kubernetes
SLI/SLO/SLA
Apply
In office • 6+ years exp • London
JavaScript
TypeScript
Node JS
Frontend
Next.js
React.js
Turborepo
DevOps
Azure
CI/CD
AWS
Kubernetes
SLI/SLO/SLA
Apply
In office • London
Python
Java
TypeScript
Java
Spring Boot
Databases
Apache Kafka
DevOps
gRPC
Azure
AWS
Kubernetes
SLI/SLO/SLA
Apply
$100k – $192k per year (Estimated) • In office • 8+ years exp • London
Java
Rust
Databases
Apache Kafka
DevOps
Azure
AWS
Kubernetes
SLI/SLO/SLA
Apply
$84k – $189k per year (Estimated) • In office • 10+ years exp • London
Apply
$129k – $203k per year • Remote/Hybrid • Full-Time • 10+ years exp • London • Paris
Apply
$92k – $246k per year (Estimated) • Remote/Hybrid • Full-Time • London
Apply
$34k per year • In office • Internship • London
Apply
$105k – $184k per year (Estimated) • In office • Full-Time • 10+ years exp • London • Northampton
Apply
See all jobs
This is one of many
431,271 more open roles from verified company boards, updated every day.