368,746open jobs
9,444companies
47,506added this week
Browse all
Salary
$45k – $132k per year (Estimated)
Location
In office (London)
Seniority
Middle · 3+ years exp
Employment
Full-Time
Overview
Company
Impact
Profile match

Cxm

Capitalising on their many years of expertise across the global markets such as the Asia Pacific, the US, and Europe, CXM Direct&s;s leading team delivered a multi-award-winning hub of services that can satisfy even the most demanding traders. Hav...

Join our Platform & Production Reliability team and help ensure the reliability, performance, and availability of our mission-critical trading systems. As an Application Site Reliability Engineer (SRE), you will own the day-to-day reliability of our .NET/C# services running on Windows, starting with our in-house liquidity bridge that connects MetaTrader trading servers to external liquidity providers. Over time, you will expand your impact across related trading and back-office services.

This is a hands-on role for an engineer who enjoys solving production challenges, improving observability, automating operations, and building resilient systems where uptime directly impacts customer experience.

Position Details

Team Platform & Production Reliability

Location Remote (Americas, LatAm preferred)

Working Hours Americas time zones (UTC-3 to UTC-8)

On-call Rotation aligned with the London trading day

Employment Type Full-time, Permanent

Experience Level Mid-Level (3-5 years)

Technology Stack.NET/C#, Windows Server, AWS, Aurora PostgreSQL, Prometheus, Grafana, Terraform

About the Role

Our trading platform powers every customer interaction, making reliability a first-class product concern. You will be responsible for maintaining and improving the operational reliability of our .NET/C# services on Windows, ensuring they remain highly available, observable, and resilient.

You'll collaborate closely with software engineers to improve monitoring, deployment safety, automation, fault isolation, and incident response, while driving continuous improvements in platform reliability and operational excellence.

What You'll Do

  • Participate in the on-call rotation for production trading systems and lead incident response during service disruptions.
  • Investigate production incidents, perform root cause analysis, and implement preventive actions to eliminate recurring issues.
  • Build and maintain Grafana dashboards, Prometheus alerts, and operational health views across applications, infrastructure, and databases.
  • Instrument .NET services to improve telemetry, metrics, logging, and visibility into service health and customer impact.
  • Define, implement, and monitor Service Level Indicators (SLIs), Service Level Objectives (SLOs), and error budgets.
  • Troubleshoot issues across:
    • .NET/C# applications
    • Windows Server
    • Aurora PostgreSQL databases
    • AWS infrastructure
    • CI/CD pipelines and deployments
  • Improve deployment safety, release automation, and rollback strategies.
  • Partner with developers to improve application operability, resilience, and fault isolation.
  • Automate operational tasks through scripting and infrastructure automation.
  • Create and maintain runbooks, operational documentation, and incident response procedures.
  • Continuously improve monitoring, alert quality, automation, and platform reliability.

Requirements

Required Technical Skills

.NET & Windows

  • Strong experience debugging and supporting .NET/C# applications in production.
  • Hands-on experience with Windows Server environments.

Scripting & Automation

  • Strong PowerShell scripting skills.
  • Experience with Python or Bash.

Observability

  • Experience with Grafana, Prometheus, and Loki (or equivalent monitoring and observability tools).
  • Solid understanding of metrics, logging, tracing, and alerting best practices.

CI/CD & DevOps

  • Experience with modern CI/CD pipelines.
  • Knowledge of deployment strategies, release automation, and rollback mechanisms.

Cloud & Infrastructure

  • Experience working with AWS.
  • Hands-on experience with Terraform or other Infrastructure as Code (IaC) tools.

Databases

  • Experience troubleshooting and supporting Aurora PostgreSQL or other relational database platforms.

Reliability Engineering

  • Practical experience with:
    • SLIs & SLOs
    • Error Budgets
    • Incident Response
    • Root Cause Analysis (RCA)
    • Alert Design
    • Production Operations

Preferred Qualifications

  • Experience supporting high-availability or low-latency financial or trading systems.
  • Familiarity with MetaTrader environments or financial technology platforms.
  • Experience with distributed systems and microservices.
  • Knowledge of OpenTelemetry or similar observability frameworks.
  • Exposure to Docker, Kubernetes, or containerized environments.

Benefits

Why Join Us?

  • Work on mission-critical trading infrastructure that directly impacts customers.
  • Solve challenging reliability and scalability problems in a real-time environment.
  • Build world-class observability, automation, and deployment practices.
  • Collaborate with experienced engineers in a modern engineering culture.
  • Influence reliability strategy and engineering best practices across the platform.

If you're passionate about production engineering, automation, and building reliable systems at scale, we'd love to hear from you.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
368,746 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
London
POS TP.NET Leader 1 day ago
$28k – $73k per year (Estimated) • Remote/Hybrid • 5+ years exp • Bengaluru
SQL
C#
C#
.NET
Databases
MS SQL
DevOps
SLI/SLO/SLA
Management
ServiceNow
Apply
$25k – $42k per year • Equity 0–0.2% • Remote • Full-Time • 3+ years exp
Bash
Go
JavaScript
Python
TypeScript
DevOps
AWS
Azure
CI/CD
Datadog
Docker
GCP
GitHub Actions
GitLab CI
Grafana
Incident Management
Kubernetes
Platform Engineering
Prometheus
Terraform
Amazon CloudWatch
GitHub
GitLab
IAM
Cybersecurity
Least Privilege
Apply
$100k – $210k per year • Equity 0–0.5% • Remote • Full-Time • 3+ years exp • San Francisco
Bash
Go
JavaScript
Python
TypeScript
DevOps
AWS
Azure
CI/CD
Datadog
Docker
GCP
GitHub Actions
GitLab CI
Grafana
Incident Management
Kubernetes
Platform Engineering
Prometheus
Terraform
Amazon CloudWatch
GitHub
GitLab
IAM
Cybersecurity
Least Privilege
Apply
$19k – $28k per year (net) • Remote • Full-Time • Moscow
C#
C++
C++
CMake
DevOps
CI/CD
Git
Management
Jira
Slack
Apply
Team Lead DevOps 1 day ago
$23k – $62k per year (Estimated) • Remote • 5+ years exp • Moscow
Bash
Python
Erlang
Erlang
EMQX
Databases
Apache Kafka
ClickHouse
PostgreSQL
RabbitMQ
Redis
Redpanda
Trino
DevOps
Ansible
AWS
AWX
FinOps
HAProxy
Hetzner
Kubernetes
SLI/SLO/SLA
Terraform
Yandex Cloud
Amazon S3
Apply
Remote/Hybrid • Full-Time • Bachelor's Degree
PowerShell
Python
DevOps
Azure
SLI/SLO/SLA
Cybersecurity
GDPR
Microsoft Entra ID
Wireshark
Apply
Full-Stack QA Engineer 2 months ago
Remote • Full-Time • Bachelor's Degree
JavaScript
Python
TypeScript
QA
Playwright
Apply
Graphic Designer 3 months ago
Remote/Hybrid • Full-Time • Bachelor's Degree • Bangkok
Design
Adobe Photoshop
Apply
$20k – $48k per year (Estimated) • Remote • Full-Time • Bachelor's Degree
Python
DevOps
AWS
CloudFormation
Kubernetes
Nagios
Terraform
Amazon CloudWatch
Amazon ECS
Apply
$65k – $155k per year (Estimated) • In office • Internship • Bachelor's Degree • London
Go
JavaScript
Ruby
Scala
Apply
In office • Internship • 1+ year exp • Bachelor's Degree • London
Go
JavaScript
Ruby
Scala
Apply
$221k – $370k per year • In office • Full-Time • 5+ years exp • London
AI/ML
OpenAI
Apply
$72k – $136k per year (Estimated) • Equity • Remote/Hybrid • 5+ years exp • London
Python
SQL
Databases
Snowflake
AI/ML
AI Agents
DevOps
Kibana
Marketing
Salesforce
Apply
$77k – $148k per year (Estimated) • Equity • In office • Master's Degree • London
JavaScript
Python
Scala
AI/ML
AI Agents
DevOps
GitHub
Apply
See all jobs
This is one of many
368,746 more open roles from verified company boards, updated every day.