606,098open jobs
31,949companies
86,799added this week
Browse all
Salary
$117k – $223k per year (Estimated)
Location
In office (Lehi)
Seniority
Senior · 5+ years exp
Overview
Company
Impact
Profile match

MX is a fintech company on a mission to empower the world to be financially strong. We build technology that helps banks, credit unions, and fintechs deliver smarter, more intuitive financial experiences to millions of people.

Like many startups, we’ve navigated real growth challenges - and we’ve come out stronger on the other side. Today, MX is in a phase of renewed momentum and scale, with a solid foundation and a clear vision for what’s next. This is a place where thoughtful execution matters, innovation is encouraged, and individuals have real ownership over their work.

Our culture values curiosity, accountability, and impact. We give people the space to question assumptions, design better solutions, and help shape how the company grows. If you’re looking to do meaningful work, influence outcomes, and grow alongside a company that’s ready to move fast, you’ll feel at home at MX.

At MX, reliability is a product. Our infrastructure powers financial applications used by millions of people and processes billions of transactions for major financial institutions, and customers feel every second of downtime.

We're building a new observability function that runs the way we run incident response: the system does the heavy lifting, and people handle judgment, customers, and the exceptions. As a Senior Observability Engineer, you build and operate an observability control plane. You scaffold baselines, score coverage, and turn every real incident into the detection the platform should have caught. This is a multiplier role: you raise the bar for every team through standards and automation instead of building each team's dashboards by hand.

We call it the shepherd model. You shepherd Datadog and partner with our product engineering teams so they observe the right signals for their products. Service owners get real signal instead of noise, and leadership gets coverage and health as a program metric.

This role shares the team pager. Observability and incident response run one on-call roster. You take shifts with the rest of the team and act as Incident Commander when an incident needs one. It is core to the role, not an afterthought.

Engineering at MX runs hybrid infrastructure (AWS and bare metal) with services in Ruby, Go, and Java, messaging over NATS and RabbitMQ, and data on PostgreSQL and Redis. Datadog is our observability platform and incident.io is our incident response platform.

What you'll do:

  • Build and operate an observability control plane: automate baseline monitors, dashboards, and tagging standards through the Datadog API and Terraform.

  • After significant incidents, produce detection and dashboard gap packs grounded in Datadog and MX investigation patterns, with queries ready to apply.

  • Define what "good" looks like for a Ruby, Go, or Java service on Datadog (tags, golden signals, alert quality, dashboard contracts), then audit services against that standard and accept or reject readiness.

  • Validate, don't own. Service owners keep their alerts and dashboards; you confirm they are complete and correct, then move on. Escalate to engineering managers when coverage fails or an owner is missing.

  • Own the monthly observability and service-catalog health report: departed owners, stale dashboards, services with no monitors, SLO gaps, and coverage trends.

  • Run maturity assessments (baseline through SLO, launch-ready, self-serve) and track them over time.

  • Tune alerting toward zero false SEV1/2 pages and actionable SEV3/4 alerts, and coach teams on Datadog cost and cardinality.

  • Build self-serve onboarding so new services get baseline observability on day one, without a multi-week embed.

  • Share the team pager. Rotate on the shared IR & Observability on-call, triage and investigate live incidents with Datadog and MX investigation patterns, and take Incident Commander or supporting technical roles as the incident needs.

  • After incidents, close the detection loop (gap packs, new monitors, dashboards) so the pager gets quieter over time.

  • Run high-value launch and production-readiness reviews as a checkpoint, not a permanent staffing model.

Basic Requirements

  • BS in Computer Science or equivalent experience

  • 5+ years running production observability, SRE, or DevOps

  • 5+ years automation-first engineering in Python, Bash, Go, and/or Terraform, plus Kubernetes proficiency

  • AI- and workflow-literate. You've used or built scripted and AI-assisted workflows to scale reviews, audits, and docs

  • Distributed-systems debugging across microservices: latency, connection pools, queues, and cascading failure on Kubernetes and bare metal, with NATS, RabbitMQ, Postgres, and Redis

  • Shared on-call, Incident Commander-capable

Preferred Requirements

  • Fintech experience with MX-like architectures

  • Datadog preferred; strong Grafana/Prometheus, Splunk, or New Relic experience counts if you can ramp on Datadog fast

  • Google SRE practices: toil elimination, incident management, automation for self-healing

  • Cross-functional influence without authority. You've improved teams that don't report to you

  • Governance and reporting: you can produce a monthly health and compliance report leadership reads (orphans, stale entries, gaps, trends)

  • OpenTelemetry instrumentation

  • Incident response platforms (incident.io, PagerDuty, OpsGenie); prior formal Incident Commander experience

  • Golang and Ruby on Rails (the MX stack)

MX is proudly committed to recruiting and retaining a diverse and inclusive workforce. As an Equal Opportunity Employer, we never discriminate based on race, religion, color, national origin, gender (including pregnancy, childbirth, or related medical conditions), sexual orientation, gender identity, gender expression, age, military or veteran status, status as an individual with a disability, or other applicable legally protected characteristics. We particularly welcome applications from veterans and military spouses. All your information will be kept confidential according to EEO guidelines. You may request reasonable accommodations by sending an email to [email protected].

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
606,098 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
Lehi
$98k – $244k per year (Estimated) • In office • Full-Time • 5+ years exp • Bachelor's Degree • Tunis
Python
Python
SQLAlchemy
FastAPI
Pydantic
Databases
PostgreSQL
AI/ML
Model Context Protocol
Prompt Engineering
AI Agents
LLM
DevOps
Kong
Azure
CI/CD
Docker
Kubernetes
API Gateway
Cybersecurity
Microsoft Entra ID
Apply
$66k – $152k per year (Estimated) • In office • Full-Time • Bachelor's Degree • Plano
Python
Java
SQL
Python
Flask
FastAPI
Java
Spring Boot
Databases
Pinecone
FAISS
OpenSearch
AI/ML
LangGraph
AutoGen
LangChain
Embeddings
Prompt Engineering
AI Agents
AWS Bedrock
TensorFlow
PyTorch
CrewAI
LLM
RAG
OpenAI
Amazon SageMaker
LLMOps
Google AI Studio
DevOps
Rest API
GCP
Azure
CI/CD
Git
AWS
Docker
Management
Agile
Scrum
Apply
In office • TS/SCI • 2+ years exp
JavaScript
Java
TypeScript
SQL
Java
Apache Tomcat
Databases
PostgreSQL
Oracle
Apache Kafka
Frontend
Angular
DevOps
ZooKeeper
CI/CD
Jenkins
Git
Docker
Kubernetes
Management
Agile
QA
Selenium
Apply
In office • 3+ years exp • PhD
Python
SQL
DevOps
GCP
Docker
Apply
$56k – $164k per year (Estimated) • In office • Full-Time • 3+ years exp • Bachelor's Degree • Tunis
Python
AI/ML
Model Context Protocol
AI Agents
DevOps
Helm
Azure DevOps
GitHub Actions
Azure
CI/CD
GitOps
ArgoCD
AWS
Docker
Kubernetes
Amazon EKS
Azure AKS
Apply
$136k – $244k per year (Estimated) • In office • 8+ years exp • Bachelor's Degree • Lehi
Apply
$79k – $180k per year (Estimated) • In office • 4+ years exp • Lehi
Apply
In office • Bachelor's Degree • Lehi
Marketing
Salesforce
HubSpot
Apply
$73k – $156k per year (Estimated) • In office • 3+ years exp • Lehi
SQL
Analytics
Domo
Marketing
Salesforce
Marketo
Apply
$119k – $227k per year (Estimated) • In office • 5+ years exp • Lehi
Go
Java
Ruby
DevOps
Terraform
GCP
Crossplane
CI/CD
Kubernetes
Chaos Engineering
Apply
$136k – $244k per year (Estimated) • In office • 8+ years exp • Bachelor's Degree • Lehi
Apply
$79k – $180k per year (Estimated) • In office • 4+ years exp • Lehi
Apply
In office • Bachelor's Degree • Lehi
Marketing
Salesforce
HubSpot
Apply
$73k – $156k per year (Estimated) • In office • 3+ years exp • Lehi
SQL
Analytics
Domo
Marketing
Salesforce
Marketo
Apply
$119k – $227k per year (Estimated) • In office • 5+ years exp • Lehi
Go
Java
Ruby
DevOps
Terraform
GCP
Crossplane
CI/CD
Kubernetes
Chaos Engineering
Apply
See all jobs
This is one of many
606,098 more open roles from verified company boards, updated every day.