368,634open jobs
9,437companies
50,578added this week
Browse all
Salary
$120k – $140k per year
Location
Remote/Hybrid (Wilmington, United States)
Employment
Full-Time
Overview
Company
Impact
Profile match
Best Egg Best Egg is a leading online lender that provides personal loans, credit cards, and other financial products to consumers. We are committed to providing our customers with the best possible rates and terms, and we strive to make the borrowing process simple and transparent. We believe that everyone deserves access to affordable and convenient financial products.

The Job

As Site Reliability Engineer, you will serve as a reliability subject matter expert who leads major incident recovery, drives observability and reliability improvements, mentors associate engineers, reduces operational toil, influences technical decisions, and improves resiliency standards.

This role requires depth across production systems, telemetry, batch operations, automation, and incident response. You will be expected to guide technical direction for reliability improvements and help teams prevent recurring failures.

Employees joining Best Egg's Information Technology organization can expect a culture centered on Continuous Delivery, Total Quality Management, Knowledge Sharing, Personal and Career Advancement, Empowerment, Innovation, and Collective Ownership.

Duties & Responsibilities

    • Lead technical recovery efforts for major incidents, coordinating triage, evidence review, restoration actions, and validation.
    • Optimize observability strategy, alert quality, dashboard standards, and telemetry coverage across multiple services.
    • Drive reliability initiatives that reduce recurring failures, noisy alerts, manual work, and operational risk.
    • Mentor associate engineers on troubleshooting methods, RCA evidence, runbook quality, and production support judgment.
    • Influence engineering decisions by identifying reliability risks, missing telemetry, supportability gaps, and resiliency patterns.
    • Improve JAMS, GoAnywhere, Datadog, xMatters, and service support practices through automation and standards.
    • Partner with leaders and technical teams to prioritize remediations based on customer impact, business impact, and operational exposure.

Required

    • Hands-on familiarity with production support, monitoring, alerting, and incident response practices.
    • Working knowledge of Datadog dashboards, monitors, logs, metrics, and APM concepts.
    • Ability to troubleshoot application, infrastructure, batch, or file transfer issues using runbooks and telemetry.
    • Exposure to AWS or cloud operations and scripting with Python, PowerShell, Bash, or similar tools.
    • Clear communication skills during incidents, service requests, and post-incident follow-through.
    • Strong experience leading production incident recovery and cross-system reliability investigations.
    • Ability to mentor engineers and influence technical decisions without direct authority.

Recommended

    • Datadog, AWS, ITIL, Linux, or automation certification.
    • Experience with JAMS, GoAnywhere, xMatters, ServiceNow/Jira, or CI/CD environments.
    • Exposure to AIOps, anomaly detection, operational automation, or reliability engineering.
    • Familiarity with financial services controls, secure file transfer, or regulated operations.

Skill Set

    • Serves as a reliability SME across observability, incident response, batch operations, and operational platforms.
    • Leads major incident recovery with calm command of telemetry, dependencies, impact, and restoration options.
    • Reduces operational toil through automation, standards, better alerting, and durable remediation.
    • Mentors engineers and improves the quality of technical support practices across the team.
    • Influences design and readiness decisions that improve resiliency and operational resilience.

Success Metrics in First 90 Days

    • Lead a major incident or complex reliability investigation with clear recovery and follow-through.
    • Deliver an observability or automation improvement that measurably reduces alert noise, toil, or repeat issues.
    • Mentor associate engineers through troubleshooting reviews, runbook improvements, or incident debriefs.
    • Identify and influence remediation of a meaningful resiliency or supportability gap.
    • Improve standards or patterns for telemetry, escalation, batch support, or operational validation.
Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
368,634 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
Wilmington
$76k – $165k per year (Estimated) • In office • Full-Time • Bachelor's Degree • Canada
PowerShell
SQL
Databases
MS SQL
MySQL
Oracle
PostgreSQL
DevOps
AWS
Azure
CI/CD
Analytics
ETL/ELT
Apply
In office • Internship • 1+ year exp • Bachelor's Degree • Farmington Hills
Python
Rust
TypeScript
AI/ML
Time Series Forecasting
DevOps
AWS
CI/CD
Git
gRPC
IoT
MQTT
Apply
$140k – $263k per year (Estimated) • Remote
DevOps
Ansible
CI/CD
GitOps
Kubernetes
OpenShift
Rancher
Terraform
Apply
$180k – $225k per year • Equity • In office • Full-Time • 4+ years exp • Bachelor's Degree • Seattle
C#
Go
Java
Kotlin
TypeScript
JavaScript
AI/ML
Claude
Claude Code
OpenAI Codex
Frontend
React.js
DevOps
AWS
AWS Step Functions
Apply
$14k – $35k per year (Estimated) • Remote/Hybrid • Full-Time • Bachelor's Degree • Phoenix
SQL
AI/ML
AI Agents
Edge AI
DevOps
AWS
Apply
$120k – $140k per year • Remote/Hybrid • Full-Time • 3+ years exp • Master's Degree • Wilmington
Python
SQL
AI/ML
ChatGPT
Claude
NumPy
Scikit-learn
Apply
$100k – $135k per year • Remote/Hybrid • Full-Time • 4+ years exp • Wilmington
Java
Python
SQL
Databases
Amazon DocumentDB
DynamoDB
PostgreSQL
DevOps
Amazon EKS
AWS
CI/CD
Datadog
Docker
Kubernetes
Amazon CloudWatch
Amazon ECS
Apply
$85k – $100k per year • Remote/Hybrid • Full-Time • 3+ years exp • Bachelor's Degree • Wilmington
Python
SQL
Apply
$150k – $165k per year • Remote/Hybrid • Full-Time • 12+ years exp • Wilmington
Marketing
Zendesk
Apply
$90k – $105k per year • Remote/Hybrid • Full-Time • 1+ year exp • Bachelor's Degree • Wilmington
Python
SQL
Apply
IAM Coordinator 1 hour ago
$60k – $70k per year • Remote/Hybrid • Full-Time • 3+ years exp • Wilmington
DevOps
IAM
Cybersecurity
Least Privilege
Apply
$142k – $224k per year • In office • Full-Time • Bachelor's Degree • Wilmington
Apply
Lead AI Engineer 8 hours ago
$186k – $286k per year • Equity • In office • Full-Time • 10+ years exp • Bachelor's Degree • Wilmington
Java
Python
C#
TypeScript
JavaScript
Java
Spring Boot
C#
.NET
AI/ML
Claude
Copilot
LLM
Anthropic
OpenAI
Frontend
Angular
DevOps
AWS
Azure
CI/CD
Docker
GCP
GitLab CI
Jenkins
Kubernetes
OpenShift
Rest API
GitLab
Apply
$120k – $150k per year • In office • Full-Time • 5+ years exp • Bachelor's Degree • Wilmington
Management
Confluence
Jira
Trello
Apply
ML Engineer 1 day ago
$180k – $220k per year • In office • Full-Time • 10+ years exp • Bachelor's Degree • Wilmington
Python
SQL
AI/ML
LightGBM
LLM
PyTorch
Scikit-learn
Spark
TensorFlow
XGBoost
Amazon SageMaker
Feature Store
LLM Guardrails
DevOps
AWS
Analytics
A/B Testing
Apply
See all jobs
This is one of many
368,634 more open roles from verified company boards, updated every day.