368,634open jobs
9,437companies
50,578added this week
Browse all
Salary
$109k – $212k per year (Estimated)
Location
Remote (United States)
Seniority
Senior · 8+ years exp
Employment
Full-Time
Overview
Company
Impact
Profile match
Filevine is a leading provider of legal case management software. We are committed to helping law firms streamline their workflows and improve their efficiency. Filevine is a global organization with a mission to help law firms streamline their workflows and improve their efficiency.

Responsibilities

  • Design and improve the monitoring, logging, distributed tracing, dashboards, alerting, SLIs,

and SLOs that give teams meaningful visibility into production health and customer impact.

  • Build and maintain automation, internal tools, and CI/CD systems that increase engineering

efficiency, reduce toil, and support reliable deployments at scale. Take responsibility for the

quality and reliability of tools and services you support.

  • Drive the implementation and continuous improvement of reliable systems for building,

deploying, testing, and operating Filevine products, proactively identifying and resolving

reliability, performance, scalability, and security risks before they impact customers.

  • Own complex production incidents through detection, triage, communication, resolution,

and follow-up. Turn incident learning into durable corrective actions, stronger runbooks and

operating practices, and improvements that reduce recurring incidents and operational

burden.

Lead significant technical initiatives from problem definition and design through

implementation and adoption. Coordinate work across engineers and teams, communicate

tradeoffs and risks, and help ensure the work delivers the intended results.

  • Mentor other Site Reliability Engineers through design reviews, incident follow-ups, paired

problem-solving, and meaningful delegation. Help engineers develop stronger technical

judgment and become increasingly capable of handling complex production work

independently.

  • Participate in the shared on-call rotation and help ensure production systems are prepared

to operate reliably at scale through capacity planning, operational readiness, and

continuous improvements to resilience and recovery.

  • Apply AI and machine learning to analyze operational signals, identify patterns, forecast

reliability and capacity risks, and implement improvements that make systems more

reliable, efficient, and easier to operate.

Qualifications

  • 8+ years of hands-on experience in software engineering, cloud infrastructure, platform

engineering, DevOps, or related technical roles, including at least 5 years in a Site Reliability

Engineering or reliability-focused role.

  • Strong knowledge of distributed systems and hands-on experience operating Kubernetes

workloads and cloud infrastructure in AWS or a comparable platform, with proficiency in

Infrastructure as Code, monitoring, logging, alerting, distributed tracing, SLIs, and SLOs.

  • Strong proficiency with Python, Go, Bash, or a similar language, with demonstrated

experience building and maintaining production tooling, automation, CI/CD pipelines, and

deployment systems that reduce toil, improve reliability, and simplify ongoing operations.

  • Demonstrated ability to lead troubleshooting, incident response, root cause analysis, and

long-term reliability improvements for complex production systems, including the

elimination of recurring incidents and operational work.

  • Proven ability to mentor Site Reliability Engineers, help others build stronger technical

judgment, communicate clearly with technical and business stakeholders, and lead complex

initiatives from planning through delivery.

  • Demonstrated experience applying AI and machine learning to operational data and

engineering workflows to identify patterns, forecast reliability or capacity risks, and

implement measurable improvements with appropriate safeguards.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
368,634 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
United States
$23k – $57k per year (Estimated) • Remote • Full-Time • 6+ years exp
Python
Databases
OpenSearch
DevOps
AIOps
ArgoCD
AWS
Azure
CI/CD
Datadog
Docker
Dynatrace
GCP
GitHub Actions
Grafana
Incident Management
Jaeger
Jenkins
Kubernetes
New Relic
OpenTelemetry
Platform Engineering
Prometheus
SLI/SLO/SLA
Splunk
Terraform
GitHub
Apply
$37k – $80k per year (Estimated) • Remote/Hybrid • Full-Time • 7+ years exp • Pune
PowerShell
Python
AI/ML
Anomaly Detection
DevOps
AppDynamics
AWS
Azure
CI/CD
GCP
Grafana
Incident Management
Prometheus
Splunk
Apply
$27k – $69k per year (Estimated) • Remote • Full-Time
JavaScript
Node JS
Python
TypeScript
AI/ML
Embeddings
LLM
RAG
LLM Guardrails
AI Agents
Function Calling
Model Context Protocol
Frontend
Angular
DevOps
AWS
Azure
CI/CD
Docker
GCP
Vector
Apply
$23k – $58k per year (Estimated) • In office • Full-Time • 7+ years exp • Gurgaon
Python
Databases
Apache Kafka
ElasticSearch
Redis
DevOps
Error Budget
Grafana
Incident Management
Kubernetes
OpenShift
Platform Engineering
Prometheus
Self-Healing
Cybersecurity
HashiCorp Vault
Cryptography
Vault
Apply
$36k – $78k per year (Estimated) • Remote/Hybrid • Full-Time • 6+ years exp • Pune
PowerShell
Python
SQL
DevOps
AppDynamics
AWS
Azure
GCP
Grafana
Prometheus
Splunk
Apply
$99k – $201k per year (Estimated) • In office • Full-Time • 6+ years exp • Bachelor's Degree • Salt Lake City
SQL
Databases
Snowflake
Analytics
Power BI
Tableau
Marketing
HubSpot
Salesforce
Apply
Remote/Hybrid • Contractor • 3+ years exp • Prague
Python
TypeScript
JavaScript
Python
Django
Frontend
React.js
Svelte
DevOps
AWS
Kubernetes
Apply
Remote/Hybrid • Contractor • 3+ years exp • Bratislava
Python
TypeScript
JavaScript
Python
Django
Frontend
React.js
Svelte
DevOps
AWS
Kubernetes
Apply
$76k – $164k per year (Estimated) • In office • Full-Time • 3+ years exp • Salt Lake City
Management
Slack
Apply
$175k – $195k per year • Remote • Full-Time • 8+ years exp
Python
DevOps
AWS
CI/CD
Incident Management
Kubernetes
Platform Engineering
SLI/SLO/SLA
Apply
$78k – $130k per year • In office • Full-Time • 4+ years exp • Bachelor's Degree • Cary • San Jose
Analytics
Power BI
Apply
$91k – $182k per year (Estimated) • In office • Full-Time • 5+ years exp • Bachelor's Degree • United States
Design
SolidWorks
Apply
$45k – $60k per year • Remote • Full-Time • 2+ years exp • PhD • United States
Cybersecurity
HIPAA
Apply
$117k – $258k per year (Estimated) • Equity • Remote • Full-Time • United States
C++
Java
Python
Cybersecurity
Crowdstrike
Apply
$90k – $184k per year (Estimated) • Equity • Remote • Full-Time • 3+ years exp • United States
AI/ML
Red Teaming
Cybersecurity
Burp Suite
Cobalt Strike
Crowdstrike
Metasploit
MITRE ATT&CK
Nessus
Nmap
Apply
See all jobs
This is one of many
368,634 more open roles from verified company boards, updated every day.