Role Summary:
As a Site Reliability Engineer you will be embedded with a cross functional team who has key responsibilities for certain portions of our systems. Over the course of the first year you will gain the valuable context needed to be truly effective and move at speed in the Filevine environment. During your successive years you will be given specific mission critical objectives that help build out and improve our autonomous systems and simultaneously build out your personal brand as an exceptional engineer who has built and maintained amazing systems that can grow to internet scale.
Responsibilities
Observability & Alerting: Design and improve monitoring, logging, tracing, dashboards, and SLI/SLOs for production visibility.
Automation & CI/CD: Build internal tools and delivery pipelines to boost efficiency, eliminate toil, and ensure reliable deployments.
Reliability & System Quality: Drive continuous improvements in system performance, scalability, and security to mitigate customer impact.
Incident Management: Lead production incident response from triage to resolution, turning lessons into durable runbooks and preventatives.
Technical Leadership: Guide major technical initiatives, align cross-team engineering efforts, and manage technical risks.
Mentorship: Elevate SRE team capability through design reviews, paired problem-solving, and incident post-mortems.
Operations & On-Call: Join the on-call rotation while leading capacity planning and disaster recovery readiness.
AI/ML Operationalization: Leverage operational AI/ML tools to forecast capacity risks, detect patterns, and automate system health.
Qualifications
Experience: 8+ years in software/platform engineering or DevOps, including 5+ years dedicated to Site Reliability Engineering.
Infrastructure & Observability: Expertise in cloud platforms (AWS), Kubernetes, IaC, and full-stack observability (tracing, logging, SLI/SLOs).
Automation & Scripting: Proficient in Python, Go, or Bash for building CI/CD pipelines, production tools, and toil-reducing automation.
Incident Leadership: Proven track record in root cause analysis, high-severity incident response, and long-term reliability engineering.
Leadership & Mentorship: Strong communication skills with a history of mentoring engineers, leading cross-functional projects, and setting technical strategy.
Operational AI/ML: Hands-on experience using AI/ML on telemetry data to predict capacity risks, spot anomalies, and optimize system workflows.

