368,941open jobs
9,452companies
47,951added this week
Browse all
Location
In office
Seniority
Senior · 6+ years exp
Overview
Company
Impact
Profile match
KLDiscovery provides technology-enabled services and software to help law firms, corporations, government agencies and consumers solve complex data challenges. The company, with 1,000+ employees in 30+ locations across 18 countries, is a global leader in delivering best-in-class eDiscovery, information governance and data recovery solutions to support the litigation, regulatory compliance, internal investigation and data recovery and management needs of our clients. Serving clients for over 30 years, KLDiscovery offers data collection and forensic investigation, early case assessment, electronic discovery and data processing, application software and data hosting for web-based document reviews, and managed document review services.

The Senior Infrastructure Operations Engineer owns the design, administration, and continuous improvement of KLDiscovery's compute, storage, and cloud infrastructure globally. This role takes end-to-end technical ownership of the physical server, virtualization, block/file/object storage, enterprise backup, and Azure IaaS environments underpinning KLDiscovery's production client systems, and serves as the primary technical escalation point and design authority within the Compute & Storage team. The Senior Infrastructure Operations Engineer leads infrastructure design decisions, drives standards development, partners with Enterprise Architecture and IT Security on architecture and compliance, and provides mentoring and technical direction to Infrastructure Operations Engineers. This role participates in 24x7x365 on-call rotation.

Key Responsibilities

Compute, Virtualization & Storage Architecture:

  • Hold end-to-end technical ownership of KLDiscovery’s physical server and virtualized compute environments; define and enforce configuration standards, capacity thresholds, and operational procedures; lead the design of significant compute changes, cluster expansions, and platform upgrades; review and approve significant configuration changes before implementation

  • Own the design standards, capacity planning, performance management, and operational procedures for block, file, and object storage environments globally; monitor performance and capacity; identify constraints and drive remediation before they impact production; engage vendor support directly for complex issues

Backup, Recovery & Azure IaaS:

  • Own the enterprise backup platform and recovery strategy - job standards, recovery validation procedures, and RPO/RTO alignment; ensure recovery procedures are documented, tested on a defined schedule, and executable by any team member; manage the formal backup coverage request process for consuming teams

  • Own the administration and governance of Azure IaaS infrastructure; define Azure infrastructure standards in conjunction with Enterprise Architecture; monitor and govern cloud costs; surface anomalies and optimization opportunities proactively; partner with Enterprise Architecture on hybrid cloud architecture direction and roadmap input

OS Standards, Patching & Provisioning:

  • Own Windows Server and Linux configuration standards, patching cadence, and hardening baselines across all infrastructure globally; own the server patching function for all server infrastructure - end-user endpoints are excluded and owned by Enterprise Platforms

  • Define and own the server build runbook, provisioning standards, sizing guidelines, and handoff checklists in conjunction with Enterprise Architecture; ensure the provisioning process delivers a configured, network-connected OS at the domain-join boundary with sufficient fidelity to eliminate rework at handoff

Security, Observability & Escalation:

  • Embed security controls into infrastructure design from inception - network segmentation, least-privilege access, encryption at rest and in transit, and audit logging - in conjunction with Enterprise Architecture and IT Security; support audit and compliance reviews as needed; maintain working familiarity with applicable security and compliance frameworks (ISO 27001, CIS Controls) as applied to infrastructure configuration and access management

  • Partner with the Automation & Observability team to ensure all owned infrastructure is covered by monitoring and alerting; serve as the primary escalation point within the Compute & Storage team for complex or time-critical infrastructure incidents; participate in the 24x7x365 on-call rotation; conduct root cause analysis for significant incidents, lead post-incident reviews and blameless retrospectives, and drive systemic remediation

Continual Service Improvement & Automation:

  • Evaluate existing infrastructure for improvement, consolidation, and modernization opportunities; make specific, cost-weighted technical recommendations to management for platform upgrades, replacements, or cloud migrations

  • Partner with the Automation & Observability team to identify, prioritize, and define requirements for infrastructure automation; serve as the infrastructure subject matter expert in the development of IaC solutions; own the execution, validation, and operational adoption of approved automated solutions within the Compute & Storage environment.

  • Own and govern infrastructure performance KPIs - availability, capacity utilization, incident response SLAs, and patching compliance; proactively surface trends and recommend corrective actions

Strategy, Cross-Team Coordination, Projects & Documentation:

  • Serve as the primary technical liaison to Database, Enterprise Platforms, Automation & Observability, Networking, and DC Operations; resolve cross-team disagreements with peer team leads; escalate unresolved issues per the established 48-hour escalation model

  • Translate architectural designs from Enterprise Architecture into operational infrastructure standards and procedures; provide input into overall infrastructure roadmap planning and technology lifecycle decisions

  • Lead technical delivery of infrastructure projects from inception to completion with measurable milestones and managed scope; engage and manage third-party vendors on complex infrastructure issues, support renewals, and service escalations

  • Own the documentation standard for the Compute & Storage team; drive creation and maintenance of the server build runbook, capacity and resource catalog, infrastructure configuration standards, provisioning handoff checklist, backup coverage map, and recovery runbooks

Decision Scope & Accountability: Authorized to make independent infrastructure configuration and architectural decisions within defined scope and standards. Exercises discretion on decisions with broader impact - consults the Manager, IT - Compute & Storage before committing to changes affecting cross-team dependencies, security posture, Azure spend, or architectural direction.

Budgetary Awareness: Weighs cost into all infrastructure recommendations - hardware, Azure consumption, licensing, and vendor renewals. Surfaces cost considerations and optimization opportunities proactively to the Manager, IT - Compute & Storage. Does not independently commit spend.

Mentoring & Development: Provides significant mentoring and coaching to Infrastructure Operations Engineers. Actively contributes to cross-training with Database, Automation & Observability, Enterprise Platforms, and DC Operations teams.

Skills & Qualifications

  • 6+ years in infrastructure engineering with hands-on ownership across compute, storage, and cloud in a production enterprise environment; prior senior or lead individual contributor experience preferred

  • Expert-level administration of VMware vSphere and Nutanix in enterprise production environments; experience with cluster design, capacity planning, and lifecycle management

  • Deep experience with block and file storage technologies - RAID, SAN (Fibre Channel or iSCSI), NAS protocols, and enterprise storage array administration; Hitachi or equivalent required

  • Expert-level administration of Veeam or equivalent enterprise backup platform - architecture, job design, recovery validation, and RPO/RTO management

  • Strong working knowledge of Azure IaaS architecture - virtual machines, managed disks, virtual networking, and cloud cost governance; hybrid cloud experience required

  • Expert-level Windows Server administration - Active Directory, DNS, DHCP, Group Policy, clustering, and server hardening

  • Strong Linux (Ubuntu) administration - installation, configuration, patching, hardening, and troubleshooting in a production environment

  • Advanced PowerShell and/or Bash scripting

  • Working knowledge of IaC concepts and tooling (Ansible, Terraform, or equivalent) sufficient to define infrastructure requirements, review IaC solutions developed by the Automation & Observability team, and execute and validate automated workflows in production

  • Working familiarity with security and compliance frameworks (ISO 27001, CIS Controls) as applied to infrastructure configuration and access management; CompTIA Security+ or equivalent understanding expected

  • Familiarity with container technologies (Docker, Kubernetes) and their infrastructure dependencies in an enterprise environment

  • Proficient in ITSM processes and ITIL-based Incident, Problem, Change, and Capacity Management

  • Experience delivering and owning technical projects end-to-end, including vendor management and stakeholder communication

  • Effective communication with peers, management, vendors, and internal customers at all levels; ability to convey complex technical concepts to non-technical stakeholders

  • Experience supporting 24x7 global production environments; on-call availability required

  • Education: Bachelor’s degree in computer science, Information Technology, or equivalent experience

Preferred: ITIL V3/4; VMware VCP, Nutanix NCP, or Azure Administrator/Solutions Architect (AZ-104/AZ-305); Microsoft Certified: Azure Administrator Associate; Veeam certification; CompTIA Security+ or CISSP; prior experience in eDiscovery, legal technology, or a similarly regulated environment

Driving Career Growth, Benefit Excellence: The KLD Advantage

At KLD we invest in employees and their families by placing their wellbeing first. We offer competitive total compensation that includes base pay, bonus potential, inclusive benefits, wellness programs, and perks. We use market and industry data to inform pay decisions while considering geography and labor markets, individual experience, and business needs. India compensation is based upon the local competitive market.

  • Paid time off, that offers various time off options to help employees maintain a work-life balance, such as Casual, Earned, Sick, Special Leave, and Holidays

  • Ongoing learning and development, a focus on continuous professional development through various training and education reimbursement programs

  • A diverse and inclusive workplace where we all learn, grow, and achieve the greatest heights…together

  • A surrounding team of mission-driven individuals who genuinely love what they do

  • Free, fun, interactive and incentivized global wellness program that promotes the wellbeing of our employees

  • India compensation is based upon the local competitive market

Who We Are

KLDiscovery provides technology-enabled services and software to help law firms, corporations, and government agencies solve complex data challenges. With offices in 26 locations across 17 countries, KLDiscovery is a global leader in delivering best-in-class data management, information governance, and eDiscovery solutions to support the litigation, regulatory compliance, and internal investigation needs of clients. Our Nebula Ecosystem provides powerful end-to-end eDiscovery and enterprise-grade information governance. Through its global Ontrack data recovery business, KLDiscovery delivers world-class data recovery, disaster recovery, email extraction and restoration, data destruction, and tape management.

We Provide Equal Employment Opportunity.

At KLDiscovery we believe that inclusion and diversity make us stronger. We are committed to fostering an inclusive environment for all employees that enhances wellbeing and belonging. We welcome and celebrate individuals of all backgrounds, experiences, and perspectives.

We do not discriminate on the basis of race, color, religion, gender, pregnancy, gender identity, sexual orientation, national origin, age, disability, genetic information, veteran status, or any other protected status. We are happy to support you with any accommodation request at any stage in our hiring process

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
368,941 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
In your city
$173k – $314k per year • In office • Full-Time • 12+ years exp • Bachelor's Degree • San Francisco
Apex
JavaScript
Node JS
Python
SQL
TypeScript
Apex
Lightning Web Components
AI/ML
Agentforce
AI Agents
Claude
Claude Code
Copilot
Cursor
LLM
RAG
DevOps
AWS
Azure
CI/CD
Docker
GCP
GitHub
Grafana
gRPC
Kubernetes
New Relic
Prometheus
Splunk
Marketing
Salesforce
QA
Cypress
JMeter
k6
Locust
Playwright
Postman
Rest-Assured
Selenium
Apply
$29k – $73k per year (Estimated) • Remote • Full-Time • 5+ years exp • Bachelor's Degree • Guadalajara
Python
DevOps
Amazon CloudWatch
Azure
CI/CD
Datadog
Docker
Kubernetes
Prometheus
Terraform
Apply
$38k – $91k per year (Estimated) • Remote • Full-Time • 8+ years exp • PhD • Guadalajara
PowerShell
Python
Databases
Amazon Aurora
DynamoDB
AI/ML
Amazon SageMaker
AWS Bedrock
AWS Bedrock AgentCore
Ray
DevOps
Amazon CloudWatch
Amazon EC2
Amazon ECS
Amazon EKS
Amazon EventBridge
Amazon S3
AWS
AWS Lambda
AWS Step Functions
Azure
CI/CD
Datadog
FinOps
GCP
Git
GitLab
GitLab CI
IAM
Jenkins
JFrog Artifactory
Kubernetes
New Relic
Service Mesh
Splunk
Terraform
Cybersecurity
HIPAA
ISO 27001
PCI DSS
SOC 2
Apply
$129k – $231k per year (Estimated) • In office • Full-Time • 8+ years exp • Bachelor's Degree • Atlanta
SQL
AI/ML
AI Agents
Claude
Claude Code
DevOps
Azure
Design
Figma
Apply
Senior ML Engineer 4 hours ago
$149k – $224k per year • In office • Full-Time • 5+ years exp • Master's Degree • San Francisco • Washington • Palo Alto
Python
Python
pySpark
Databases
Apache Kafka
AI/ML
AI Agents
Agentforce
Airflow
Anomaly Detection
Feature Store
Flink
Ray
Red Teaming
Spark
DevOps
CI/CD
Docker
Kubernetes
Cybersecurity
MITRE ATT&CK
Marketing
Salesforce
Apply
$144k – $311k per year (Estimated) • Remote • 15+ years exp • Bachelor's Degree
AI/ML
ISO 42001
NIST AI RMF
DevOps
AWS
Azure
CI/CD
GCP
SLI/SLO/SLA
IAM
Cybersecurity
Crowdstrike
ISO 27001
SOC 2
Threat Modeling
Apply
Software Engineer II 22 days ago
Remote/Hybrid • 3+ years exp
C#
SQL
TypeScript
JavaScript
C#
ASP.NET Core
Blazor
Entity Framework Core
Databases
MS SQL
AI/ML
Claude
Claude Code
Frontend
Angular
DevOps
Azure
Azure DevOps
GitHub
Cybersecurity
SonarQube
Apply
$74k – $155k per year (Estimated) • Equity • Remote • Bachelor's Degree
Management
Microsoft Project
Apply
$140k – $160k per year • In office
C#
SQL
TypeScript
JavaScript
C#
.NET
Databases
PostgreSQL
Frontend
Angular
DevOps
Azure
Azure DevOps
GitHub
Apply
$190k – $225k per year • In office • 10+ years exp • Bachelor's Degree
DevOps
CI/CD
Cybersecurity
ISO 27001
SOC 2
Apply
See all jobs
This is one of many
368,941 more open roles from verified company boards, updated every day.