683,639open jobs
39,571companies
98,002added this week
Browse all
Salary
$116k – $254k per year (Estimated)
Location
In office (Berkeley)
Overview
Company
Impact
Profile match
The Essence of Innovation takes redefining what’s possible, to own the challenge and the solution. Solutions VAR/ITVAR The right technology will enable your agency to make smarter decisions, achieve mission-critical goals faster, and streamline operations.

Work Location: Onsite - California

Schedule: Full-Time | 5 Days Per Week | Midnight-8:00 AM (Owl Shift)

This position does not offer sponsorship. Must be authorized to work in the United States.

Position Overview

Essnova Solutions, Inc. is seeking an experienced Site Reliability Engineer (SRE) to support the National Energy Research Scientific Computing Center (NERSC), a mission-critical high-performance computing (HPC) and data environment supporting scientific research for the U.S. Department of Energy (DOE) Office of Science.

The Site Reliability Engineer will work as part of a 24/7 operations environment responsible for maintaining the accessibility, reliability, security, and operational health of large-scale computing and data systems.

This is a highly hands-on position combining Linux systems administration, infrastructure monitoring, incident response, programming and scripting, automation, networking, ServiceNow, and physical data center operations.

IMPORTANT SCHEDULE REQUIREMENT: This position requires working onsite five days per week on the midnight-8:00 AM shift. Candidates must be willing and able to consistently work this overnight schedule.

Key Responsibilities

  • Monitor high-performance computing systems, storage infrastructure, networks, and other data center and facility-related systems.
  • Review and respond to infrastructure and system alerts, perform initial triage, and engage appropriate on-call personnel when escalation is required.
  • Respond to alerts across multiple systems to help ensure monitoring and data collection remain operational 24/7.
  • Troubleshoot system, application, network, monitoring, and infrastructure issues affecting system reliability.
  • Develop solutions that improve operational processes, prevent recurring issues, and automate responses to routine service conditions.
  • Identify opportunities to improve monitoring capabilities, alerting, incident triage, and operational automation.
  • Develop and maintain tools within the monitoring pipeline in collaboration with operations personnel.
  • Develop software and integrations capable of generating alerts and notifications from HPC system APIs into monitoring pipelines.
  • Build and maintain application and tool configurations to ensure reliable operation as data volumes and user demands increase.
  • Utilize ServiceNow to support incident management, trouble-ticketing, operational workflows, and service management activities.
  • Collaborate across technical teams to identify and resolve operational bottlenecks and maintain system reliability.
  • Coordinate with technical groups during center-wide maintenance activities.
  • Manage diagnostic, monitoring, and notification software during planned maintenance periods.
  • Perform regular physical and logical walkthroughs of the data center floor.
  • Monitor environmental conditions, power distribution units (PDUs), cooling infrastructure, and other facility systems supporting reliable data center operations.
  • Maintain accurate trouble-ticket documentation for outages, incidents, maintenance activities, troubleshooting actions, and operational updates.
  • Analyze problems of varying complexity and evaluate technical data to determine appropriate troubleshooting and remediation methods.
  • Exercise independent technical judgment when selecting methods and approaches for resolving operational issues.

Compensation

$80.00 per hour

The anticipated pay rate for this position is $80.00 per hour. Actual compensation may be determined based on job-related factors including experience, qualifications, skills, contractual requirements, and applicable law.

Equal Employment Opportunity

Essnova Solutions, Inc. is an Equal Opportunity Employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, creed, sex, pregnancy, childbirth or related medical conditions, sexual orientation, gender, gender identity or expression, national origin, ancestry, age, physical or mental disability, medical condition, genetic information, marital status, military or veteran status, or any other characteristic protected by applicable federal, state, or local law.

Essnova Solutions, Inc. is committed to providing reasonable accommodations to qualified individuals with disabilities and applicants with disabilities throughout the recruitment and employment process.

Requirements

Required Qualifications

  • 5+ years of relevant professional experience in Site Reliability Engineering, systems/infrastructure engineering, DevOps, data center operations, HPC operations, network/system operations, or a closely related technical environment.
  • Strong hands-on experience working with Linux, including Linux shell and command-line environments such as SSH.
  • Programming and/or scripting experience using one or more languages such as:
    • Python
    • C
    • C++
    • Perl
    • Java
    • Comparable scripting or programming languages
  • Knowledge of standard software development practices.
  • Experience supporting large-scale IT infrastructure, highly available systems, data centers, critical installations, or comparable technical environments.
  • Knowledge of large data communications networks and common network protocols.
  • Network security experience, including knowledge of firewalls and access control lists (ACLs).
  • Experience troubleshooting infrastructure, application, system, network, or operational issues.
  • Experience responding to monitoring alerts and performing technical incident triage.
  • Ability to analyze operational and system data to identify problems and determine appropriate solutions.
  • Experience collaborating across multiple technical teams to resolve operational issues and maintain system reliability.
  • Strong written and verbal communication skills.
  • Ability to independently learn and apply new technologies in a complex technical environment.
  • Ability and willingness to work within a 24/7 operational environment.
  • Ability and willingness to work onsite five days per week from midnight-8:00 AM.

Education

  • Bachelor's degree in Computer Science, Information Technology, Engineering, or a related technical discipline preferred.
  • An equivalent combination of education, technical training, certifications, and relevant professional experience may be considered.

Technical Environment

Candidates may work with technologies and platforms including:

  • Linux / SSH
  • Python
  • C / C++
  • Perl
  • Java
  • ServiceNow
  • Kubernetes
  • Prometheus
  • VictoriaMetrics
  • Alertmanager
  • HPC systems
  • Monitoring and alerting pipelines
  • Network protocols
  • Firewalls and ACLs
  • Building management systems
  • Data center power and cooling infrastructure
  • Infrastructure and operational automation

Candidates are not necessarily expected to have prior experience with every technology listed above but should possess the technical foundation and learning ability necessary to work effectively within a complex computing and data center environment.

Preferred Qualifications

  • Experience implementing, configuring, or customizing ServiceNow.
  • Familiarity with IT Service Management (ITSM) best practices and service lifecycle management.
  • Experience supporting high-performance computing (HPC) environments.
  • Experience supporting scientific computing, large-scale data centers, critical infrastructure, or other highly available environments.
  • Hands-on experience with Kubernetes.
  • Experience with monitoring technologies such as Prometheus, VictoriaMetrics, Alertmanager, or comparable platforms.
  • Experience developing monitoring, alerting, infrastructure automation, or incident-response tools.
  • Experience developing integrations with system or infrastructure APIs.
  • Understanding of data center environmental monitoring, cooling systems, power utilization, and/or building management systems.
  • Practical experience developing or deploying Agentic AI or autonomous automation tools to streamline technical operations.
  • Experience building autonomous-agent solutions capable of automating technical decision-making, optimizing workflows, or enhancing proactive system monitoring.
Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
683,639 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
Berkeley
$38k – $73k per year (Estimated) • In office • Moscow
Python
SQL
Scala
Python
pySpark
Databases
PostgreSQL
Apache Kafka
AI/ML
Hadoop
Spark
AI Agents
NLP
LLM
RAG
DevOps
OpenShift
Prometheus
CI/CD
Jenkins
Docker
Kubernetes
Bitbucket
Apply
$20k – $52k per year (Estimated) • In office • Full-Time • 1+ year exp • Bachelor's Degree • Bengaluru
AI/ML
Copilot
ChatGPT
Claude Code
Model Context Protocol
Dagster
XGBoost
Triton Inference Server
Scikit-learn
Function Calling
AI Agents
Kubeflow
TensorFlow
PyTorch
LLM
TorchServe
Feature Store
ONNX Runtime
Agentic Workflows
Tool Use
DevOps
CI/CD
Git
Kubernetes
Apply
In office • Full-Time • 15+ years exp • Bachelor's Degree • Ahmedabad
DevOps
Incident Management
Apply
$69k – $157k per year (Estimated) • Equity • Remote • 5+ years exp
SQL
AI/ML
dbt
DevOps
GCP
Azure
AWS
Kubernetes
FinOps
Analytics
Looker
Apply
$36k – $88k per year (Estimated) • In office • Full-Time • Athens
Python
Java
SQL
Databases
Apache Kafka
AI/ML
Spark
Airflow
DevOps
Terraform
GCP
CI/CD
Git
Docker
Kubernetes
Analytics
ETL/ELT
Apply
$41k – $98k per year (Estimated) • In office • TS/SCI • Full-Time • Bachelor's Degree • Washington
Apply
In office • TS/SCI
Apply
$96k – $214k per year (Estimated) • Remote/Hybrid • Bethesda
DevOps
Azure
Cybersecurity
Microsoft Defender
Apply
$160k per year • In office • Berkeley
Apply
$97k – $224k per year (Estimated) • In office • Full-Time • Berkeley
Design
AutoCAD
Apply
$50k – $121k per year (Estimated) • In office • Full-Time • High School Diploma • Berkeley
Apply
$100k – $150k per year • Equity • In office • 3+ years exp • Bachelor's Degree • Berkeley
Analytics
Microsoft Excel
Apply
$80k – $110k per year • Equity • In office • 4+ years exp • Bachelor's Degree • Berkeley
Analytics
Microsoft Excel
Apply
$72k – $191k per year (Estimated) • In office • Full-Time • Bachelor's Degree • Berkeley
Apply
$39k – $96k per year (Estimated) • In office • Berkeley
C#
C#
.NET
Apply
See all jobs
This is one of many
683,639 more open roles from verified company boards, updated every day.