Overview
Company
Profile match
Impact
Conditions
Benefits
Hiring process
Similar jobs

The Voleon Group

The Voleon Group is a premier quantitative investment management firm headquartered in Berkeley, California, that specializes in machine learning-driven trading strategies. Founded in 2007 by scientists Michael Kharitonov and Jon McAuliffe, the hedge fund combines advanced data science, statistical modeling, and computational AI to forecast financial markets and execute systematic algorithmic trades.

Voleon is a technology company that applies state-of-the-art AI and machine learning techniques to real-world problems in finance. For nearly two decades, we have led our industry and worked at the frontier of applying AI/ML to investment management. We have become a multibillion-dollar asset manager, and we have ambitious goals for the future.

Your colleagues will include internationally recognized experts in artificial intelligence and machine learning research as well as highly experienced finance and technology professionals. In addition to our enriching and collegial working environment, we offer highly competitive compensation and benefits packages, technology talks by our experts, a beautiful modern office, daily catered lunches, and more.

As a Senior Cluster Site Reliability Engineer (SRE), you will help scale our research compute cluster to meet our growing needs, and you will leverage engineering skills to ensure high degrees of uptime, reliability, and robustness. Our research clusters are at the core of our R&D, and you will be directly responsible for keeping this key resource available and performant. Your work will provide a world-class HPC platform for researchers to focus on cutting-edge machine learning problems at scale. You will support both on-prem and cloud infrastructure, and work to provide the best experience to our technical staff. You will leverage IaC, Automation, and SRE principles to refine and hone a product that operates 24/7 to support Voleon.

The Cluster Operations team works on the frontline to triage and mitigate real-time operational issues. You will be an integral member of this team, solving day-to-day issues with high urgency, while also engineering systemic improvements and architectural fixes to prevent recurring issues. You will collaborate with engineering teams to develop improvements to monitoring/telemetry. You will help design and oversee operational frameworks to ensure the cluster operates within a set of rigorous SLAs.

Responsibilities

  • Be a first responder in the event of cluster outages or issues. Triage and resolve urgent issues as they arise

  • Ensure a high degree of cluster uptime (measured in multiple nines), and define + track SLAs to quantify reliability

  • Diagnose systemic/recurring patterns of problems, and engineer precision solutions to them in collaboration with engineering teams

  • Develop robust metrics and observability for cluster health and use those metrics to inform your work. Build out custom observability mechanisms when off-the-shelf ones won't do

  • Help software and research teams design policies around fair cluster usage, and help develop enforcement mechanisms for said policies

  • Assist in forecasting cluster growth, and help select appropriate scale-up strategies. Help optimize operations across dimensions of cost and usability

Requirements

  • 5+ years of experience in SRE or DevOps roles, preferably working as a senior engineer or tech lead

  • Knowledge of HPC/batch compute frameworks (Slurm, Kueue, AWS/GCP Batch) and/or machine learning training systems (Kubeflow, MLflow, Horovod)

  • Ability to develop scripts and utilities of moderate complexity in a common scripting language (Python, Ruby, etc.)

  • Familiarity with infrastructure-as-code and configuration management tools (Terraform, Ansible)

  • Experience with cloud infrastructure (AWS or GCP)

  • Familiarity designing and implementing modern observability stacks (Prometheus, Grafana, Loki, ELK, OpenTelemetry)

  • Experience with distributed storage technologies (Lustre, Ceph, S3)

  • Embodies a "system engineer" rather than "system administrator" mindset, thinking systematically and leveraging automation

  • Bachelor degree in computer science

Preferred Qualifications

  • Hands-on experience with HPC frameworks (Slurm, Grid Engine) and Kubernetes-based job orchestrators (Airflow, Kueue, Kubeflow Pipelines), along with other distributed computing frameworks (Ray, Modin, Dask, Spark)

  • Familiarity with ML frameworks (PyTorch/Tensorflow, JAX, Horovod, DeepSpeed)

  • Familiarity with hybrid/on-prem environments

  • Experience with containerization (Docker, Podman, Singularity), particularly for HPC/batch compute environments

  • Experience with HPC networking (InfiniBand, RDMA)

  • Solid security/IAM foundations (Identity management systems, AWS/GCP IAM, Zero Trust)

“Friends of Voleon” Candidate Referral Program

If you have a great candidate in mind for this role and would like to have the potential to earn $15,000 if your referred candidate is successfully hired and employed by The Voleon Group, please use this form to submit your referral. For more details regarding eligibility, terms and conditions please make sure to review the Voleon Referral Bonus Program.

Equal Opportunity Employer

The Voleon Group is an Equal Opportunity employer. Applicants are considered without regard to race, color, religion, creed, national origin, age, sex, gender, marital status, sexual orientation and identity, genetic information, veteran status, citizenship, or any other factors prohibited by local, state, or federal law.

Recommended for you based on this role

Similar stack
Same company
In your city
$225k – $255k per year • Remote • Full-Time • 5+ year exp • Berkeley
C++
Go
Java
Python
DevOps
Bazel
CI/CD
Docker
Apply
$230k – $295k per year • Remote • Full-Time • 3+ year exp • Bachelor's Degree
Python
SQL
Databases
Presto
AI/ML
Spark
DevOps
AWS
SLURM
Apply
$200k – $255k per year • Remote • Full-Time • 6+ year exp • Bachelor's Degree
Python
SQL
Databases
Presto
AI/ML
Spark
DevOps
AWS
SLURM
Apply
$235k – $300k per year • Remote • Full-Time • 3+ year exp • Bachelor's Degree
C++
Go
Java
Python
Databases
DynamoDB
Google BigQuery
Snowflake
Trino
DevOps
Docker
Kubernetes
SLURM
Apply
$315k – $405k per year • Remote • Full-Time • 8+ year exp • Berkeley
Python
Databases
Trino
AI/ML
Dagster
Flink
Spark
DevOps
Platform Engineering
Apply
$225k – $255k per year • Remote • Full-Time • 5+ year exp • Berkeley
Python
Databases
Trino
AI/ML
Dagster
Flink
Spark
DevOps
Platform Engineering
Apply
$160k – $200k per year • Remote • Full-Time • 3+ year exp • Bachelor's Degree • Berkeley
Python
SQL
Apply
$315k – $405k per year • Remote • Full-Time • 10+ year exp • Bachelor's Degree
Go
Python
Databases
Cassandra
DuckDB
MySQL
PostgreSQL
AI/ML
Feast
Ray
Spark
DevOps
Grafana
gRPC
Prometheus
Apply
$290k – $395k per year • Remote • Full-Time • 5+ year exp • Bachelor's Degree
C++
Python
C++
PyTorch C++
TensorFlow C++
AI/ML
NumPy
Pandas
PyTorch
Scikit-learn
SciPy
TensorFlow
Apply
$225k – $310k per year • Remote • Full-Time • 5+ year exp • Bachelor's Degree • New York
C++
Go
Python
C++
PyTorch C++
Databases
Cassandra
DuckDB
DynamoDB
MySQL
PostgreSQL
SQLite
AI/ML
PyTorch
DevOps
gRPC
Apply
Career impact
Discover how this job can transform your career
Get a personal career forecast for this job - salary uplift, next-level role, skill boost and a 3-year financial impact, all calculated from your profile.
Personal salary uplift vs. your current pay
Your 3-year career trajectory
Skills you will level up in this role
3-year financial impact in dollars
Create free account
Free forever • Less than a minute • No credit card

Work setup

Location
Berkeley
Remote work
Remote
Employment
Full-Time

Compensation

Salary
$205k – $235k per year