1,152,325open jobs
65,679companies
205,349added this week
Browse all
Salary
≈ $159k – $347k per year (Estimated)
Location
In office (San Francisco)
Employment
Full-Time

Confirmed on the employer's own hiring board on Oct 3, 2026. First seen by Alion on Jun 15, 2026. Spellbrush scores D on the Alion truth index.

Overview
Company
Impact
Profile match
Spellbrush builds generative models for anime art and the games made with them, including the illustration tool used by its community. Its research focuses on stylised character generation rather than photorealism. The studio ships its own titles alongside the underlying model work.
Backed by Y Combinator

We’re looking for an experienced HPC infrastructure engineer to lead bringup, administration, and operations on is probably the largest anime AI training cluster in the world. You’ll serve as the bridge between our researchers and the bare GPU machines, helping to make sure that SLURM jobs are running, parallel filesystems are serving, network is transmitting, and that the anime models are training.

You may be a good fit if:

You love anime and the anime aesthetic.

This probably one of the only jobs in the world where you will get to combine your love of anime and large-scale GPU systems.

You’re familiar with the modern HPC software landscape

Once upon a time, our team could install SLURM on a few bare metal nodes and get away with it. Now the landscape has become unbelievable complex, with SLURM deploys through Slinky on K8s, provisioning through warewulf/MAAS/ansible, filesystems through WEKA/VAST/Ceph, VPN and access through tailscale, and monitoring via the Grafana/Prometheus stack. We’re looking for someone with relevant experience up and down the stack (and maybe a papercut or two to show for it!)

As well as the traditional sysadmin landscape

Bringing up and managing cluster still requires good old linux sysadmin skills, including wrangling ldap, triaging dmesg, and setting sticky bits on directories for misbehaving users and tools.

You're not afraid of physical computers

We’re building out edge datacenters and our CEO is still personally racking, stacking, and provisioning HGX-based nodes in our living room. Also his VLAN design sucks and he’s bad at fiber routing. Please send help.

And you're comfortable working on small, fast-paced teams.

We currently have a very tiny research team, and you’ll be directly helping some of the AI researchers in the world train the best anime image model in the world.

We also believe in the unmatched speed of in-person teams, and prefer on-site collaboration in either our primary research office in Tokyo (downtown Akihabara), or San Francisco (dogpatch!). Bay area is strongly preferred as we have physical hardware in the Bay Area. Visa sponsorships are available.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
1,152,325 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

AI/ML
Similar stack
Same company
San Francisco
≈ $27k – $81k per year (Estimated) • In office • Beijing
Python
Java
C++
C++
TensorFlow C++
PyTorch C++
AI/ML
NLP
TensorFlow
PyTorch
LLM
Apply
≈ $26k – $78k per year (Estimated) • In office • Hangzhou
Apply
In office • 2+ years exp • Beijing
AI/ML
LangChain
LlamaIndex
Model Context Protocol
Function Calling
RAG
Google AI Studio
Tool Use
Apply
≈ $33k – $99k per year (Estimated) • In office • Shanghai
Apply
In office • Hangzhou
Python
Java
Apply
In office
Python
SQL
Bash
DevOps
Terraform
Ansible
OpenShift
Helm
Loki
Prometheus
ArgoCD
Jenkins
Git
Kubernetes
Grafana
KVM
Incident Management
Linux
Management
Jira
Agile
Scrum
Apply
$120k – $185k per year • Equity • In office • TS/SCI • Arvada
Python
Go
Java
Rust
C++
DevOps
GCP
Loki
FluxCD
Prometheus
CI/CD
GitOps
ArgoCD
Kubernetes
Grafana
FinOps
Apply
≈ $73k – $143k per year (Estimated) • In office • Full-Time • 7+ years exp • Bachelor's Degree • United States
Python
Ruby
Perl
DevOps
Terraform
Puppet
Ansible
GCP
GitLab CI
Azure
CI/CD
Jenkins
Git
AWS
Kubernetes
SaltStack
Platform Engineering
Shift-Left
OpenStack
Cybersecurity
Shift-Left Security
Apply
≈ $34k – $102k per year (Estimated) • Hybrid • Full-Time • Bachelor's Degree • Islamabad
Python
SQL
Python
FastAPI
AI/ML
LangGraph
LangChain
LoRA
Fine-tuning
Embeddings
NLP
NER
PEFT
Transformers
TensorFlow
Keras
PyTorch
LLM
RAG
Tokenization
Agentic Workflows
DevOps
GCP
Azure
CI/CD
Git
AWS
Docker
Kubernetes
Apply
≈ $93k – $177k per year (Estimated) • Remote (United Kingdom) • Full-Time • 1+ year exp • United Kingdom
Python
JavaScript
TypeScript
Node JS
Databases
PostgreSQL
Redis
ClickHouse
Apache Iceberg
OpenSearch
Amazon Redshift
AI/ML
Spark
DevOps
Terraform
GitHub Actions
OpenTelemetry
CircleCI
CI/CD
AWS
Kubernetes
Amazon EKS
GitHub
Cybersecurity
ISO 27001
Analytics
AWS Glue
Apply
≈ $174k – $379k per year (Estimated) • In office • Full-Time • San Francisco
AI/ML
RLHF
LLM
Post-training
Google AI Studio
Apply
≈ $164k – $357k per year (Estimated) • In office • Full-Time • San Francisco
AI/ML
Grok
LLM
Google AI Studio
Apply
AI Anime Researcher 2 years ago
In office • Full-Time • Tokyo • San Francisco
AI/ML
JAX
Diffusion Models
PyTorch
TPU
Apply
≈ $103k – $245k per year (Estimated) • Hybrid • Full-Time • San Francisco
AI/ML
LLM
Google AI Studio
Machine Learning
Game Dev
Unity
Apply
Level Designer 1 year ago
≈ $96k – $227k per year (Estimated) • In office • Full-Time • San Francisco
C#
Game Dev
Unity
Apply
$120k – $150k per year • Equity 1–3% • In office • Full-Time • 1+ year exp • San Francisco
Python
JavaScript
Node JS
AI/ML
Edge AI
DevOps
GCP
Azure
AWS
Apply
≈ $189k – $390k per year (Estimated) • In office • Full-Time • San Francisco
Python
JavaScript
AI/ML
Fine-tuning
AI Agents
LLM
RAG
GPT-5
Frontend
Next.js
React.js
DevOps
WebRTC
WebSockets
Apply
$100k – $150k per year • Equity 0.2–0.5% • In office • Full-Time • 1+ year exp • San Francisco
Marketing
YouTube
Reddit
Apply
$75k – $100k per year • Equity 0–0.1% • Remote (United States) • Full-Time • 1+ year exp • San Francisco
Apply
$80k – $100k per year • Equity 0.1–0.2% • Remote (United States) • Full-Time • 1+ year exp • San Francisco
Apply
See all jobs
This is one of many
1,152,325 more open roles from verified company boards, updated every day.