612,151open jobs
29,363companies
85,745added this week
Browse all
Salary
$210k – $240k per year
Location
In office (San Francisco)
Seniority
Senior · 8+ years exp
Employment
Full-Time
Overview
Company
Impact
Profile match

About Us

Alembic is the pioneering Causal AI platform. We help the world's largest enterprises move past correlation to prove what actually drives business outcomes - the question marketing and growth teams have never been able to answer with confidence. Fortune 100 companies including Nvidia, Delta Air Lines, and Mars use Alembic to make multimillion-dollar decisions on trusted, causal evidence.

We're backed by a $145M Series B from WndrCo (founded by Jeffrey Katzenberg), Jensen Huang, Joe Montana, Prysm Capital, and Accenture. Our models run on our own NVIDIA DGX SuperPOD built on Grace Blackwell infrastructure - one of the fastest private supercomputers in the world. (We've melted GPUs getting here.)

About the Role

We're building infrastructure that has to perform under real-world scale, reliability, and security demands - and we're looking for an engineer who wants to own the foundation it runs on. This isn't a traditional "keep the lights on" role.

You'll design and operate the global network and reliability layer behind one of the world's fastest private supercomputers - the fabric powering distributed compute, ML workloads, real-time analytics, and mission-critical enterprise systems. You'll work across networking, systems, automation, observability, and reliability engineering to scale a platform where performance genuinely matters, with real influence over architecture decisions.

It's a strong fit if you like solving deep infrastructure problems, building resilient systems, automating everything repetitive, and owning architecture rather than just maintaining it.

What You'll Do

  • Architect and operate scalable, secure network architecture for high-security requirements and large-scale machine learning workloads.

  • Own network device configuration management end to end, ensuring consistency and reliability across the fleet.

  • Improve system and network reliability and performance through automation, observability, and proactive capacity planning.

  • Implement and manage complex network protocols and connectivity, including BGP, VPNs, and WAN circuits and external peering.

  • Build and maintain comprehensive monitoring, alerting, and incident response - SLOs, runbooks, and on-call rotations - and drive post-incident analysis and continuous improvement.

  • Ensure security, compliance, and operational readiness across our network and cloud infrastructure.

  • Partner across engineering and data science to drive a culture of performance and reliability.

What Will Help You Succeed

  • 8+ years in network or infrastructure engineering, including 5+ years in datacenter operations and/or systems and network administration.

  • A strong background in network security, architecture, design, and operations.

  • Extensive hands-on experience with network devices (firewalls, switches, load balancers) and large-scale architectures and protocols - BGP, QoS, MPLS, and IPsec VPNs.

  • Experience designing and operating modern datacenter network fabrics (spine-leaf, EVPN/VXLAN, ECMP).

  • Network automation and IaC tooling (Ansible, Terraform, Nornir, or similar), plus IPAM/DCIM platforms (NetBox, Infoblox, or similar).

  • WAN engineering - carrier circuit provisioning and external network peering.

  • Familiarity with Kubernetes networking (CNI plugins, ingress, service networking, network policy) and strong operational experience with Linux-based production infrastructure.

  • Experience with monitoring and observability stacks (Prometheus, Grafana, Datadog, ELK, OpenTelemetry).

  • Solid scripting (Python, Bash) to debug complex network and system issues and automate solutions, plus excellent cross-functional communication.

Also Helpful

  • NVIDIA networking technologies - Cumulus Linux, InfiniBand, Spectrum-X, and BlueField DPUs (this is the fabric behind our SuperPOD).

  • Familiarity with data-intensive platforms (Spark, Airflow, Kafka) and storage network protocols (NFS, LustreFS, iSCSI).

  • Security practices for applications and infrastructure, and experience in high-compliance or SOC 2 environments.

The Role Is Right for You If

  • You want to own mission-critical network and infrastructure end to end - from architecture to incident management - not just keep it running.

  • You'd rather build and automate than direct from a distance, and you want meaningful influence over how a high-performance platform scales.

Why You Might Be Excited About Alembic

  • Hard problems with real impact: You'll own the network and reliability layer behind systems that influence multimillion-dollar decisions at Fortune 100 companies.

  • Cutting-edge technology: Operate our own NVIDIA DGX SuperPOD on Grace Blackwell - one of the fastest private supercomputers in the world - and run a fabric (InfiniBand, Spectrum-X, BlueField) almost no company has in-house.

  • Technical autonomy: Ownership over architecture decisions and the freedom to solve hard infrastructure problems your way.

  • Elite team: Join top engineers who thrive on hard problems and high-impact work.

  • Series B momentum, real ownership: Meaningful equity at a Series B company that's raised $145M, with proven product-market fit and Fortune 100 traction.

Why You Might Not Be Excited

  • If you only want to tell people what to build instead of building and automating alongside them, this isn't the environment for you.

  • You prefer companies with 100% built-out process for every detail.

  • You prefer static over dynamic - projects and priorities adapt as we grow. We have real paying customers and a playbook, and we still move at startup speed at Series B scale.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
612,151 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
San Francisco
$18k – $43k per year (Estimated) • In office • Full-Time • 8+ years exp • Bachelor's Degree • Pune
Python
AI/ML
Copilot
AI Agents
OpenAI
DevOps
Rest API
Terraform
Ansible
Azure
AWS
AIOps
Management
ServiceNow
Apply
$19k – $45k per year (Estimated) • In office • Full-Time • 5+ years exp • Indore
Python
DevOps
Rest API
Terraform
Ansible
CI/CD
Git
Cybersecurity
Zscaler
Apply
$17k – $42k per year (Estimated) • Remote • Full-Time • Moscow
Python
Bash
Databases
OpenSearch
AI/ML
Model Context Protocol
vLLM
Ollama
LLM
DevOps
Terraform
Docker Compose
Helm
Prometheus
Yandex Cloud
GitLab CI
CI/CD
GitOps
ArgoCD
Docker
Kubernetes
Grafana
Harbor
GitLab
Cybersecurity
SBOM
Apply
$26k – $61k per year (Estimated) • Remote/Hybrid • 7+ years exp • Bachelor's Degree • Pune
Python
SQL
Scala
Databases
Snowflake
Apache Kafka
AI/ML
Cursor
Spark
Claude Code
OpenAI Codex
DevOps
GitHub Actions
CI/CD
Jenkins
AWS
Kubernetes
Amazon EKS
AWS Lambda
Amazon S3
IAM
Amazon CloudWatch
Amazon Kinesis
Apply
$68k – $196k per year (Estimated) • Remote/Hybrid • Full-Time • Sydney
SQL
C#
C#
.NET
Databases
Azure Cosmos DB
DevOps
Rest API
Terraform
Azure DevOps
Azure
CI/CD
Git
Bicep
Apply
$140k – $186k per year • In office • Full-Time • San Francisco
Python
Python
Alembic
Cybersecurity
Okta
Least Privilege
Management
Google Workspace
Apply
Personal Trainer 23 days ago
$95k – $115k per year • In office • Full-Time • San Francisco
Apply
Technical Recruiter 2 months ago
$125k – $160k per year • In office • Full-Time • San Francisco
C++
AI/ML
CUDA Toolkit
CUDA
Apply
Technical Sourcer 2 months ago
$110k – $135k per year • In office • Full-Time • PhD • San Francisco
C++
AI/ML
CUDA Toolkit
CUDA
Semantic Search
DevOps
GitHub
HPC
Marketing
LinkedIn
Apply
Patent Counsel 3 months ago
$175k – $200k per year • In office • Full-Time • 4+ years exp • San Francisco
Python
C++
Python
Alembic
AI/ML
ChatGPT
DevOps
HPC
Apply
$115k – $180k per year • Remote/Hybrid • Full-Time • 4+ years exp • San Francisco • Seattle • Raleigh • New York
Apply
$97k – $124k per year • In office • Full-Time • 1+ year exp • San Francisco
Design
Canva
Management
Slack
Google Workspace
Apply
$200k – $240k per year • Remote/Hybrid • San Francisco
AI/ML
AI Agents
LLM
RAG
Apply
$152k – $301k per year (Estimated) • In office • Full-Time • Bachelor's Degree • Dallas • Austin • San Francisco • Fort Worth • Los Angeles
Marketing
Salesforce
Apply
$285k – $335k per year • Equity • In office • Full-Time • 10+ years exp • San Francisco
DevOps
VMWare
containerd
Kubernetes
KVM
QEMU
Hyper-V
Apply
See all jobs
This is one of many
612,151 more open roles from verified company boards, updated every day.