657,963open jobs
38,241companies
93,447added this week
Browse all
Salary
$200k – $240k per year
Location
Remote/Hybrid (Palo Alto, United States)
Seniority
Staff
Employment
Full-Time
Overview
Company
Impact
Profile match
Run full-resolution simulations in minutes. Vinci’s foundation model for physics unites AI acceleration with verified solvers for as-built accuracy.

About Vinci

Every physical thing you touch exists because somebody successfully navigated the laws of physics: the chips in your phone, the vehicles on the road, the data centers powering AI. Physics determines what can be built, how well it performs, and where it breaks. Yet the tools engineers use to understand physical behavior are too slow and too specialized to use continuously while designing, so critical decisions get made with only a partial view of how a system will behave.

Our mission is to make physical reasoning as accessible to engineers as language became through modern AI. This is not an attempt to build slightly better engineering software. It is an attempt to change how physical products are designed. Our technology is used today by many of the world's most advanced semiconductor and electronics organizations, including nearly half of the twenty largest companies in the industry. We are backed by Khosla Ventures and Eclipse Ventures.

The role

Training a foundation model for physics means holding petabytes of simulation data and keeping GPU clusters saturated with it. We are hiring the engineer who will own that layer: the storage where our training data lives, and the compute the AI team trains and validates models on.

This is a founding role for the discipline, with a wide surface and a small team. You will set the compute and data architecture, and you will also be the person who finds out why a job has been sitting in pending for two hours. The AI team decides what to train and how to judge the result. Your job is to make sure the compute and the data are there when they need them, that the runs finish, and that we can afford it. The measure of the role is what the AI team can do with what you build: how many experiments they can run, how quickly results come back, how reliably a long run finishes without an infrastructure failure, and how much training we get for what we spend.

Nothing in this layer stays fixed. Our training data will grow, and it will take on new sources and new types alongside what we have today. GPU hardware and the market for it move faster than most infrastructure, and our appetite for compute keeps increasing. We are hiring the person who leads us through that on the compute and data side: who identifies the next step before we are forced into it, makes the case for what it costs, and then builds it.

What you'll do

  • Compute architecture. Define how our GPU capacity is organized, provisioned, and grown, and how that architecture holds up as both the fleet and the team get larger.

  • Workload scheduling. Own how competing work claims that capacity: queueing, priority, preemption, gang scheduling, and quota between training runs, validation jobs, and data preparation. This starts as a judgment call among a handful of stakeholders and becomes an allocation problem worth solving algorithmically. Recognizing when that transition arrives, and building for it, is part of the role.

  • Data infrastructure at petabyte scale. Design the storage, ingestion, and access paths that keep training jobs fed at full throughput, and keep datasets versioned and reproducible as they change. Build for new sources and data types arriving at similar or larger scale, not only for what we hold today.

  • Training and validation platform. Own the systems the AI team uses to launch, checkpoint, resume, and monitor runs, and make sure validation workloads get the compute and data access they need without competing with training for it.

  • Reliability, utilization, and cost. Set and meet targets for cluster utilization, job success rate, and time from submission to first batch. Own the cost of a training run and be able to account for where it goes.

  • Hands-on engineering. Write and review production code, lead architecture reviews, and take the hardest debugging problems yourself.

  • Mentorship and teaching. Mentor the engineers who join this team as it grows, and help hire them. Review designs outside your own work. Make the AI team more capable with the infrastructure than they were before, and write things down so the answer to a recurring question lives somewhere other than in your head.

  • Technical direction. Work with the AI team on what the next model will require from infrastructure before it is required, and with leadership on capacity planning and compute investment. Say what our compute and data footprint needs to look like a year out and what it will cost. The leadership in this role is technical, not managerial.

What we're looking for

  • 10+ years building large-scale distributed systems, including several years running GPU infrastructure for large model training

  • Direct experience operating multi-node distributed training on modern accelerators: runs that held dozens or hundreds of GPUs for days at a time, with the scheduling, interconnect behavior, checkpointing, and failure recovery that requires. You have found out why a run was slower than the hardware allowed and fixed it.

  • Direct experience serving training data at petabyte scale, where throughput and storage cost were both constraints you had to answer for

  • Experience setting scheduling and quota policy on shared GPU capacity, with a view on the tradeoff between fleet utilization and how long people wait in the queue

  • A track record of building infrastructure whose users are researchers and engineers, and of being measured on what those users were able to do with it, not on the system itself

  • Infrastructure you built that survived a substantial change in scale or in the character of the workload, with a clear account of what held up and what you had to replace

  • A history of mentoring engineers and of teaching people outside your specialty enough to work on their own

  • The ability to set technical direction, make the case for it to people outside infrastructure, and then implement it. At this stage the role is one engineer and a small team, not one engineer directing several.

  • Advanced degree in computer science or a related field, or equivalent industry experience

Nice to have

  • Kubernetes, Ray, Slurm, Airflow, or comparable orchestration running ML workloads in production

  • Hybrid or multi-cloud GPU capacity, including capacity planning and cost negotiation with providers

  • Fluency in a modern training stack: PyTorch distributed, FSDP or DeepSpeed, mixed precision, GPU profiling

  • Applied optimization for scheduling or allocation problems, such as linear programming, graph algorithms, or heuristic solvers

  • Experiment tracking and model artifact management

  • Work with large numerical or geometric datasets, such as simulation, EDA, CAD, imaging, or autonomous systems

Is this you?

This role writes and owns the code. If your recent work has been directing an infrastructure organization, setting standards for other teams, or managing vendor relationships, this is probably not the right fit. If you have been the person who owns a cluster and everything that runs on it, it likely is.

Much of what a platform team would have handled for you at a larger company does not exist here yet. Building it is the job.

We are also hiring for the scale you have already worked at. This role assumes production training runs on large multi-GPU clusters, and petabyte-scale data serving under a budget. Single-node or lab-scale model work does not transfer to these problems.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
657,963 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
Palo Alto
$176k – $264k per year • Equity • In office • Full-Time • 8+ years exp • Pleasanton
Python
TypeScript
Python
Flask
FastAPI
Django
Databases
PostgreSQL
ElasticSearch
AI/ML
AI Agents
DevOps
CI/CD
AWS
Docker
Kubernetes
Management
Agile
Apply
$13k – $31k per year (Estimated) • In office • Full-Time • 1+ year exp • Gurgaon
Python
Java
C#
C#
.NET
Databases
Weaviate
Pinecone
AI/ML
LangGraph
LangChain
Claude
Model Context Protocol
Prompt Engineering
Function Calling
Chain-of-Thought
AI Agents
Semantic Kernel
CrewAI
Gemini
LLM
RAG
Semantic Search
GPT-4
Structured Outputs
Semantic Search
LLM Guardrails
Agentic Workflows
Tool Use
Frontend
GraphQL
DevOps
Rest API
GCP
Azure
CI/CD
GitOps
Git
AWS
Docker
Kubernetes
Vector
Management
ServiceNow
Apply
$16k – $37k per year (Estimated) • In office • Full-Time • 1+ year exp • Pune
Python
Java
C#
C#
.NET
Databases
Weaviate
Pinecone
AI/ML
LangGraph
LangChain
Claude
Model Context Protocol
Prompt Engineering
Function Calling
Chain-of-Thought
AI Agents
Semantic Kernel
CrewAI
Gemini
LLM
RAG
Semantic Search
GPT-4
Structured Outputs
Semantic Search
LLM Guardrails
Agentic Workflows
Tool Use
Frontend
GraphQL
DevOps
Rest API
GCP
Azure
CI/CD
GitOps
Git
AWS
Docker
Kubernetes
Vector
Management
ServiceNow
Apply
$61k – $147k per year (Estimated) • Remote/Hybrid • Full-Time • 3+ years exp • United Kingdom
Python
Java
Databases
Snowflake
Databricks
Apache Kafka
Google BigQuery
BigQuery
AI/ML
Spark
Flink
DevOps
Terraform
GCP
Azure DevOps
GitHub Actions
CloudFormation
Azure
CI/CD
Jenkins
Git
AWS
Docker
Kubernetes
Amazon EKS
AWS Lambda
Amazon Kinesis
Analytics
ETL/ELT
Management
Agile
Apply
$152k – $279k per year (Estimated) • Remote/Hybrid • Full-Time • 12+ years exp • Associate's Degree • Chicago • Milwaukee • Dallas • Columbus • Kirkland
Python
JavaScript
Node JS
AI/ML
LLM
Frontend
React.js
DevOps
Terraform
AWS CDK
Azure
AWS
Docker
Kubernetes
Platform Engineering
Amazon EKS
AWS Fargate
AWS Lambda
Amazon EC2
HPC
Apply
$138k – $150k per year • Remote • Full-Time • 15+ years exp • PhD • Paris
Python
C++
AI/ML
Edge AI
Apply
$110k – $140k per year • Remote • Full-Time • 15+ years exp • PhD • Paris
Python
C++
AI/ML
Edge AI
Apply
$77k – $177k per year (Estimated) • In office • Full-Time • 15+ years exp • Brussels
Python
C++
AI/ML
Edge AI
Apply
$180k – $220k per year • Remote/Hybrid • Full-Time • 5+ years exp • Palo Alto
Python
JavaScript
TypeScript
C++
Cython
C++
Protobuf
STL
Cython
PyBind11
Databases
PostgreSQL
AI/ML
CUDA Toolkit
CUDA
ROCm
Frontend
React.js
DevOps
gRPC
OpenTelemetry
Prometheus
SLURM
CI/CD
Kubernetes
HPC
Apply
$180k – $220k per year • Remote/Hybrid • Full-Time • 6+ years exp • Palo Alto
Python
C++
AI/ML
CUDA Toolkit
VLM
LLM
CUDA
DevOps
gRPC
Docker
Kubernetes
HPC
Apply
$73k – $150k per year (Estimated) • In office • Full-Time • 2+ years exp • PhD • Palo Alto
Apply
$184k – $371k per year (Estimated) • Equity • In office • Full-Time • 10+ years exp • Bachelor's Degree • Palo Alto
Python
SQL
Databases
Snowflake
Databricks
Apache Kafka
AI/ML
LangGraph
LangChain
Fine-tuning
AI Agents
CrewAI
LLM
RAG
Multi-Agent Systems
DevOps
GCP
Azure
AWS
Platform Engineering
Vector
Cybersecurity
PCI DSS
GDPR
HIPAA
Analytics
Tableau
ETL/ELT
Looker
Dimensional Modeling
Management
Outlook
Apply
People Ops Generalist 3 hours ago
$80k – $100k per year • Equity • Remote • Full-Time • 2+ years exp • Bachelor's Degree • Chicago • Boston • Austin • New York • Palo Alto
Apply
Director of Data 3 hours ago
$230k – $345k per year • Remote/Hybrid • Full-Time • 15+ years exp • Palo Alto
Databases
Snowflake
Analytics
Fivetran
Looker
Apply
$46k per year • In office • Full-Time • 1+ year exp • Palo Alto
Management
Google Docs
Apply
See all jobs
This is one of many
657,963 more open roles from verified company boards, updated every day.