1,358,174open jobs
79,249companies
207,309added this week
Browse all
Salary
≈ $132k – $286k per year (Estimated)
Location
Hybrid (San Francisco, United States)
Employment
Full-Time

Confirmed on the employer's own hiring board on Oct 8, 2026. First seen by Alion on Oct 7, 2026.

Overview
Company
Impact
Profile match
Prometheus is a niche investment boutique with a global perspective, serving from the Dubai International Financial Center (DIFC) and designing innovative and high performing tailored products. We take care of our clients and create customized financial solutions that suit their needs.

The Role

We’re hiring a Pre-Training Data & Acquisition Engineer to build the data systems powering Prometheus’s foundation models for the physical world. You’ll work closely with research and infrastructure teams to acquire, process, and deliver large-scale training datasets across engineering, scientific, and multimodal domains. This role spans distributed crawling, source integration, data processing, and production operation, with end-to-end ownership from raw content to training-ready datasets.

What You’ll Do

  • Identify and integrate valuable data sources across engineering and scientific domains.

  • Build distributed crawlers, API integrations, and bulk ingestion systems with effective scheduling, rate limiting, retries, and incremental updates.

  • Develop pipelines for parsing, extraction, normalization, deduplication, quality filtering, and tokenization across heterogeneous formats.

  • Optimize throughput and cost across networking, compute, storage, and databases as acquisition and processing workloads scale.

  • Build monitoring and tooling to track source coverage, ingestion failures, processing throughput, and usable data yield.

  • Work closely with pre-training researchers to translate data requirements into reliable pipelines and deliver datasets ready for large-scale training.

  • Own dataset reproducibility, versioning, provenance, and recovery from acquisition through delivery.

What We’re Looking For

  • Experience building and operating large-scale distributed systems, web crawlers, or data processing pipelines.

  • Strong programming ability in Python and Rust, Go, C++, or a comparable systems language.

  • A practical understanding of web infrastructure, including HTTP, DNS, concurrency, caching, and common failure modes.

  • Hands-on experience with databases, object storage, and distributed batch or streaming processing.

  • Experience designing fault-tolerant systems that handle partial failures, resume interrupted work, and prevent unintended duplication or data loss.

  • Ability to profile and debug performance across CPU, memory, disk, and network usage.

  • Strong technical judgment when integrating unfamiliar sources, APIs, and file formats.

  • Bias toward fast iteration and end-to-end ownership, from initial implementation through reliable production operation.

  • Experience with search indexing, document extraction, or foundation-model data pipelines is a plus.

Why Join Us

  • Work with world-class researchers on frontier AI systems for the physical world.

  • Build the acquisition and processing systems that supply engineering, scientific, and multimodal data to large-scale model training.

  • Competitive compensation and flexible work arrangements.

  • High-impact, mission-driven environment.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
1,358,174 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Data Science
Similar stack
Same company
San Francisco
≈ $81k – $155k per year (Estimated) • In office • Full-Time • 3+ years exp • Bachelor's Degree • Oklahoma City
Python
SQL
PowerShell
Databases
Snowflake
MS SQL
DevOps
Rest API
Prometheus
Azure
Cortex
IAM
Analytics
Tableau
Power BI
ETL/ELT
Matillion
Management
Agile
Apply
≈ $66k – $193k per year (Estimated) • In office • Internship • PhD • Columbus
Python
SQL
AI/ML
Scikit-learn
TensorFlow
Pandas
NumPy
Machine Learning
Analytics
Tableau
Power BI
Apply
$132k per year • Hybrid • Internship • 1+ year exp • San Mateo
Python
SQL
AI/ML
Spark
Machine Learning
Apply
≈ $99k – $193k per year (Estimated) • In office • 7+ years exp • Bachelor's Degree • Chino
Python
SQL
Databases
Snowflake
Apache Iceberg
Delta Lake
AI/ML
dbt
DevOps
Azure
Analytics
Power BI
Fivetran
Azure Data Factory
Master Data Management
Apply
Senior Data Engineer 2 hours ago
$160k – $185k per year • Equity • Hybrid • New York
Python
SQL
Databases
Google BigQuery
BigQuery
AI/ML
Airflow
dbt
DevOps
Terraform
GCP
CI/CD
Git
IAM
Cybersecurity
Least Privilege
Analytics
ETL/ELT
Superset
Apply
Senior Data Engineer 3 hours ago
$221k – $245k per year • In office • Full-Time • 4+ years exp • San Francisco
Python
SQL
Databases
Google BigQuery
BigQuery
AI/ML
dbt
Anomaly Detection
Machine Learning
Analytics
Tableau
Looker
Management
Discord
Apply
$115k – $130k per year • Hybrid • San Francisco
Design
AutoCAD
Management
Intercom
Microsoft Office
Apply
≈ $115k – $269k per year (Estimated) • In office • 8+ years exp • Bachelor's Degree • San Francisco
Marketing
GA4
YouTube
Apply
Senior FAS / Sales 2 hours ago
Remote (Europe, United States) • Full-Time • Master's Degree • San Francisco
Apply
≈ $51k – $84k per year (Estimated) • In office • Internship • Bachelor's Degree • San Francisco
Apply
See all jobs
This is one of many
1,358,174 more open roles from verified company boards, updated every day.