368,634open jobs
9,437companies
50,578added this week
Browse all
Salary
$150k – $230k per year
Location
In office (San Diego)
Seniority
Staff
Employment
Full-Time
Overview
Company
Impact
Profile match
Shield AI is an American defence technology company founded in 2015 that builds autonomy software and uncrewed aircraft for military missions. Its Hivemind autonomy stack lets aircraft fly, navigate and complete objectives without GPS, a remote pilot or a communications link, which matters in contested environments where those links are jammed. Headquartered in San Diego, California, the company fields the V-BAT vertical take-off aircraft with United States and allied forces, licenses Hivemind to other manufacturers and is developing the X-BAT jet for high-end missions.

Job Description:

We are looking for a Staff Data Platform Engineer to help define and build the data foundation of the AI Factory.

The Data Platform provides a unifying, knowledge-graph-centered API layer for human and agentic workflows. It connects configurations, requirements, software versions, test executions, files, signals, training data, and results through stable identities and typed relationships. It also provides consistent access to the storage and compute systems behind those data products.

This is a hands-on technical leadership role. You will design platform architecture, implement production software, evaluate storage and compute technologies, establish data-modeling patterns, and work directly with teams collecting and consuming mission-critical data. Success requires balancing developer productivity, semantic clarity, operational reliability, system performance, portability, and long-term maintainability.

What you'll do:

  • Develop a unifying Graph API: Lead the architecture and implementation of the knowledge graph and multi-modal API layer that serves as the backbone for human, service, and agentic workflows.
  • Own DataOps infrastructure: Research, optimize, and maintain the storage, indexing, query, ingestion, and compute infrastructure used throughout the data lifecycle.
  • Establish best-practices: Establish durable, best-practice patterns for schema modeling, relationships, lineage, and schema evolution.
  • Turbocharge agentic data access: Build APIs that enable agents to retrieve structured, connected, and explainable context rather than relying only on keyword or vector similarity.
  • Develop reference architectures: Establish recommended storage and compute profiles, deployment patterns, benchmarks, and operational guidance for both internal and customer-managed infrastructure.
  • Advise downstream teams: Partner directly with autonomy, ML, test, infrastructure, product, and customer-facing teams to turn real workflows into reusable platform capabilities from modeling to integrations.
  • Build first-party integrations: Deliver integrations that make important data easy to collect and aggregate, including data produced by simulations, test infrastructure, training systems, and edge devices.
  • Improve developer experience: Create self-service APIs, SDKs, tools, examples, and diagnostics that make correct data modeling and ingestion the easiest path.
  • Drive technical direction: Evaluate emerging data and AI infrastructure technologies, make principled build-versus-buy decisions, and guide implementation across team boundaries.
  • Raise operational quality: Establish expectations for observability, performance, reliability, security, data integrity, disaster recovery, and lifecycle management.

Key outcomes:

  • Human and agentic workflows use one coherent API for discovering data, traversing relationships, and accessing specialized payloads.
  • Teams spend their time deciding how to model and use data rather than repeatedly deciding where and how to store it.
  • Data produced at the edge, in simulation, during testing, and in training flows into reusable platform models with minimal integration friction.
  • Portable and operational platform capabilities across all deployment environments.
  • Downstream teams can adopt the platform through stable APIs and SDKs instead of custom point-to-point integrations.

Required qualifications:

  • Significant experience designing and operating distributed data solutions, storage systems, or data-intensive backend services.
  • Strong software engineering skills and a record of delivering production systems in languages such as Go and Python.
  • Deep understanding of data modeling, API design, schema evolution, identity, consistency, indexing, query planning, and data lifecycle concerns.
  • Experience working across multiple storage modalities, such as relational or graph databases, object storage, analytical or columnar systems, and file storage.
  • Experience designing reliable ingestion and access paths for high-volume or operationally important data.
  • Strong understanding of Kubernetes, Linux, networking, security, storage, observability, and distributed-systems fundamentals.
  • Experience deploying data infrastructure across cloud or customer-managed environments using modern Infrastructure as Code and platform engineering practices.
  • Ability to evaluate technologies through prototypes, benchmarks, operational requirements, and total lifecycle cost rather than feature lists alone.
  • Experience defining architecture and technical standards while remaining hands-on in implementation and debugging.
  • Demonstrated ability to collaborate with ML researchers, autonomy engineers, test teams, platform engineers, and product stakeholders.
  • Clear technical communication and the ability to make complex data architecture understandable to both specialists and downstream users.

Bonus qualifications and relevant technologies:

    Experience in any of the following is beneficial but not required:

  • Experience with specialized modern databases (eg graph, OLAP, etc...)
  • Graph-backed retrieval, agent tooling, structured RAG, provenance-aware context construction, or explainable retrieval systems.
  • S3-compatible APIs, cloud object storage, content-addressable storage, multipart transfer, or large-file lifecycle management.
  • Apache Arrow, Parquet, columnar formats, time-series data, or high-performance analytical query systems.
  • OpenAPI, AsyncAPI, WebSockets, generated SDKs, and long-lived public API contracts.
  • Kubernetes storage and data operators, Terraform, Helm, GitOps, and repeatable platform distribution.
  • Distributed execution technologies such as Ray and experience connecting workflow execution to data lineage and artifact management.
  • Streaming and ingestion technologies such as Kafka, NATS, Redpanda, or comparable event-driven systems.
  • Data and analysis in a robotics or AI domain.
  • ML data lifecycle systems, experiment tracking, dataset management, evaluation infrastructure, feature or artifact stores, and model versioning.
  • Observability tools, distributed tracing, and benchmarking.
  • Data security, authorization, governance, retention, classification, and auditability across shared platforms.

Why join us:

    The Data Platform is foundational to how Shield AI builds, evaluates, certifies, and improves autonomy. This role offers the opportunity to shape both a core internal platform and a reference architecture delivered into demanding customer environments.

    Your work will determine how effectively engineers and agents can find trusted context, understand provenance, aggregate data from distributed systems, and turn operational experience into better intelligent systems. You will work at the intersection of data systems, AI infrastructure, autonomy, distributed computing, and defense while helping establish the architecture that supports the next generation of mission-critical AI development.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
368,634 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
San Diego
Team Lead DevOps 1 day ago
$23k – $62k per year (Estimated) • Remote • 5+ years exp • Moscow
Bash
Python
Erlang
Erlang
EMQX
Databases
Apache Kafka
ClickHouse
PostgreSQL
RabbitMQ
Redis
Redpanda
Trino
DevOps
Ansible
AWS
AWX
FinOps
HAProxy
Hetzner
Kubernetes
SLI/SLO/SLA
Terraform
Yandex Cloud
Amazon S3
Apply
$25k – $42k per year • Equity 0–0.2% • Remote • Full-Time • 3+ years exp
Bash
Go
JavaScript
Python
TypeScript
DevOps
AWS
Azure
CI/CD
Datadog
Docker
GCP
GitHub Actions
GitLab CI
Grafana
Incident Management
Kubernetes
Platform Engineering
Prometheus
Terraform
Amazon CloudWatch
GitHub
GitLab
IAM
Cybersecurity
Least Privilege
Apply
$100k – $210k per year • Equity 0–0.5% • Remote • Full-Time • 3+ years exp • San Francisco
Bash
Go
JavaScript
Python
TypeScript
DevOps
AWS
Azure
CI/CD
Datadog
Docker
GCP
GitHub Actions
GitLab CI
Grafana
Incident Management
Kubernetes
Platform Engineering
Prometheus
Terraform
Amazon CloudWatch
GitHub
GitLab
IAM
Cybersecurity
Least Privilege
Apply
$100k – $252k per year (Estimated) • In office • 6+ years exp • Bachelor's Degree • Herzliya
Java
Databases
Apache Kafka
MySQL
PostgreSQL
RabbitMQ
AI/ML
AI Agents
DevOps
AWS
Azure
Docker
GCP
Kubernetes
Apply
$120k – $220k per year • Equity 0.2–0.8% • Remote/Hybrid • Full-Time • 6+ years exp • San Francisco
C++
Go
JavaScript
Python
TypeScript
Apply
$230k – $350k per year • Remote • Full-Time • 8+ years exp • Bachelor's Degree
Apply
$240k – $350k per year • Remote • Full-Time • 5+ years exp • Bachelor's Degree
Design
SolidWorks
Apply
$160k – $240k per year • In office • Full-Time • 10+ years exp • Bachelor's Degree • Boston
C++
Python
AI/ML
AI Agents
Robotics
Motion Planning
Swarm Robotics
Apply
In office • Internship • 3+ years exp • Bachelor's Degree • Kyiv
C++
Python
Robotics
ArduPilot
MAVLink
QGroundControl
ROS
Apply
$88k – $191k per year (Estimated) • In office • Internship • 7+ years exp • Bachelor's Degree • London
C++
Python
C
C++
CMake
Conan
C
ZeroMQ
Databases
ActiveMQ
DevOps
CI/CD
Docker
gRPC
Rest API
Apply
$70k – $206k per year • In office • Full-Time • 12+ years exp • Associate's Degree • Chicago • Milwaukee • Dallas • Columbus • Kirkland
AI/ML
AI Agents
Apply
$94k – $294k per year • In office • Full-Time • 12+ years exp • Associate's Degree • Chicago • Portland • Milwaukee • Dallas • Columbus
Apply
$70k – $206k per year • In office • Full-Time • 12+ years exp • Associate's Degree • Chicago • Milwaukee • Dallas • Columbus • Kirkland
AI/ML
AI Agents
Apply
$145k – $170k per year • Equity • In office • Full-Time • 3+ years exp • Bachelor's Degree • San Diego
AI/ML
Human-in-the-Loop
Management
Jira
Apply
$143k – $258k per year (Estimated) • Remote/Hybrid • Full-Time • 12+ years exp • Associate's Degree • Chicago • Milwaukee • Dallas • Columbus • Kirkland
JavaScript
Python
TypeScript
Python
pySpark
AI/ML
Prompt Engineering
Spark
DevOps
AWS
Azure
CI/CD
GCP
Git
Jenkins
GitHub
GitLab
Analytics
ETL/ELT
Apply
See all jobs
This is one of many
368,634 more open roles from verified company boards, updated every day.