Overview
BigBear.ai is seeking an AI/ML Data Engineer to architect, build, and operate enterprise-scale AI data platforms enabling vector databases, semantic search, Retrieval-Augmented Generation (RAG), and agentic AI systems. This role will establish the technical and governance foundations for production AI, including data lineage, source attribution, prompt/context traceability, explainability, and evaluation in mission environments.
What you will do
What You’ll Do
- Architect and evolve an enterprise AI data platform enabling vector search, semantic retrieval, RAG, and agentic workflows.
- Design cloud-native, distributed data systems optimized for performance, scale, security, reliability, and cost.
Establish and implement controls for:
- Data quality, lineage, provenance/source attribution
- Prompt + context traceability and auditability
- Explainability and evaluation of AI outputs
- Partner across Data Science, ML Engineering, Software, Cybersecurity, and Enterprise Architecture to translate AI requirements into production-grade capabilities.
- Lead architectural decisions; drive reuse across organizations and eliminate duplication. Implement monitoring/observability/alerting and operational excellence best practices.
- Mentor engineers and raise engineering standards (design reviews, coding standards, CI/CD discipline).
- Evaluate emerging AI technologies (vector DBs, retrieval frameworks, evaluation stacks) and recommend adoption paths.
What you should have
- Active Top Secret / SCI with Polygraph is mandatory.
- Bachelor's in CS/Engineering/Math/Data Science (or equivalent experience).
- 20+ years in software engineering, data engineering, distributed systems, cloud architecture, or AI/ML platform development.
- Proven delivery of enterprise-scale AI/ML / GenAI / agentic systems.
- Track record of architecting production cloud-native data platforms supporting AI workloads.
- Experience in complex enterprises with security constraints, dependencies, governance, and competing priorities.
- Deep experience with data pipelines supporting ML models, vector databases, semantic search, and GenAI apps.
- End-to-end delivery ownership from strategic requirements through operational deployment.
Technical:
- Expert in Python and SQL; strong software engineering practices.
- Deep experience with AWS, Azure, or GCP data/AI platforms.
- Strong understanding of distributed systems, cloud-native architecture, MLOps, platform engineering.
- Hands-on with vector DBs, embeddings, retrieval systems, RAG.
- Experience with CI/CD, orchestration, IaC/automation, observability.
Leadership & Communication:
- Exceptional written/verbal communication.
- Able to translate complex AI/data architecture concepts into mission impact, risk, and tradeoffs.
- Demonstrated ability to influence and drive consensus across stakeholders.
What we'd like you to have
- Prior roles as Principal Engineer, Lead Data Engineer, Solutions Architect, Technical Lead.
- Built/operated enterprise vector search/knowledge management / RAG / LLM platforms.
- Large-scale distributed processing (e.g., Spark/Flink/Beam-whatever aligns to your stack).
- Experience supporting AI adoption in government/defense/intelligence or highly regulated environments.
- Familiarity with AI governance, evaluation frameworks, explainability, responsible AI.
About BigBear.ai
BigBear.ai is a leading provider of AI-powered decision intelligence solutions for national security, supply chain management, and digital identity. Customers and partners rely on Bigbear.ai’s predictive analytics capabilities in highly complex, distributed, mission-based operating environments. Headquartered in McLean, Virginia, BigBear.ai is a public company traded on the NYSE under the symbol BBAI. For more information, visit https://bigbear.ai/ and follow BigBear.ai on LinkedIn: @BigBear.ai and X: @BigBearai.
BigBear.ai is an Equal opportunity employer all protected groups, including protected veterans and individuals with disabilities.

