431,605open jobs
14,862companies
60,035added this week
Browse all
Salary
$146k – $256k per year (Estimated)
Location
Remote (Virginia, Pennsylvania, United States)
Seniority
Senior · 5+ years exp
Employment
Full-Time
Overview
Company
Impact
Profile match

Air

Air, formerly known as Govini, is a defense technology company headquartered in Arlington, Virginia, and founded in 2011. The company provides an AI-native Enterprise Readiness platform that integrates commercial data and enterprise systems to optimize acquisition, supply chain management, and maintenance for national security organizations. It serves various branches of the United States Department of Defense, including the Army and Air Force, by providing real-time data orchestration and predictive analytics to close operational readiness gaps.

Company Description

Air is the leader in Enterprise Readiness. Our mission is to establish readiness as a real-time condition that is continuously achieved. Today, a dangerous Readiness Gap exists between what the front line needs and what is delivered. Our AI-native platform, Air Enterprise Readiness, aligns development, production, delivery, and sustainment into one coordinated execution system for government agencies and industrial suppliers. By revealing true capacity, exposing real constraints, coordinating resources, and executing at the speed of operational demands, the front line gets what it needs to succeed.

Job Description

We are seeking an experienced Senior Machine Learning Engineer to join our AI/ML team and build the infrastructure that powers the development, evaluation, deployment, and continuous improvement of our language models and AI systems.

As our AI capabilities expand, we need robust infrastructure for moving models from experimentation into production. This role will own critical parts of that lifecycle, including LLMOps, fine-tuning infrastructure, model evaluation, dataset pipelines, experiment management, model serving, and production observability.

In order to do this job well: This is an engineering-heavy ML role. You will build platforms and infrastructure that allow AI engineers and researchers to rapidly experiment with models, datasets, and training techniques while maintaining the reproducibility, scalability, and reliability required for production systems.

You will work across the full model lifecycle - from dataset creation and experimentation through training, evaluation, deployment, monitoring, and iteration.

This role is a full-time position based in our Pittsburgh, PA office or open to Remote Opportunities.

This role may require up to 25% travel, including periodic travel to our Pittsburgh, PA and Arlington, VA offices for team collaboration, planning activities, and in-person meetings.

Scope of Responsibilities

  • Design and build LLMOps infrastructure supporting the development, evaluation, deployment, and continuous improvement of production language models.
  • Build scalable training and fine-tuning infrastructure for commercial and open-weight language models.
  • Develop pipelines supporting supervised fine-tuning, parameter-efficient fine-tuning, preference optimization, and other post-training techniques.
  • Build infrastructure for distributed training and GPU-accelerated ML workloads.
  • Develop data pipelines for training, fine-tuning, evaluation, and synthetic data generation.
  • Build systems for dataset versioning, lineage, quality validation, transformation, and reproducible experimentation.
  • Develop experiment management infrastructure that enables engineers to compare models, datasets, hyperparameters, prompts, and training techniques.
  • Build automated evaluation pipelines that determine whether new models or model versions are ready for production deployment.
  • Design model registries, artifact management, versioning, and promotion workflows across development and production environments.
  • Build and operate scalable model-serving and inference infrastructure for open-weight and fine-tuned models.
  • Develop abstractions that allow product and AI engineering teams to use multiple models and inference providers without tightly coupling applications to a single model or vendor.
  • Build observability for model training and inference, including metrics, tracing, logging, resource utilization, model quality, latency, throughput, and cost.
  • Optimize training and inference workloads for GPU utilization, throughput, latency, reliability, and infrastructure cost.
  • Build automated workflows for model deployment, rollback, canarying, and production validation.
  • Investigate model and infrastructure failures across data pipelines, training jobs, inference services, distributed systems, and production environments.
  • Evaluate emerging models, training techniques, inference frameworks, and ML infrastructure and determine where they can improve our production systems.
  • Partner closely with AI engineers building agentic systems to provide the model, evaluation, and training infrastructure required to continuously improve those systems.

Qualifications

  • U.S. Citizenship is required

Required Skills: 

  • 5+ years of experience building production machine learning systems, ML infrastructure, distributed systems, or similar technical systems.
  • Deep experience designing, building, and operating production ML infrastructure or ML platforms.
  • Experience building infrastructure for training, fine-tuning, evaluating, deploying, and monitoring large language models or other large-scale deep learning models.
  • Experience with LLM fine-tuning and post-training workflows, including techniques such as supervised fine-tuning, LoRA/QLoRA or other parameter-efficient approaches, and preference optimization.
  • Strong understanding of the modern LLM lifecycle, including data preparation, training, evaluation, model artifacts, deployment, inference, monitoring, and iteration.
  • Experience building reproducible ML pipelines involving dataset versioning, experiment tracking, model versioning, and automated evaluation.
  • Experience building and operating production GPU infrastructure across AWS, GCP, Azure, or dedicated GPU providers, including training and/or inference workloads.
  • Strong understanding of distributed systems and the challenges involved in running computationally intensive ML workloads at scale.
  • Strong programming experience in Python and experience building production-quality software.
  • Deep experience with containers, Kubernetes, and cloud platforms such as AWS, GCP, or Azure.
  • Experience designing scalable APIs, services, asynchronous workloads, and data-processing pipelines.
  • Strong understanding of observability and operational reliability for production ML systems.
  • Comfortable debugging failures across training code, datasets, models, GPUs, distributed systems, and cloud infrastructure.
  • Able to move between ML experimentation and infrastructure engineering, understanding the needs of researchers and AI engineers while building systems that make those workflows scalable and reproducible.
  • Comfortable working in a rapidly evolving field where tooling, models, and best practices change quickly.
  •  

Desired Skills: 

  • Current possession of a U.S. security clearance, or the ability to obtain one with our sponsorship
  • Experience in or exposure to the nuances of a startup or other entrepreneurial environment
  • Experience building secure code execution environments or sandboxes for AI agents.
  • Experience with multi-agent architectures, agent-to-agent communication, or distributed agent execution.
  • Experience with fine-tuning, post-training, reinforcement learning, or synthetic data generation.
  • Experience building AI observability, tracing, and debugging infrastructure.
  • Experience optimizing inference latency, throughput, GPU utilization, or model-serving costs.
  • Experience with AI security, adversarial testing, or securing agentic systems.
  • Experience working in government, defense, or other mission-critical environments.

We firmly believe that past performance is the best indicator of future performance.  If you thrive while building solutions to complex problems, are a self-starter, and are passionate about making an impact in global security, we’re eager to hear from you.

Air is an Equal Opportunity Employer.  All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, national origin, sexual orientation, gender identity, disability and protected veterans status or any other characteristic protected by law.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
431,605 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
Arlington
$35k – $87k per year (Estimated) • Remote • Full-Time • 3+ years exp • Bachelor's Degree
Python
SQL
Databases
MySQL
PostgreSQL
Snowflake
AI/ML
Spark
Scikit-learn
AI Agents
NLP
TensorFlow
Pandas
NumPy
PyTorch
BERT
Sentiment Analysis
Hugging Face
Amazon SageMaker
DevOps
GCP
Azure
AWS
AWS Lambda
Amazon S3
Analytics
Tableau
Power BI
Matplotlib
ETL/ELT
Management
Microsoft Teams
Apply
Remote • Full-Time • 5+ years exp • Bachelor's Degree
Python
JavaScript
Node JS
Databases
Snowflake
ElasticSearch
DevOps
Rest API
GCP
Kibana
Azure
AWS
Docker
Kubernetes
Amazon EKS
AWS Fargate
AWS Lambda
Amazon ECS
Analytics
Tableau
ETL/ELT
Apply
$22k – $54k per year (Estimated) • Remote • Full-Time • 3+ years exp
Python
JavaScript
AI/ML
AI Agents
DevOps
Rest API
Terraform
CI/CD
Platform Engineering
Management
ServiceNow
Apply
$109k – $145k per year • Remote/Hybrid • Full-Time • 8+ years exp
Python
Go
DevOps
Terraform
Puppet
Ansible
GCP
Chef
Pulumi
CI/CD
GitOps
AWS
Platform Engineering
Chaos Engineering
Configuration Management
SLI/SLO/SLA
Apply
$142k – $235k per year (Estimated) • Remote • Full-Time • 5+ years exp • Bachelor's Degree
Python
AI/ML
Copilot
Cursor
AI Agents
LLM
DevOps
Rest API
Azure DevOps
GitHub Actions
GitLab CI
Azure
CI/CD
Jenkins
Git
GitHub
GitLab
Management
Linear
Jira
Apply
Senior AI Engineer 2 days ago
$143k – $251k per year (Estimated) • Remote • Full-Time • 5+ years exp • Arlington
Python
AI/ML
Fine-tuning
Reinforcement Learning
Function Calling
AI Agents
LLM
Synthetic Data
Post-training
Structured Outputs
Context Engineering
Agentic Workflows
Multi-Agent Systems
Tool Use
DevOps
GCP
Azure
AWS
Kubernetes
Apply
AI Engineer 7 days ago
$96k – $223k per year (Estimated) • In office • Full-Time • 1+ year exp • Bachelor's Degree • Pittsburgh
Python
AI/ML
Claude
Claude Code
Function Calling
AI Agents
LLM
OpenAI
OpenAI Agents SDK
Multi-Agent Systems
Tool Use
DevOps
Git
AWS
Apply
Executive Assistant 7 days ago
$125k – $140k per year • In office • 5+ years exp • Bachelor's Degree • San Francisco
AI/ML
Claude
Gemini
Apply
$139k – $261k per year (Estimated) • In office • Full-Time • 15+ years exp • Bachelor's Degree • Arlington
Apply
$133k – $265k per year (Estimated) • In office • Full-Time • 7+ years exp • Bachelor's Degree • Arlington
Marketing
Salesforce
Apply
$169k – $237k per year • Equity • In office • Full-Time • 8+ years exp • Bachelor's Degree • Seattle • Arlington
Apply
$73k – $153k per year (Estimated) • In office • Full-Time • 7+ years exp • High School Diploma • Arlington
Apply
$112k – $118k per year • In office • Full-Time • 3+ years exp • Bachelor's Degree • Arlington
Apply
Data Scientist 1 day ago
$76k – $127k per year • Remote/Hybrid • Public Trust • 2+ years exp • Bachelor's Degree • Arlington
Python
Databases
Databricks
AI/ML
Copilot
NLP
OpenAI
DevOps
Azure DevOps
Azure
Git
GitHub
Cybersecurity
FedRAMP
Apply
$204k – $284k per year • In office • Top Secret • 15+ years exp • Bachelor's Degree • Arlington
Python
C++
Assembly
Assembly
Binary Ninja
DevOps
QEMU
Cybersecurity
Ghidra
IDA Pro
Apply
See all jobs
This is one of many
431,605 more open roles from verified company boards, updated every day.