818,177open jobs
52,655companies
132,591added this week
Browse all
Salary
$100k – $150k per year
Location
Remote (United States)
Seniority
Staff · 10+ years exp

Confirmed on the employer's own hiring board on Sep 26, 2026. First seen by Alion on Sep 26, 2026.

Overview
Company
Impact
Profile match
Bright Vision Technologies is an IT consulting, enterprise technology services, and workforce solutions enterprise. Headquartered in Bridgewater, New Jersey, United States, the minority-owned firm specializes in technology staffing, cybersecurity, application management, and digital product engineering. Founded in 2020, the enterprise delivers specialized staffing and IT services alongside proprietary automation and AI software - including its flagship enterprise talent intelligence platform, Lumina - serving clients across information technology, defense, healthcare, government, and manufacturing sectors.

AI Data Engineer - Remote

Bright Vision Technologies is a technology consulting and software development company delivering cloud, AI, data, and enterprise solutions across the United States.

This is a fantastic opportunity to join an established and well-respected organization offering tremendous career growth potential.

Job Title: AI Data Engineer

Location: 100% Remote (U.S.)

Position Type: Full-time, Direct W2

Salary Range: $100,000-$150,000 Annually

Experience Required: 10+ years

Sponsorship: U.S. Citizens, Green Card Holders, EAD Holders, and H-1B transfer candidates are encouraged to apply. We are unable to sponsor new H-1B visa petitions for this position.

Job Summary:

We are seeking an AI Data Engineer to build and operate the large-scale data systems that power modern AI training and evaluation pipelines. The role combines deep data engineering expertise with a strong understanding of AI workloads, focusing on ingestion, transformation, quality assurance, lineage, and high-throughput delivery of data to training jobs across diverse modalities. The ideal candidate has experience operating petabyte-scale data systems, strong software engineering fundamentals, and clear understanding of how data infrastructure choices propagate into model quality and training efficiency.

Key Responsibilities

  • Design and operate large-scale data pipelines supporting AI training, evaluation, and continual improvement workflows.
  • Build ingestion systems for diverse modalities including text, image, audio, video, and structured signals.
  • Implement data cleaning, deduplication, filtering, and quality assurance at petabyte scale.
  • Develop dataset versioning, lineage, and provenance tracking systems suitable for reproducible training.
  • Build high-throughput data loading systems that maximize GPU utilization during training.
  • Implement labeling workflows, active learning pipelines, and human-in-the-loop data improvement systems.
  • Design storage architectures balancing cost, throughput, and latency across data tiers.
  • Build evaluation dataset construction pipelines with strict integrity and contamination controls.
  • Implement data privacy, redaction, and consent enforcement throughout the pipeline.
  • Collaborate with ML researchers and engineers to align data systems with model development needs.
  • Drive observability of data quality, drift, and pipeline health across the AI data estate.
  • Optimize cost and performance through compression, format selection, and caching strategies.
  • Document data systems, schemas, and operational procedures for broad internal use.
  • Stay current with AI data infrastructure research and emerging open-source tools.
Required Qualifications
  • Bachelor’s or Master’s degree in Computer Science or a related field.
  • Ten or more years of data engineering experience, with significant work supporting ML or AI workloads.
  • Strong proficiency in Python and at least one JVM or systems language.
  • Deep experience with modern data processing frameworks such as Spark, Ray, or Beam.
  • Hands-on experience operating petabyte-scale storage and pipeline systems.
  • Strong understanding of distributed systems, data modeling, and storage formats.
  • Experience with dataset versioning, lineage, and reproducibility for ML workflows.
  • Familiarity with high-throughput data loading for accelerator-based training.
  • Strong software engineering practices including testing, CI/CD, and code review.
  • Excellent communication and cross-functional collaboration skills.
Preferred Qualifications
  • Experience with multimodal datasets at large scale.
  • Familiarity with data quality tooling and dataset evaluation methodology.
  • Exposure to privacy-preserving data systems and regulated data handling.
  • Open-source contributions to data infrastructure projects.
  • Experience supporting frontier model training pipelines.
How to Apply

Would you like to know more about this opportunity? For immediate consideration, please send your resume to [email protected] or contact us at (908) 505-3899. Learn more about Bright Vision Technologies at www.bvteck.com.

Bright Vision Technologies is an Equal Opportunity Employer.

Equal Employment Opportunity (EEO) Statement

Bright Vision Technologies (BV Teck) is committed to equal employment opportunity (EEO) for all employees and applicants without regard to race, color, religion, sex, sexual orientation, gender identity or expression, national origin, age, genetic information, disability, veteran status, or any other protected status as defined by applicable federal, state, or local laws. This commitment extends to all aspects of employment, including recruitment, hiring, training, compensation, promotion, transfer, leaves of absence, termination, layoffs, and recall.

BV Teck expressly prohibits any form of workplace harassment or discrimination. Any improper interference with employees' ability to perform their job duties may result in disciplinary action up to and including termination of employment.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
818,177 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Data Science
Similar stack
Same company
In your city
Senior Data Engineer 1 month ago
≈ $121k – $209k per year (Estimated) • Remote (United States) • Full-Time • 10+ years exp • Bachelor's Degree • United States
Python
SQL
Databases
Snowflake
Databricks
Apache Kafka
AI/ML
Airflow
DevOps
Azure
Analytics
ETL/ELT
Azure Data Factory
Dimensional Modeling
Apply
$140k per year • Remote (United States) • 5+ years exp • Bachelor's Degree • Somerville
Python
SQL
Databases
PostgreSQL
AI/ML
LangChain
Dagster
dbt
AI Agents
AWS Bedrock
NumPy
Anomaly Detection
Human-in-the-Loop
DevOps
Azure
AWS
Analytics
ETL/ELT
Management
Agile
QA
Pytest
Apply
$145k per year • Remote (United States) • 5+ years exp • Bachelor's Degree • Washington
Python
SQL
Databases
Amazon Aurora
AI/ML
Airflow
Dagster
AI Agents
DevOps
Terraform
Helm
Azure
CI/CD
Git
AWS
Docker
Kubernetes
AWS Lambda
Apply
$145k per year • Remote (United States) • 5+ years exp • Bachelor's Degree • Somerville
Python
SQL
Databases
Amazon Aurora
AI/ML
Airflow
Dagster
AI Agents
DevOps
Terraform
Helm
Azure
CI/CD
Git
AWS
Docker
Kubernetes
AWS Lambda
Apply
≈ $65k – $97k per year (Estimated) • Equity • Remote (Poland) • Full-Time • 5+ years exp • Warsaw
Python
JavaScript
TypeScript
SQL
Python
FastAPI
pySpark
Databases
PostgreSQL
Databricks
OpenSearch
AI/ML
Claude
Spark
Claude Code
Anthropic
Frontend
Zustand
Next.js
React.js
Radix UI
Storybook
DevOps
GCP
Azure
AWS
Platform Engineering
Management
Agile
QA
Playwright
Apply
$160k – $180k per year • Remote (United States) • 6+ years exp • Bachelor's Degree
Python
SQL
Databases
Snowflake
Databricks
Apache Kafka
Google BigQuery
BigQuery
AI/ML
Hadoop
Spark
MLFlow
Reinforcement Learning
Scikit-learn
Computer Vision
Kubeflow
TensorFlow
Pandas
NumPy
PyTorch
Explainable AI
Feature Store
Recommender Systems
Machine Learning
DevOps
GCP
Azure
CI/CD
Git
AWS
Docker
Kubernetes
Analytics
ETL/ELT
Management
Agile
Apply
$145k – $165k per year • Remote (United States) • 6+ years exp • Bachelor's Degree
Python
Java
SQL
Scala
Databases
Databricks
Apache Iceberg
HBase
Apache Kafka
Apache Hudi
Trino
AI/ML
Hadoop
Spark
Airflow
Flink
Machine Learning
DevOps
Azure
CI/CD
AWS
Kubernetes
Analytics
Collibra
Apply
$155k – $175k per year • Remote (United States) • 6+ years exp • Bachelor's Degree
Python
Java
SQL
Scala
Databases
MySQL
PostgreSQL
Snowflake
Databricks
Oracle
Cassandra
RabbitMQ
Apache Kafka
AI/ML
Spark
Airflow
Flink
Machine Learning
DevOps
Terraform
GCP
Helm
Azure
CI/CD
Git
AWS
Docker
Kubernetes
Analytics
ETL/ELT
Apache NiFi
Dimensional Modeling
Master Data Management
Management
Agile
Apply
≈ $28k – $60k per year (Estimated) • In office • Full-Time • India
Python
JavaScript
TypeScript
C#
Node JS
C#
.NET
AI/ML
Anomaly Detection
Frontend
React.js
DevOps
Splunk
GCP
New Relic
OpenTelemetry
Datadog
Dynatrace
Azure
CI/CD
AWS
Grafana
Self-Healing
AppDynamics
AIOps
Incident Management
Apply
$100k – $180k per year • Remote (United States) • 10+ years exp • Bachelor's Degree
Python
Go
JavaScript
Java
Node JS
Bash
Node JS
Commander.js
DevOps
GCP
Istio
OpenTelemetry
Consul
Datadog
Linkerd
Prometheus
Azure
CI/CD
AWS
Kubernetes
Grafana
Chaos Engineering
Service Mesh
Linux
Apply
$160k – $180k per year • Remote (United States) • 6+ years exp • Bachelor's Degree
Python
SQL
Databases
Snowflake
Databricks
Apache Kafka
Google BigQuery
BigQuery
AI/ML
Hadoop
Spark
MLFlow
Reinforcement Learning
Scikit-learn
Computer Vision
Kubeflow
TensorFlow
Pandas
NumPy
PyTorch
Explainable AI
Feature Store
Recommender Systems
Machine Learning
DevOps
GCP
Azure
CI/CD
Git
AWS
Docker
Kubernetes
Analytics
ETL/ELT
Management
Agile
Apply
$145k – $165k per year • Remote (United States) • 6+ years exp • Bachelor's Degree
Python
Java
SQL
Scala
Databases
Databricks
Apache Iceberg
HBase
Apache Kafka
Apache Hudi
Trino
AI/ML
Hadoop
Spark
Airflow
Flink
Machine Learning
DevOps
Azure
CI/CD
AWS
Kubernetes
Analytics
Collibra
Apply
$155k – $175k per year • Remote (United States) • 6+ years exp • Bachelor's Degree
Python
Java
SQL
Scala
Databases
MySQL
PostgreSQL
Snowflake
Databricks
Oracle
Cassandra
RabbitMQ
Apache Kafka
AI/ML
Spark
Airflow
Flink
Machine Learning
DevOps
Terraform
GCP
Helm
Azure
CI/CD
Git
AWS
Docker
Kubernetes
Analytics
ETL/ELT
Apache NiFi
Dimensional Modeling
Master Data Management
Management
Agile
Apply
$120k – $180k per year • Remote (United States) • 6+ years exp • PhD
Python
Java
SQL
PowerShell
Databases
Oracle
DevOps
Rest API
Azure DevOps
Azure
CI/CD
Jenkins
Git
AWS
Analytics
ETL/ELT
Dimensional Modeling
Management
Agile
Scrum
Apply
$100k – $150k per year • Remote (United States) • 6+ years exp • PhD
JavaScript
SQL
Apex
Apex
MuleSoft
Databases
Oracle
DevOps
Rest API
CI/CD
SOAP
Apply
See all jobs
This is one of many
818,177 more open roles from verified company boards, updated every day.