1,432,392open jobs
84,555companies
217,679added this week
Browse all
Salary
$100k – $150k per year
Location
Remote (United States)
Seniority
Senior · 6+ years exp
Visa
No sponsorship (stated in the posting)

Confirmed on the employer's own hiring board on Oct 10, 2026. First seen by Alion on Oct 3, 2026.

Overview
Company
Impact
Profile match
Bright Vision Technologies is an IT consulting, enterprise technology services, and workforce solutions enterprise. Headquartered in Bridgewater, New Jersey, United States, the minority-owned firm specializes in technology staffing, cybersecurity, application management, and digital product engineering. Founded in 2020, the enterprise delivers specialized staffing and IT services alongside proprietary automation and AI software - including its flagship enterprise talent intelligence platform, Lumina - serving clients across information technology, defense, healthcare, government, and manufacturing sectors.

ML Performance Engineer - Remote

Bright Vision Technologies is a technology consulting and software development company delivering cloud, AI, data, and enterprise solutions across the United States.

This is a fantastic opportunity to join an established and well-respected organization offering tremendous career growth potential.

Job Title: ML Performance Engineer

Location: 100% Remote (U.S.)

Position Type: Full-time, Direct W2

Salary Range: $100,000-$150,000 Annually

Experience Required: 6+ years

Sponsorship: U.S. Citizens, Green Card Holders, EAD Holders, and H-1B transfer candidates are encouraged to apply. We are unable to sponsor new H-1B visa petitions for this position.

Job Summary

We are seeking an AI Performance Optimization Engineer to focus on extracting maximum throughput, minimizing latency, and reducing cost across training and inference workloads for large neural network systems. The role spans the full stack from low-level kernel optimization to distributed system tuning, requiring deep understanding of GPU architecture, model parallelism, memory management, and compiler-level optimization. The ideal candidate has demonstrated impact on production AI workloads, with strong instrumentation and measurement discipline that enables rigorous, data-driven optimization decisions. In this role you will work closely with cross-functional partners - product, design, engineering, operations, and business stakeholders - to translate ambiguous requirements into well-engineered solutions, and will be expected to raise the bar through code review, design review, and mentorship of more junior engineers. The successful candidate brings strong engineering discipline, a clear communication style, and a track record of shipping meaningful work that holds up well in production.

Key Responsibilities

  • Profile and optimize end-to-end AI training and inference pipelines for throughput, latency, and cost.
  • Identify and eliminate bottlenecks across data loading, model compute, communication, and memory.
  • Implement and tune quantization, sparsity, and pruning strategies to reduce model footprint and accelerate inference.
  • Optimize distributed training using tensor parallelism, pipeline parallelism, FSDP, and ZeRO-style sharding.
  • Tune attention implementations using FlashAttention, paged attention, and related techniques.
  • Implement KV cache optimization, continuous batching, and speculative decoding for LLM serving.
  • Drive compiler-level optimizations using Triton, XLA, TorchInductor, or TVM, working with the broader ML framework community to land improvements that translate into measurable end-to-end performance gains.
  • Optimize data pipelines, sharding strategies, and storage access patterns for high-throughput training.
  • Build and maintain rigorous benchmark suites and regression frameworks across workloads.
  • Collaborate with ML and platform engineering teams to embed best practices in standard pipelines.
  • Drive cost-efficiency improvements through model architecture, hardware selection, and scheduling strategies.
  • Evaluate new hardware and software offerings, and advise on adoption.
  • Document performance tuning playbooks and share findings broadly across engineering teams.
  • Stay current with AI systems research and translate advances into production improvements.
Required Qualifications
  • Bachelor’s or Master’s degree in Computer Science, Computer Engineering, or a related field.
  • Six or more years of experience in performance engineering, ML systems, or HPC.
  • Strong proficiency in Python and C++.
  • Hands-on experience optimizing deep learning workloads on modern GPUs.
  • Deep understanding of distributed training and inference techniques.
  • Experience with profiling tools across CPU, GPU, and distributed systems.
  • Familiarity with model compression techniques and their accuracy implications.
  • Strong grasp of memory hierarchies, communication primitives, and parallelism strategies.
  • Excellent measurement, debugging, and analytical reasoning skills.
  • Strong communication and collaboration skills.
Preferred Qualifications
  • Experience optimizing LLM inference at production scale.
  • Contributions to vLLM, TensorRT-LLM, DeepSpeed, or similar projects.
  • Familiarity with custom kernel authoring in Triton or CUTLASS.
  • Experience with FinOps for AI workloads.
  • Publications or talks on AI systems performance.
How to Apply

Would you like to know more about this opportunity? For immediate consideration, please send your resume to [email protected] or contact us at (908) 505-3544. Learn more about Bright Vision Technologies at www.bvteck.com.

Bright Vision Technologies is an Equal Opportunity Employer.

Equal Employment Opportunity (EEO) Statement

Bright Vision Technologies (BV Teck) is committed to equal employment opportunity (EEO) for all employees and applicants without regard to race, color, religion, sex, sexual orientation, gender identity or expression, national origin, age, genetic information, disability, veteran status, or any other protected status as defined by applicable federal, state, or local laws. This commitment extends to all aspects of employment, including recruitment, hiring, training, compensation, promotion, transfer, leaves of absence, termination, layoffs, and recall.

BV Teck expressly prohibits any form of workplace harassment or discrimination. Any improper interference with employees' ability to perform their job duties may result in disciplinary action up to and including termination of employment.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
1,432,392 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

AI/ML
Similar stack
Same company
In your city
$170k – $241k per year • Remote (United States) • Full-Time • 5+ years exp • Bachelor's Degree • Sunnyvale
Python
AI/ML
TensorFlow
PyTorch
FSDP
Machine Learning
DevOps
GCP
Azure
AWS
Apply
$159k – $231k per year • Remote (United States) • Full-Time • Master's Degree • Sunnyvale • Washington
Python
SQL
AI/ML
Spark
Fine-tuning
Reinforcement Learning
Multimodal AI
Pandas
NumPy
PyTorch
Self-Supervised Learning
Synthetic Data
SFT
Pre-training
Embodied AI
Machine Learning
Mobile
AVFoundation
Robotics
Sim-to-Real
Imitation Learning
Reinforcement Learning
Apply
$172k – $304k per year • Remote (United States) • Full-Time • 8+ years exp • Sunnyvale • Washington
Python
Go
JavaScript
TypeScript
SQL
AI/ML
Computer Vision
Agentic Workflows
Machine Learning
Frontend
Redux
GraphQL
React.js
DevOps
gRPC
CI/CD
Analytics
A/B Testing
Apply
AI Engineer 2 hours ago
≈ $48k – $120k per year (Estimated) • Remote (Argentina) • Full-Time • 7+ years exp
Python
Databases
PostgreSQL
Weaviate
pgvector
Pinecone
AI/ML
Cursor
Claude Code
Embeddings
Function Calling
AI Agents
LLM
RAG
OpenAI
Anthropic
OpenAI Codex
Human-in-the-Loop
Structured Outputs
Tool Use
DevOps
Rest API
GCP
Azure
AWS
Apply
AI Engineer 2 hours ago
≈ $48k – $119k per year (Estimated) • Remote (Mexico) • Full-Time • 7+ years exp • Mexico City
Python
Databases
PostgreSQL
Weaviate
pgvector
Pinecone
AI/ML
Cursor
Claude Code
Embeddings
Function Calling
AI Agents
LLM
RAG
OpenAI
Anthropic
OpenAI Codex
Human-in-the-Loop
Structured Outputs
Tool Use
DevOps
Rest API
GCP
Azure
AWS
Apply
$124k – $196k per year • In office • Full-Time • 2+ years exp • Bachelor's Degree • Santa Clara
Python
C++
C++
PyTorch C++
AI/ML
CUDA Toolkit
JAX
Transformers
PyTorch
LLM
CUDA
Triton
OpenAI
cuDNN
Apply
$212k – $286k per year • In office • Secret • Full-Time • 22+ years exp • Bachelor's Degree • El Segundo
Python
JavaScript
Rust
C++
Perl
MATLAB
Visual Basic
MATLAB
Simulink
DevOps
Azure
CI/CD
Git
AWS
Linux
Management
Agile
Apply
$106k – $144k per year • In office • Top Secret • Full-Time • 13+ years exp • Bachelor's Degree • Seal Beach
Python
Java
SQL
C++
Databases
Oracle
Teradata
DevOps
Rest API
CI/CD
Linux
Analytics
ETL/ELT
Management
Agile
Apply
$99k – $135k per year • In office • Secret • Full-Time • 13+ years exp • Bachelor's Degree • Saint Charles
Python
C#
C++
Ada
DevOps
Git
Bitbucket
Management
Confluence
Jira
Agile
Apply
$152k – $242k per year • In office • Full-Time • 5+ years exp • Master's Degree • Santa Clara
Python
C
C++
C
Pthreads
MPI
AI/ML
CUDA Toolkit
AI Agents
OpenMP
CUDA
DevOps
CI/CD
HPC
Apply
$130k – $180k per year • Remote (United States) • 10+ years exp • Master's Degree
Python
AI/ML
LangGraph
LangChain
DeepSpeed
LlamaIndex
LoRA
Fine-tuning
RLHF
Multimodal AI
AI Agents
NLP
PEFT
QLoRA
Transformers
PyTorch
LLM
RAG
Ray
Hallucination
Synthetic Data
DPO
SFT
PPO
FSDP
Knowledge Graph
Machine Learning
DevOps
GCP
Azure
AWS
Docker
Kubernetes
Apply
$130k – $180k per year • Remote (likely United States) • 10+ years exp • Bachelor's Degree
Python
AI/ML
LangGraph
AutoGen
LangChain
Claude
LlamaIndex
LoRA
Fine-tuning
Embeddings
Prompt Engineering
AI Agents
PEFT
QLoRA
Semantic Kernel
Llama
Mistral
CrewAI
Gemini
LLM
RAG
Hallucination
OpenAI
Anthropic
LLM Guardrails
Multi-Agent Systems
Tool Use
Machine Learning
DevOps
GCP
Azure
CI/CD
AWS
Kubernetes
Cybersecurity
PCI DSS
SOC 2
HIPAA
FedRAMP
Apply
$130k – $180k per year • Remote (United States) • 10+ years exp • Bachelor's Degree
Python
Databases
Weaviate
Milvus
Pinecone
FAISS
AI/ML
LangGraph
LangChain
LlamaIndex
Fine-tuning
Embeddings
JAX
Prompt Engineering
Multimodal AI
Computer Vision
AI Agents
NLP
TensorFlow
PyTorch
LLM
RAG
Ray
Recommender Systems
Machine Learning
DevOps
GCP
Azure
CI/CD
AWS
Docker
Kubernetes
Apply
$130k – $180k per year • Remote (United States) • 10+ years exp • Bachelor's Degree
Python
C++
C++
PyTorch C++
AI/ML
DeepSpeed
vLLM
CUDA Toolkit
Triton Inference Server
Quantization
TensorRT
TensorRT-LLM
PyTorch
LLM
TensorBoard
Ray
CUDA
NCCL
ROCm
CUTLASS
Speculative Decoding
KV Cache
Machine Learning
DevOps
GCP
Azure
AWS
FinOps
HPC
Apply
$100k – $150k per year • Remote (United States) • 6+ years exp • Bachelor's Degree
Python
Rust
C++
AI/ML
vLLM
Quantization
Knowledge Distillation
TensorRT
TensorRT-LLM
LLM
Recommender Systems
Speculative Decoding
KV Cache
Model Distillation
Machine Learning
DevOps
Kubernetes
Platform Engineering
FinOps
Apply
See all jobs
This is one of many
1,432,392 more open roles from verified company boards, updated every day.