738,299open jobs
44,403companies
105,260added this week
Browse all
Salary
$100k – $150k per year
Location
Remote (United States)
Seniority
Senior · 6+ years exp

Confirmed on the employer's own hiring board on Sep 22, 2026. First seen by Alion on Sep 19, 2026.

Overview
Company
Impact
Profile match
Bright Vision Technologies is an IT consulting, enterprise technology services, and workforce solutions enterprise. Headquartered in Bridgewater, New Jersey, United States, the minority-owned firm specializes in technology staffing, cybersecurity, application management, and digital product engineering. Founded in 2020, the enterprise delivers specialized staffing and IT services alongside proprietary automation and AI software - including its flagship enterprise talent intelligence platform, Lumina - serving clients across information technology, defense, healthcare, government, and manufacturing sectors.

AI Performance Engineer - Remote

Bright Vision Technologies is a technology consulting and software development company delivering cloud, AI, data, and enterprise solutions across the United States.

This is a fantastic opportunity to join an established and well-respected organization offering tremendous career growth potential.

Job Title: AI Performance Engineer

Location: 100% Remote (U.S.)

Position Type: Full-time, Direct W2

Salary Range: $100,000-$150,000 Annually

Experience Required: 6+ years

Sponsorship: U.S. Citizens, Green Card Holders, EAD Holders, and H-1B transfer candidates are encouraged to apply. We are unable to sponsor new H-1B visa petitions for this position.

Key Responsibilities

  • Profile and optimize end-to-end AI training and inference pipelines for throughput, latency, and cost.
  • Identify and eliminate bottlenecks across data loading, model compute, communication, and memory.
  • Implement and tune quantization, sparsity, and pruning strategies to reduce model footprint and accelerate inference.
  • Optimize distributed training using tensor parallelism, pipeline parallelism, FSDP, and ZeRO-style sharding.
  • Tune attention implementations using FlashAttention, paged attention, and related techniques.
  • Implement KV cache optimization, continuous batching, and speculative decoding for LLM serving.
  • Drive compiler-level optimizations using Triton, XLA, TorchInductor, or TVM, working with the broader ML framework community to land improvements that translate into measurable end-to-end performance gains.
  • Optimize data pipelines, sharding strategies, and storage access patterns for high-throughput training.
  • Build and maintain rigorous benchmark suites and regression frameworks across workloads.
  • Collaborate with ML and platform engineering teams to embed best practices in standard pipelines.
  • Drive cost-efficiency improvements through model architecture, hardware selection, and scheduling strategies.
  • Evaluate new hardware and software offerings, and advise on adoption.
  • Document performance tuning playbooks and share findings broadly across engineering teams.
  • Stay current with AI systems research and translate advances into production improvements.
Required Qualifications
  • Bachelor’s or Master’s degree in Computer Science, Computer Engineering, or a related field.
  • Six or more years of experience in performance engineering, ML systems, or HPC.
  • Strong proficiency in Python and C++.
  • Hands-on experience optimizing deep learning workloads on modern GPUs.
  • Deep understanding of distributed training and inference techniques.
  • Experience with profiling tools across CPU, GPU, and distributed systems.
  • Familiarity with model compression techniques and their accuracy implications.
  • Strong grasp of memory hierarchies, communication primitives, and parallelism strategies.
  • Excellent measurement, debugging, and analytical reasoning skills.
  • Strong communication and collaboration skills.
Preferred Qualifications
  • Experience optimizing LLM inference at production scale.
  • Contributions to vLLM, TensorRT-LLM, DeepSpeed, or similar projects.
  • Familiarity with custom kernel authoring in Triton or CUTLASS.
  • Experience with FinOps for AI workloads.
  • Publications or talks on AI systems performance.
How to Apply

Would you like to know more about this opportunity? For immediate consideration, please send your resume to [email protected] or contact us at (908) 505-3899. Learn more about Bright Vision Technologies at www.bvteck.com.

Bright Vision Technologies is an Equal Opportunity Employer.

Equal Employment Opportunity (EEO) Statement

Bright Vision Technologies (BV Teck) is committed to equal employment opportunity (EEO) for all employees and applicants without regard to race, color, religion, sex, sexual orientation, gender identity or expression, national origin, age, genetic information, disability, veteran status, or any other protected status as defined by applicable federal, state, or local laws. This commitment extends to all aspects of employment, including recruitment, hiring, training, compensation, promotion, transfer, leaves of absence, termination, layoffs, and recall.

BV Teck expressly prohibits any form of workplace harassment or discrimination. Any improper interference with employees' ability to perform their job duties may result in disciplinary action up to and including termination of employment.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
738,299 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

AI/ML
Similar stack
Same company
In your city
Senior AI Engineer 2 hours ago
$113k – $246k per year (Estimated) • Remote (location not specified) • 5+ years exp
Python
AI/ML
LangChain
Fine-tuning
Embeddings
Scikit-learn
AI Agents
TensorFlow
PyTorch
LLM
RAG
Agentic Workflows
Machine Learning
DevOps
GCP
Azure
CI/CD
AWS
Cloudflare
Management
Notion
ServiceNow
Apply
$117k – $256k per year (Estimated) • Remote (likely Hungary) • 5+ years exp
AI/ML
Computer Vision
TensorFlow
PyTorch
Ray
Feature Store
Recommender Systems
DevOps
CI/CD
Apply
$115k – $250k per year (Estimated) • Remote (likely Hungary) • 5+ years exp
Python
AI/ML
LoRA
vLLM
Quantization
AI Agents
PEFT
Transformers
LLM
RAG
LLM Evaluation
LLM Guardrails
Tool Use
DevOps
Rest API
gRPC
GCP
Azure
CI/CD
AWS
Apply
$171k – $309k per year (Estimated) • Remote (United States) • Full-Time • 6+ years exp
Python
Go
JavaScript
Python
Flask
FastAPI
Pydantic
Databases
PostgreSQL
Weaviate
Milvus
pgvector
Pinecone
AI/ML
LangGraph
LangChain
DSPy
LlamaIndex
Quantization
Multimodal AI
Function Calling
AI Agents
Speech Recognition
Pydantic AI
LLM
RAG
Hybrid Search
Text-to-Speech
GraphRAG
Human-in-the-Loop
KV Cache
Multi-Agent Systems
Tool Use
Frontend
WebGPU
DevOps
Terraform
GCP
GitHub Actions
OpenTelemetry
WebRTC
CI/CD
AWS
Docker
Kubernetes
Cybersecurity
SOC 2
GDPR
HIPAA
Apply
$149k – $261k per year (Estimated) • Equity • Remote (United States) • Full-Time • 5+ years exp • Bachelor's Degree
Python
Databases
Databricks
AI/ML
Multimodal AI
PyTorch
Self-Supervised Learning
Time Series Forecasting
World Models
Machine Learning
DevOps
AWS
Apply
$20k – $44k per year (Estimated) • In office • Full-Time • 6+ years exp • Bachelor's Degree • Pune
Python
PowerShell
DevOps
Azure
Platform Engineering
Incident Management
Windows
Cybersecurity
Zero Trust
Microsoft Entra ID
Active Directory
Management
ITIL
Apply
$32k – $79k per year (Estimated) • Remote (EAEU) • Full-Time • 5+ years exp • Moscow
Python
JavaScript
PHP
TypeScript
Python
FastAPI
Databases
PostgreSQL
Redis
RabbitMQ
AI/ML
Copilot
Claude
Claude Code
AI Agents
Cline
LLM
Anthropic
Frontend
React.js
DevOps
Rest API
CI/CD
Git
Docker
Kubernetes
Gitflow
GitHub
Apply
$23k per year (gross) • In office • Full-Time • Kazan
Python
Python
SQLAlchemy
Databases
PostgreSQL
ClickHouse
Milvus
FAISS
Qdrant
AI/ML
llama.cpp
vLLM
AI Agents
Axolotl
Ollama
Open WebUI
PEFT
Transformers
PyTorch
LLM
RAG
Hugging Face
DevOps
Git
Docker
Linux
Apply
$15k – $36k per year (Estimated) • Remote/Hybrid • Saint Petersburg
Python
TypeScript
SQL
AI/ML
Cursor
Claude Code
Model Context Protocol
Function Calling
LLM
RAG
OpenAI Codex
Tool Use
DevOps
Git
Docker
Linux
Management
n8n
Apply
In office • Bengaluru
AI/ML
AI Agents
NLP
LLM
Marketing
LinkedIn
Apply
$145k – $165k per year • Remote (United States) • 6+ years exp • Bachelor's Degree
Python
Rust
C++
AI/ML
vLLM
Quantization
Knowledge Distillation
TensorRT
TensorRT-LLM
LLM
Model Distillation
Machine Learning
DevOps
Kubernetes
Platform Engineering
FinOps
Apply
$155k – $180k per year • Remote (United States) • 6+ years exp • Master's Degree
Python
AI/ML
RLHF
Reinforcement Learning
AI Agents
Reward Modeling
Machine Learning
Robotics
Reinforcement Learning
Apply
$135k – $210k per year • Remote (United States) • 6+ years exp • Master's Degree
Python
AI/ML
Fine-tuning
JAX
Multimodal AI
AI Agents
PyTorch
RAG
Machine Learning
Apply
$130k – $180k per year • Remote (United States) • 10+ years exp • Master's Degree
Python
AI/ML
LangGraph
LangChain
DeepSpeed
LlamaIndex
LoRA
Fine-tuning
RLHF
Multimodal AI
AI Agents
NLP
PEFT
QLoRA
Transformers
PyTorch
LLM
RAG
Ray
Hallucination
Synthetic Data
DPO
SFT
PPO
FSDP
Knowledge Graph
Machine Learning
DevOps
GCP
Azure
AWS
Docker
Kubernetes
Apply
$80k – $100k per year • Remote (United States) • 6+ years exp • Bachelor's Degree
Python
AI/ML
Spark
Multimodal AI
Ray
Machine Learning
DevOps
CI/CD
Apply
See all jobs
This is one of many
738,299 more open roles from verified company boards, updated every day.