368,530open jobs
9,432companies
50,439added this week
Browse all
Salary
$100k – $150k per year
Location
Remote (United States)
Seniority
Senior · 6+ years exp
Employment
Full-Time
Overview
Company
Impact
Profile match
Jobgether is an AI-powered job platform focused on remote and flexible work. It matches candidates with relevant roles using skills and preference-based algorithms, and also offers career coaching and job-search guidance.

This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for an AI Systems Performance Specialist based in United States.

The AI Systems Performance Specialist will optimize large-scale artificial intelligence systems by improving performance, efficiency, and scalability across training and inference workloads.

This role focuses on maximizing throughput, reducing latency, and lowering infrastructure costs through advanced optimization techniques.

You will work across the AI technology stack, from GPU-level optimization and distributed computing to model efficiency and production deployment.

The ideal candidate combines deep machine learning systems expertise with strong engineering discipline and a passion for measurable performance improvements.

This position offers the opportunity to solve complex challenges in AI infrastructure while collaborating with engineering teams building next-generation intelligent systems.

You will contribute to performance standards, optimization strategies, and technical innovations that directly impact production AI capabilities.

Accountabilities:

The AI Systems Performance Specialist will lead efforts to improve the efficiency and reliability of advanced AI workloads through profiling, optimization, and engineering best practices. This role requires strong technical ownership, analytical thinking, and the ability to collaborate across machine learning and infrastructure teams.

  • Profile and optimize end-to-end AI training and inference pipelines to improve throughput, latency, and cost efficiency.
  • Identify performance bottlenecks across data pipelines, model execution, memory usage, communication layers, and infrastructure components.
  • Implement optimization strategies including quantization, sparsity, pruning, and other model efficiency techniques.
  • Optimize distributed training systems using approaches such as tensor parallelism, pipeline parallelism, FSDP, and ZeRO-style sharding.
  • Improve large language model serving performance through techniques such as KV cache optimization, continuous batching, and speculative decoding.
  • Develop and apply compiler-level optimizations using technologies such as Triton, XLA, TorchInductor, or TVM.
  • Optimize data loading, storage access patterns, and dataset sharding strategies for high-performance AI workloads.
  • Build and maintain benchmarking frameworks, regression testing systems, and performance measurement tools.
  • Collaborate with machine learning and platform engineering teams to integrate optimization best practices into production workflows.
  • Drive cost optimization initiatives through improvements in model architecture, hardware utilization, and workload scheduling.
  • Evaluate emerging AI hardware and software technologies and recommend adoption strategies.
  • Create technical documentation, optimization playbooks, and knowledge-sharing materials for engineering teams.
  • Stay current with AI systems research and translate new developments into practical production improvements.

Requirements:

The successful candidate will bring extensive experience in AI systems, performance engineering, or high-performance computing, with a strong ability to analyze and optimize complex machine learning workloads. The ideal profile combines software engineering expertise, deep understanding of modern AI infrastructure, and strong problem-solving skills.

  • Bachelor's or Master's degree in Computer Science, Computer Engineering, or a related technical field.
  • 6+ years of experience in performance engineering, machine learning systems, distributed computing, or high-performance computing environments.
  • Strong programming skills in Python and C++.
  • Hands-on experience optimizing deep learning workloads on modern GPU architectures.
  • Deep understanding of distributed training and inference architectures.
  • Experience using profiling and performance analysis tools across CPU, GPU, and distributed systems.
  • Strong knowledge of memory hierarchies, communication primitives, and parallel computing strategies.
  • Familiarity with model compression techniques and understanding their impact on accuracy and performance.
  • Excellent measurement, debugging, and analytical reasoning abilities.
  • Strong communication and collaboration skills with the ability to work effectively across engineering teams.

Preferred qualifications include:

  • Experience optimizing large language model inference systems at production scale.
  • Contributions to AI infrastructure projects such as vLLM, TensorRT-LLM, DeepSpeed, or similar technologies.
  • Experience developing custom GPU kernels using Triton, CUTLASS, or related frameworks.
  • Familiarity with FinOps practices for managing AI infrastructure costs.
  • Technical publications, conference presentations, or community contributions related to AI systems performance.

Benefits:

  • Fully remote work opportunity within the Continental United States.
  • Competitive annual salary range of approximately $100,000 - $150,000, depending on experience and qualifications.
  • Full-time direct employment opportunity.
  • Opportunity to work on advanced AI systems and large-scale machine learning infrastructure.
  • Career growth opportunities within an innovative technology environment.
  • Exposure to cutting-edge AI optimization techniques and emerging technologies.
  • Collaborative culture focused on engineering excellence, learning, and continuous improvement.
  • Opportunity to contribute to impactful cloud, AI, and enterprise technology solutions.
  • Inclusive workplace committed to equal opportunity and professional development.

How Jobgether works:

We use an AI-powered matching process to ensure your application is reviewed quickly, objectively, and fairly against the role's core requirements. Our system identifies the top-fitting candidates, and this shortlist is then shared directly with the hiring company. The final decision and next steps (interviews, assessments) are managed by their internal team.

We appreciate your interest and wish you the best!

Why Apply Through Jobgether?

Data Privacy Notice: By submitting your application, you acknowledge that Jobgether will process your personal data to evaluate your candidacy and share relevant information with the hiring employer. This processing is based on legitimate interest and pre-contractual measures under applicable data protection laws (including GDPR). You may exercise your rights (access, rectification, erasure, objection) at any time.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
368,530 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
In your city
$98k – $195k per year (Estimated) • In office • Full-Time • 7+ years exp • Wellington
Java
Python
SQL
Java
Spring Boot
Databases
Apache Kafka
Databricks
Neo4j
AI/ML
Flink
Spark
Frontend
GraphQL
DevOps
Azure
CI/CD
Datadog
Dynatrace
Kibana
Kubernetes
OpenShift
Platform Engineering
Splunk
Amazon ECS
Apply
$84k – $178k per year (Estimated) • In office • Full-Time • 10+ years exp • Wellington
Java
Python
DevOps
Ansible
AWS
Azure
CI/CD
Docker
GCP
Helm
Kubernetes
Platform Engineering
Prometheus
Service Mesh
Terraform
GitLab
IAM
Apply
$175k – $195k per year • Remote • 8+ years exp
Python
Python
pySpark
AI/ML
Spark
Edge AI
DevOps
AWS
CI/CD
Incident Management
Platform Engineering
Amazon S3
Analytics
ETL/ELT
Apply
$70k – $105k per year • In office • Full-Time • 3+ years exp
Python
SQL
TypeScript
AI/ML
LLM
RAG
Function Calling
LLM Guardrails
Cybersecurity
GDPR
Management
n8n
Apply
Founding Engineer 1 day ago
$93k – $139k per year • In office • Full-Time • 3+ years exp • Munich
Python
Python
FastAPI
AI/ML
Fine-tuning
LLM
VLM
Apply
$126k – $201k per year • Equity • Remote • Full-Time • 5+ years exp • Bachelor's Degree
Analytics
A/B Testing
Apply
$84k – $166k per year (Estimated) • Remote • Full-Time • 7+ years exp • Bachelor's Degree
SQL
Apply
$80k – $190k per year • Remote • Full-Time • 2+ years exp
Apply
$134k – $223k per year (Estimated) • Remote • Full-Time • 5+ years exp • Bachelor's Degree
Bash
Python
AI/ML
Claude
Claude Code
Copilot
OpenAI Codex
DevOps
Azure
Azure DevOps
CI/CD
Gerrit
Git
Jenkins
KVM
QEMU
RTOS
VMWare
Xen
Cybersecurity
Tcpdump
Wireshark
IoT
FreeRTOS
Management
Confluence
Jira
Apply
$165k – $301k per year (Estimated) • Equity • Remote • Full-Time • 12+ years exp
AI/ML
AI Agents
Apply
See all jobs
This is one of many
368,530 more open roles from verified company boards, updated every day.