685,924open jobs
39,773companies
97,229added this week
Browse all
Salary
$161k – $333k per year (Estimated)
Location
In office
Seniority
Staff
Employment
Full-Time
Overview
Company
Impact
Profile match
Jobgether is a Belgian recruitment platform built entirely around remote and flexible work, aggregating openings from thousands of employers that allow work from outside an office. Its matching engine ranks roles against a candidate's skills, seniority and stated preferences on location and flexibility, rather than leaving people to filter a keyword search, and it verifies how genuinely remote each posting is. The company also runs an AI screening layer that shortlists applicants for employers, and publishes research and guidance on distributed work practices alongside the job marketplace itself.

This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Staff ML Engineer - AWS Trainium & SageMaker based in United States.

As a Staff ML Engineer, you will design, train, optimize, and operate production machine learning workloads on AWS Trainium and Amazon SageMaker.

You will work deeply across the ML stack, from PyTorch training code and distributed execution to accelerator hardware and compiler behavior.

The role requires understanding how training workloads behave on custom silicon rather than treating infrastructure as a black box.

You will diagnose complex training issues, optimize throughput and cost, and build reliable pipelines for real production workloads.

You will work directly with engineering teams to translate business and technical requirements into scalable training solutions.

The environment is highly hands-on, technical, and client-facing, with a focus on solving specialized problems that require deep engineering expertise.

This is an opportunity to work at the intersection of machine learning, cloud infrastructure, distributed systems, and custom AI acceleration.

Accountabilities:

    • Train and operate machine learning models using Amazon SageMaker with AWS Trainium as the underlying compute infrastructure.
    • Develop and optimize PyTorch training workloads for execution on AWS Trainium.
    • Analyze how PyTorch code compiles and executes across Trainium and NeuronCore architecture.
    • Optimize memory utilization, throughput, and other performance characteristics of training workloads.
    • Diagnose training failures and performance issues caused by hardware, compiler behavior, device configuration, data, or model code.
    • Distinguish model- and data-level problems from accelerator- and compiler-level issues during troubleshooting.
    • Translate high-level requirements for Trainium workloads into complete, production-ready training pipelines.
    • Design training workflows that balance performance, scalability, reliability, and cloud infrastructure costs.
    • Tune distributed and multi-device training workloads for throughput and cost efficiency.
    • Operate production training workloads and help ensure their reliability throughout the ML lifecycle.
    • Work directly with client engineering teams to scope, design, and deliver specialized machine learning workloads.
    • Collaborate with internal engineering teams to develop production-grade solutions rather than isolated prototypes or notebook experiments.
    • Investigate complex technical issues across the ML software and hardware stack.
    • Apply a deep understanding of accelerator behavior to improve training architecture and implementation decisions.
    • Contribute to production engineering practices around deployment, monitoring, troubleshooting, and operational reliability.
    • Help translate emerging AI infrastructure capabilities into practical production solutions.
    • Communicate technical findings and tradeoffs clearly with both technical stakeholders and client teams.
    • Requirements:

      • Strong hands-on experience with PyTorch, including a deep understanding of training workflows.
      • Experience with distributed or multi-device model training is highly desirable.
      • Production experience using Amazon SageMaker for model training and/or inference.
      • Strong understanding of how machine learning workloads interact with accelerator hardware and device-specific compilation.
      • Ability to debug issues that originate at the hardware, accelerator, compiler, or runtime layer rather than solely within model or data code.
      • Experience optimizing training workloads for performance, throughput, memory utilization, or cost.
      • Strong Python programming fundamentals.
      • Understanding of distributed training concepts and production ML infrastructure.
      • Ability to design and operate end-to-end managed training pipelines.
      • AWS experience and familiarity with cloud-based machine learning infrastructure.
      • AWS Trainium or AWS Inferentia experience and familiarity with the AWS Neuron SDK is strongly preferred.
      • Candidates without direct Trainium or Inferentia experience may also be considered if they have deep PyTorch expertise and demonstrated ability to quickly learn new hardware targets.
      • Strong analytical and problem-solving skills, particularly when diagnosing complex system-level issues.
      • Ability to reason across multiple layers of the technology stack, from model code through frameworks, compilers, accelerators, and cloud infrastructure.
      • Experience working in a production engineering environment with high standards for reliability and delivery.
      • Strong communication and collaboration skills for working directly with client and internal engineering teams.
      • Comfortable operating in ambiguous environments where requirements and technical challenges may evolve.
      • Willingness and ability to travel approximately 20% within the United States.
      • Must be legally authorized to work in the United States.
      • Benefits:

        • Opportunity to work on production AI systems using AWS Trainium and Amazon SageMaker.
        • Exposure to custom AI accelerator hardware, compiler behavior, distributed training, and advanced ML infrastructure.
        • Hands-on work with complex machine learning workloads for enterprise clients.
        • Direct collaboration with client and internal engineering teams.
        • Opportunity to solve specialized technical problems that extend beyond conventional ML application development.
        • Production-focused engineering environment emphasizing systems that ship and operate at scale.
        • Approximately 20% U.S.-based travel associated with the role.
        • Final interview and onboarding may require onsite participation.
        • Professional growth through work across machine learning, cloud infrastructure, hardware acceleration, and production engineering.
Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
685,924 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
In your city
$170k – $210k per year • Remote • Full-Time • 6+ years exp • PhD
Python
SQL
Databases
MySQL
PostgreSQL
Snowflake
Cassandra
Apache Kafka
AI/ML
Spark
XGBoost
Scikit-learn
Flink
TensorFlow
Pandas
PyTorch
Hugging Face
Metaflow
Frontend
GraphQL
DevOps
gRPC
GCP
Azure
AWS
Docker
Kubernetes
Amazon Kinesis
Apply
Equity • Remote • Full-Time • 8+ years exp
Go
SQL
AI/ML
Claude Code
LLM
DevOps
AWS
Docker
Kubernetes
Management
Agile
Apply
$83k – $204k per year (Estimated) • Remote • Full-Time • 10+ years exp
Python
DevOps
Terraform
AWS
Kubernetes
IAM
Cybersecurity
ISO 27001
SOC 2
Threat Modeling
Web3
Smart Contracts
Staking
Apply
$36k – $88k per year (Estimated) • Remote • Full-Time • 10+ years exp
Python
DevOps
Terraform
AWS
Kubernetes
IAM
Cybersecurity
ISO 27001
SOC 2
Threat Modeling
Web3
Smart Contracts
Staking
Apply
$54k – $146k per year (Estimated) • In office • Full-Time • Sydney
Python
SQL
DevOps
Azure
AWS
FinOps
Analytics
Power BI
Apply
Legal Counsel aa 2 hours ago
Remote • Full-Time
Apply
$71k – $159k per year (Estimated) • Remote • Full-Time • 2+ years exp
Analytics
Power BI
Looker
Apply
VIP Account Manager 2 hours ago
$58k – $131k per year (Estimated) • Remote • Full-Time • PhD
Apply
Equity • Remote • Full-Time • 8+ years exp
Go
SQL
AI/ML
Claude Code
LLM
DevOps
AWS
Docker
Kubernetes
Management
Agile
Apply
$83k – $204k per year (Estimated) • Remote • Full-Time • 10+ years exp
Python
DevOps
Terraform
AWS
Kubernetes
IAM
Cybersecurity
ISO 27001
SOC 2
Threat Modeling
Web3
Smart Contracts
Staking
Apply
See all jobs
This is one of many
685,924 more open roles from verified company boards, updated every day.