380,196open jobs
9,948companies
48,279added this week
Browse all
Salary
$120k – $250k per year (Estimated)
Location
Remote/Hybrid (Canada)
Seniority
Senior · 5+ years exp
Employment
Full-Time
Overview
Company
Impact
Profile match
Jobgether is an AI-powered job platform focused on remote and flexible work. It matches candidates with relevant roles using skills and preference-based algorithms, and also offers career coaching and job-search guidance.

This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Senior Machine Learning Engineer based in Canada.

This is a senior engineering opportunity focused on building the production infrastructure behind AI-powered products used at significant scale.

You will design and operate distributed systems that make machine learning and generative AI capabilities reliable, scalable, and cost-effective in production.

Working alongside Applied Scientists and engineers, you will turn advanced models and research into robust customer-facing systems.

You will help shape the long-term technical direction of AI infrastructure while tackling complex architecture, scalability, reliability, and technical debt challenges.

The role combines hands-on software engineering with technical leadership, mentoring, and cross-functional collaboration.

You will work with modern technologies including Python, SQL, Spark or Dask, AWS, Kubernetes, and emerging LLM technologies.

This is an ideal role for an experienced engineer who enjoys solving complex distributed-systems problems and making AI capabilities dependable at scale.

Accountabilities

    • Lead the design and implementation of large-scale, production-grade distributed systems that support AI and machine learning features used by millions of users.
    • Shape the longer-term technical vision for AI infrastructure in collaboration with staff and senior staff engineers, translating strategic direction into practical, deliverable initiatives.
    • Make architecture decisions that balance scalability, reliability, flexibility, operational simplicity, and cost effectiveness.
    • Own production services and pipelines, including operational health, on-call responsibilities, incident response, monitoring, and technical debt management.
    • Build infrastructure and engineering interfaces that enable Applied Scientists to safely and reliably transition machine learning and LLM models from research into production.
    • Develop and operate scalable data workloads using Python, SQL, and distributed processing technologies such as Spark or Dask.
    • Deploy and maintain production systems across AWS and Kubernetes environments, ensuring they meet appropriate reliability and performance standards.
    • Integrate production-ready generative AI and large language model capabilities into customer-facing product experiences.
    • Improve data usability and engineering practices across the AI Products organization, reducing operational toil and raising overall technical quality.
    • Provide technical leadership and mentorship to engineers, helping raise engineering standards and supporting the development of less-experienced team members.
    • Collaborate across engineering, data, science, and product teams to drive technical initiatives, resolve complex problems, and build consensus around architectural decisions.
    • Identify and address technical debt, infrastructure risks, and opportunities to improve the scalability and maintainability of the AI technology estate.
    • Requirements

      • 5+ years of experience building and operating production software services at scale, with strong proficiency in Python or an equivalent programming language.
      • Strong software engineering fundamentals, including system design, architecture, coding, testing, debugging, and production operations.
      • Proven experience owning production services or data pipelines, including operational or on-call responsibilities, incident response, and long-term technical debt management.
      • Deep understanding of distributed processing principles and practical experience with Spark, Dask, or comparable distributed computing technologies.
      • Strong SQL capabilities and experience working with large-scale data workloads.
      • Demonstrated experience integrating machine learning models or LLM-based capabilities into production systems, with the ability to work effectively alongside Applied Scientists or ML researchers.
      • Production experience with AWS and Kubernetes, including deploying and operating cloud-native workloads.
      • Familiarity with machine learning technologies such as MLFlow, TensorFlow, or PyTorch and data orchestration tools such as Airflow or Prefect is advantageous.
      • Prior experience applying or fine-tuning LLMs in a product environment is a plus, but strong production engineering expertise remains the primary requirement.
      • Strong technical leadership skills, with the ability to set direction, make sound architectural decisions, mentor engineers, and raise engineering standards.
      • Excellent communication and collaboration skills, particularly when working across multidisciplinary teams and translating complex technical concepts into practical decisions.
      • A proactive, pragmatic approach to problem solving, with the ability to navigate ambiguity and drive meaningful technical outcomes.
      • Willingness to participate in an on-call rotation and take ownership of the reliability of production systems.
      • A growth-oriented mindset, curiosity about emerging AI technologies, and enthusiasm for applying new approaches responsibly in production environments.
      • Benefits

        • Annual base salary range of CA$185,000-CA$225,000, with individual compensation determined by factors such as geography, experience, skills, and role level.
        • Eligibility for annual performance bonuses and equity through RSU programs for permanent employees.
        • Potential access to additional performance-based cash or equity incentives depending on role level and company performance.
        • Comprehensive health, wellness, and retirement programs.
        • Wellbeing days and generous paid leave.
        • Dedicated professional development budgets to support ongoing learning and career growth.
        • Flexible hybrid working model combining remote autonomy with access to modern office spaces in Toronto.
        • Collaborative “boost days” designed to support team connection, knowledge sharing, and effective delivery.
        • Opportunity to work on AI-powered products and production infrastructure serving millions of users.
        • A multidisciplinary environment bringing together engineers, Applied Scientists, product managers, analysts, and data specialists.
        • A culture that values skills, impact, curiosity, continuous learning, and diverse perspectives.
        • Inclusive and accessible workplace practices, with support and reasonable accommodations available throughout the hiring process.
Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
380,196 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
In your city
$104k – $213k per year (Estimated) • Equity • Remote • Full-Time • 4+ years exp
AI/ML
Claude
Claude Code
Copilot
Cursor
DevOps
ArgoCD
AWS
CI/CD
GitHub
GitHub Actions
Kubernetes
SLI/SLO/SLA
Cybersecurity
Cortex XSOAR
CVSS
EPSS
Semgrep
Tines
Apply
$64k – $149k per year (Estimated) • Remote/Hybrid • Full-Time • 8+ years exp • Singapore
Java
Java
Spring Boot
Spring Framework
Databases
Apache Kafka
Kafka
Oracle
AI/ML
Prompt Engineering
DevOps
Amazon S3
AWS
Azure
CI/CD
Docker
Git
Kubernetes
OpenShift
Apply
$27k – $67k per year (Estimated) • Remote • Full-Time • 5+ years exp • Bachelor's Degree
Python
Ruby
AI/ML
LLM
DevOps
AWS
Azure
GCP
Apply
$28k – $55k per year (Estimated) • Remote • Full-Time • 5+ years exp
Python
SQL
Python
pySpark
Databases
Snowflake
AI/ML
Claude
Claude Code
Cursor
dbt
Embeddings
LLM
Spark
Structured Outputs
DevOps
Amazon ECS
Amazon S3
AWS
AWS Lambda
Vector
Apply
$124k – $252k per year (Estimated) • Remote • Full-Time
DevOps
AWS
Terraform
Apply
$124k – $252k per year (Estimated) • Remote • Full-Time
DevOps
AWS
Terraform
Apply
$104k – $213k per year (Estimated) • Equity • Remote • Full-Time • 4+ years exp
AI/ML
Claude
Claude Code
Copilot
Cursor
DevOps
ArgoCD
AWS
CI/CD
GitHub
GitHub Actions
Kubernetes
SLI/SLO/SLA
Cybersecurity
Cortex XSOAR
CVSS
EPSS
Semgrep
Tines
Apply
$106k – $211k per year (Estimated) • Remote • Full-Time
AI/ML
Knowledge Graph
Cybersecurity
Open Policy Agent
Apply
$150k – $180k per year • Remote • Full-Time
AI/ML
Knowledge Graph
Cybersecurity
Open Policy Agent
Apply
$139k – $235k per year • Equity • Remote • Full-Time
Ruby
Rust
Ruby
Ruby on Rails
Frontend
GraphQL
DevOps
gRPC
Rest API
Apply
See all jobs
This is one of many
380,196 more open roles from verified company boards, updated every day.