574,821open jobs
24,300companies
79,096added this week
Browse all
Salary
$64k – $171k per year (Estimated)
Location
Remote/Hybrid (Toronto, Canada)
Employment
Full-Time
Overview
Company
Impact
Profile match

Nexxa is building the best AI systems for heavy industries - enabling machines, systems and operations to think, decide and act autonomously across manufacturing, large-scale infrastructure, logistics and legacy environments.

Our mission is to translate deep technical breakthroughs into operational reality, solving some of the hardest systems-level problems in industry.

Role Overview

We're looking for a Lead / Senior / Staff QA Engineer to own quality for Nexxa's AI agent systems - products that plan, call tools, and take multi-step actions autonomously in industrial environments. This isn't traditional UI testing: you'll be designing evaluation frameworks for non-deterministic, tool-using systems, building golden datasets, catching regressions in reasoning quality, and stress-testing agent behavior under adversarial and real-world edge-case conditions.

You'll work closely with ML engineers, backend engineers, and Forward Deployed Engineers to define what "good" looks like for an agent operating in high-stakes industrial settings, then build the infrastructure and processes to measure it continuously.

Key Responsibilities

  • Design and build evaluation harnesses and regression suites for LLM-based agents, covering reasoning quality, tool-call correctness, task completion, and multi-turn coherence.

  • Develop golden datasets and labeled test sets, including edge cases, ambiguous inputs, and adversarial prompts specific to industrial and operational contexts.

  • Define and track quality metrics beyond simple accuracy - groundedness, hallucination rate, task success rate, latency/cost tradeoffs, and safety violations.

  • Build automated pipelines that run evals on every model, prompt, or tool-integration change, and integrate them into CI/CD.

  • Conduct structured red-teaming and adversarial testing (prompt injection, jailbreaks, tool misuse, unsafe actions) in partnership with security teams.

  • Test agent behavior across the full action loop - planning, tool selection, tool execution, error recovery, and final output - not just the final response.

  • Investigate and triage failures where the root cause could be the model, the prompt, the tool/API, or the orchestration logic.

  • Partner with ML and backend engineers to translate eval failures into actionable, reproducible bug reports.

  • Establish quality bars and sign-off criteria for new agent capabilities before they reach customer environments.

  • Mentor other engineers on testing strategies specific to probabilistic, LLM-driven systems.

  • Advocate for testability and observability in agent architecture from day one.

Qualifications

  • 5+ years in QA/SDET roles, with demonstrated ownership of test strategy for complex systems.

  • Hands-on experience testing LLM-based products, chatbots, or AI agents - you understand why traditional deterministic test assertions break down for generative systems.

  • Practical experience with eval frameworks or tooling (e.g., promptfoo, DeepEval, RAGAS, LangSmith) or a track record of building your own.

  • Strong scripting/programming ability (Python preferred) to build test automation, data pipelines, and eval tooling.

  • Understanding of how LLM agents work: prompting, tool/function calling, context management, RAG, memory, and orchestration frameworks.

  • Experience designing test data and labeled datasets, including sourcing, sampling, and managing dataset drift over time.

  • Familiarity with LLM-specific failure modes: hallucination, prompt injection, context poisoning, tool misuse, goal drift, and non-determinism.

  • Comfortable operating in ambiguity - defining what "correct" means for a task when there's no single right answer.

  • Strong written communication skills for turning fuzzy quality signals into clear, actionable findings for engineering and product stakeholders.

Preferred

  • Experience with human-in-the-loop evaluation workflows (labeling pipelines, inter-rater reliability, rubric design).

  • Background in ML/data science sufficient to read model evals and statistical significance.

  • Experience red-teaming or doing adversarial/security testing on ML systems.

  • Familiarity with observability/tracing tools for LLM applications (e.g., LangSmith, Arize, Langfuse, Weights & Biases).

  • Experience testing AI systems in industrial, IoT, or operational technology (OT) environments.

  • Prior experience setting up eval infrastructure from scratch at a startup or fast-moving team.

  • What We're Looking For A QA engineer who wants to define what quality means for autonomous, real-world AI systems.

  • Someone who can build rigorous evaluation infrastructure for problems that don't have a single right answer.

  • A systems thinker who enjoys turning ambiguous agent behavior into measurable, trustworthy signals.

  • A strong collaborator who partners well with ML engineers, backend engineers, and Forward Deployed teams.

Why Join Nexxa.AI?

Innovative Environment: Play a critical role in transforming heavy industries through groundbreaking AI and automation technologies.

Collaborative Culture: Be part of a team that values innovation, discipline, and continuous improvement.

Professional Growth: Benefit from significant opportunities for career development and advancement.

Competitive Compensation: Enjoy a comprehensive salary and equity package reflective of your expertise and contributions.

If you're passionate about AI quality and eager to help define what trustworthy autonomous systems look like in heavy industry, we'd love to connect.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
574,821 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
Toronto
$122k – $354k per year (Estimated) • Remote/Hybrid • Full-Time • Bachelor's Degree • Edinburgh
Python
Go
JavaScript
Java
TypeScript
Java
Spring Boot
Frontend
Angular
DevOps
Helm
Docker
Kubernetes
Management
Agile
Scrum
Apply
$65k – $156k per year (Estimated) • In office • Full-Time • Master's Degree • Edinburgh
Python
MATLAB
MATLAB
Simulink
Apply
$28k – $66k per year (Estimated) • In office • 4+ years exp • Bengaluru
Python
Go
Rust
DevOps
GCP
AWS
Incident Management
Apply
$37k – $85k per year (Estimated) • In office • Full-Time • PhD • Mexico City
AI/ML
AI Agents
Agentforce
Analytics
Microsoft Excel
Marketing
Salesforce
Apply
In office • Full-Time • PhD • Paris
AI/ML
AI Agents
Agentforce
Marketing
Salesforce
Apply
Product Team 1 day ago
$122k – $273k per year (Estimated) • Remote/Hybrid • Full-Time • San Francisco
Design
Figma
Apply
$88k – $243k per year (Estimated) • Equity • Remote • Full-Time • Toronto
Python
Bash
AI/ML
EU AI Act
ISO 42001
DevOps
Terraform
GCP
Pulumi
CI/CD
AWS
Kubernetes
Platform Engineering
Amazon EKS
IAM
Amazon ECS
Cybersecurity
ISO 27001
SOC 2
Least Privilege
Management
Google Workspace
Apply
$129k – $287k per year (Estimated) • Equity • Remote/Hybrid • Full-Time • San Francisco
Python
Bash
AI/ML
EU AI Act
ISO 42001
DevOps
Terraform
GCP
Pulumi
CI/CD
AWS
Kubernetes
Platform Engineering
Amazon EKS
IAM
Amazon ECS
Cybersecurity
ISO 27001
SOC 2
Least Privilege
Management
Google Workspace
Apply
Staff DevOps Engineer 15 days ago
$140k – $266k per year (Estimated) • Equity • Remote • Full-Time • 6+ years exp • Toronto
Python
Go
Databases
Snowflake
Databricks
Google BigQuery
Amazon Redshift
BigQuery
AI/ML
Feature Store
DevOps
Terraform
GCP
GitHub Actions
OpenTelemetry
CircleCI
Datadog
Prometheus
Pulumi
GitLab CI
Azure
CI/CD
ArgoCD
Jenkins
AWS
Kubernetes
Grafana
Platform Engineering
Service Mesh
Incident Management
Cybersecurity
SOC 2
Zero Trust
Apply
Staff DevOps Engineer 15 days ago
$143k – $262k per year (Estimated) • Equity • Remote • Full-Time • 6+ years exp
Python
Go
Databases
Snowflake
Databricks
Google BigQuery
Amazon Redshift
BigQuery
AI/ML
Feature Store
DevOps
Terraform
GCP
GitHub Actions
OpenTelemetry
CircleCI
Datadog
Prometheus
Pulumi
GitLab CI
Azure
CI/CD
ArgoCD
Jenkins
AWS
Kubernetes
Grafana
Platform Engineering
Service Mesh
Incident Management
Cybersecurity
SOC 2
Zero Trust
Apply
$81k – $181k per year (Estimated) • In office • Full-Time • 12+ years exp • Toronto
SQL
Databases
Snowflake
DevOps
Azure
Apply
$93k – $178k per year (Estimated) • In office • Full-Time • 10+ years exp • Toronto
Apply
$56k – $58k per year • In office • Part-Time • Toronto
Apply
$48k per year • In office • Part-Time • 1+ year exp • Toronto
Apply
$34k – $58k per year (Estimated) • In office • Part-Time • 1+ year exp • PhD • Toronto
Apply
See all jobs
This is one of many
574,821 more open roles from verified company boards, updated every day.