368,611open jobs
9,439companies
50,719added this week
Browse all
Salary
$137k – $247k per year (Estimated)
Location
In office
Seniority
Senior
Employment
Full-Time
Overview
Company
Impact
Profile match
Jobgether is an AI-powered job platform focused on remote and flexible work. It matches candidates with relevant roles using skills and preference-based algorithms, and also offers career coaching and job-search guidance.

This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Senior Technical Product Manager - AI Agents, Evals & Reliability based in United States.

This is a deeply technical product leadership role focused on turning advanced AI capabilities into reliable, user-facing experiences.

You will shape systems designed to support long-running workflows, persistent context, multi-step reasoning, and real-world task completion.

Working closely with ML, backend, mobile, and product teams, you will translate technical possibilities into clear product requirements and system decisions.

A major focus will be building robust evaluation frameworks that measure quality, reliability, and user impact across AI systems.

You’ll make high-leverage trade-offs involving quality, latency, cost, reliability, and user experience in a rapidly evolving environment.

The role offers significant ownership within a high-talent-density, hands-on team where decisions move quickly and execution matters.

Your work will directly influence how AI products behave, earn user trust, and deliver practical value at global scale.

Accountabilities:

    • Define end-to-end requirements for AI-powered systems, connecting model capabilities and technical constraints to user needs and measurable product outcomes.
    • Translate model performance, data limitations, evaluation results, and user feedback into clear product and system decisions.
    • Partner closely with ML, backend, mobile, and other engineering teams on architecture, system design, evaluation, iteration, and product delivery.
    • Establish and continuously improve evaluation frameworks spanning offline metrics, online experiments, human feedback, and real-world user outcomes.
    • Make informed trade-offs across product quality, reliability, latency, cost, speed of execution, and overall user experience.
    • Drive execution through clear product specifications, disciplined prioritization, strong technical judgment, and effective cross-functional alignment.
    • Own product quality end-to-end, with particular emphasis on correctness, predictability, reliability, failure handling, and user trust.
    • Identify weaknesses in AI behavior, including hallucinations, bias, model limitations, drift, and other sources of unreliable performance, and drive solutions to address them.
    • Establish strong feedback loops that enable teams to learn quickly from production behavior and continuously improve AI-powered workflows.
    • Ensure new AI capabilities are shipped efficiently and responsibly while maintaining high standards for quality and reliability.
    • Help align product strategy, roadmaps, and priorities with user needs, AI capabilities, and broader business objectives.
    • Bring structure to ambiguous problems and independently drive complex initiatives from definition through execution and measurable outcomes.
    • Requirements:

      • Strong foundation in computer science fundamentals, including algorithms, data structures, and system design, with the ability to engage deeply in technical discussions.
      • Solid understanding of machine learning fundamentals and how modern AI systems behave in production environments.
      • Significant experience owning complex, technically sophisticated products from strategy through execution.
      • Hands-on experience with AI-powered products, particularly LLM-based systems, model evaluation, prompt or pipeline iteration, and AI feedback loops.
      • Strong intuition for AI limitations and failure modes, including hallucinations, bias, model drift, and non-deterministic behavior.
      • Ability to read, review, and discuss technical design documents and confidently collaborate with senior engineers and ML specialists.
      • Demonstrated ability to make sound product and technical decisions in ambiguous, rapidly changing environments.
      • Strong judgment and prioritization skills, with the ability to balance ambitious product goals against technical, operational, and commercial realities.
      • Excellent communication and collaboration skills, with the ability to create alignment across highly technical, cross-functional teams.
      • A hands-on, independent working style and the ability to bring structure to complex problems while maintaining a high pace of execution.
      • Experience working in high-talent-density or small, fast-moving teams where individuals are expected to take significant ownership.
      • Experience shipping AI-heavy consumer products is strongly preferred.
      • An engineering background or experience as a highly technical product manager is a plus.
      • Experience defining evaluation metrics for machine learning systems, developing AI UX patterns, or working in zero-to-one product environments is highly valued.
      • Benefits:

        • Opportunity to shape AI products designed to deliver practical value to billions of potential users.
        • High-ownership role within a small, high-talent-density, hands-on team.
        • Significant influence over AI product strategy, system behavior, evaluation, and reliability.
        • Close collaboration with experienced ML, backend, mobile, and product engineering professionals.
        • Fast-paced environment emphasizing rapid iteration, thoughtful experimentation, and high-quality execution.
        • Opportunity to work on challenging problems involving AI agents, long-running workflows, evaluation, reliability, and real-world task completion.
        • Direct impact on product quality, user trust, and the practical application of advanced AI technologies.
        • Efficient interview process consisting of approximately 3-4 interviews for candidates who advance, conducted virtually and/or onsite.
        • Prompt decision-making and a collaborative environment focused on meaningful outcomes.
Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
368,611 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
In your city
$109k – $118k per year • In office • Contractor • 5+ years exp • Plano
Bash
JavaScript
Node JS
Python
Python
FastAPI
Flask
Gunicorn
Uvicorn
Databases
Meilisearch
pgvector
PostgreSQL
Redis
AI/ML
AI Agents
AWS Bedrock
Claude
LLM
RAG
Model Context Protocol
DevOps
AWS
CI/CD
Docker
Git
Grafana
Incident Management
Jenkins
Kubernetes
Nginx
Prometheus
Rest API
Ubuntu
Amazon CloudWatch
Apply
$173k – $250k per year • In office • Full-Time • San Francisco • New York • Seattle • Austin
TypeScript
Databases
PostgreSQL
AI/ML
AI Agents
Claude
Claude Code
Cursor
LLM
LLM Guardrails
Model Context Protocol
Apply
$31k – $61k per year (Estimated) • In office • Full-Time • 5+ years exp • Bengaluru
Python
SQL
Databases
Presto
AI/ML
AI Agents
Fine-tuning
Kubeflow
LangChain
LangGraph
LangSmith
LLM
LoRA
MLFlow
PEFT
Prompt Engineering
Spark
Transformers
Amazon SageMaker
LLM Guardrails
DevOps
Git
GitHub
Apply
$36k – $81k per year (Estimated) • In office • Gurgaon
AI/ML
ElevenLabs
LangChain
LangGraph
LLM
Prompt Engineering
RAG
AI Agents
DevOps
AWS
Azure
GCP
Apply
$22k – $59k per year (Estimated) • In office • 6+ years exp • Bengaluru
Java
Python
Java
Hibernate
Spring Boot
Databases
ElasticSearch
Neo4j
PostgreSQL
AI/ML
AI Agents
Aider
Claude
Claude Code
Copilot
Cursor
LLM
Context Engineering
Devin
Model Context Protocol
OpenAI Codex
DevOps
AWS
CI/CD
Datadog
GCP
Git
Grafana
New Relic
Prometheus
Apply
$126k – $201k per year • Equity • Remote • Full-Time • 5+ years exp • Bachelor's Degree
Analytics
A/B Testing
Apply
$84k – $166k per year (Estimated) • Remote • Full-Time • 7+ years exp • Bachelor's Degree
SQL
Apply
$80k – $190k per year • Remote • Full-Time • 2+ years exp
Apply
$134k – $223k per year (Estimated) • Remote • Full-Time • 5+ years exp • Bachelor's Degree
Bash
Python
AI/ML
Claude
Claude Code
Copilot
OpenAI Codex
DevOps
Azure
Azure DevOps
CI/CD
Gerrit
Git
Jenkins
KVM
QEMU
RTOS
VMWare
Xen
Cybersecurity
Tcpdump
Wireshark
IoT
FreeRTOS
Management
Confluence
Jira
Apply
$165k – $301k per year (Estimated) • Equity • Remote • Full-Time • 12+ years exp
AI/ML
AI Agents
Apply
See all jobs
This is one of many
368,611 more open roles from verified company boards, updated every day.