Confirmed on the employer's own hiring board on Oct 9, 2026. First seen by Alion on Sep 4, 2026. ID.me scores B on the Alion truth index.
Join ID.me as a Staff Software Engineer specializing in AI Agent Evaluations. In this role, you will define and lead the discipline of testing AI agents, evaluating LLM behavior, and ensuring the reliability of agentic systems in production. You will establish quality standards for AI-powered features, build evaluation infrastructure, and drive an AI-first engineering culture. This position requires deep engineering rigor, original thinking about correctness in non-deterministic systems, and the ability to build evaluation infrastructure and developer tooling.
Missions
- Definir y liderar la disciplina de pruebas de agentes de IA, evaluando el comportamiento de los LLM y asegurando la fiabilidad de los sistemas en producción.
- Establecer estándares de calidad para cómo ID.me envía características impulsadas por IA de manera segura, y mentorear a los ingenieros en las mejores prácticas de prueba de IA.
- Diseñar y mantener las tuberías de evaluación para las salidas de LLM, el comportamiento del agente, el uso de herramientas y las interacciones de múltiples turnos.
Profil recherché
- Demonstrated experience evaluating or testing LLM-powered features or autonomous agents in production- Experience designing test infrastructure, CI/CD quality gates, or evaluation pipelines at scale
- Strong written and verbal communication across engineering, product, and leadership
- 8+ years building and operating production software systems
- Strong backend engineering fundamentals in Python, Java, Go, or equivalent
- Experience building eval frameworks for LLM agents (e.g., correctness graders, LLM-as-judge, human-in-the-loop evals, benchmark dataset curation)
- Bachelor's degree in Computer Science, Engineering, or equivalent experience
- Red-teaming or adversarial testing experience for AI models or agents
- Experience improving developer experience - building internal tooling, reducing toil, or accelerating engineering workflows
- Familiarity with agentic frameworks (Claude API / Anthropic SDK, BrainTrust, LangChain, LangGraph, CrewAI, or similar)
- Production monitoring experience for AI systems: behavioral drift detection, output sampling, shadow scoring
- Proven ability to lead cross-team technical initiatives and influence engineering standards
- Proficiency with AI-assisted development tools (Claude Code, Cursor, or equivalent) - you build with AI every day
- Background in identity verification, fraud detection, or regulated industries
- Familiarity with Anthropic's model evaluation methodology or similar published eval research
- Experience with observability tooling (Datadog, OpenTelemetry) applied to AI workloads
- Track record of building developer tooling or platforms that other teams adopt widely

