{"id":831402,"url":"https://alion.io/job/judgmentlabs-product-engineer-agent","title":"Product engineer, Agent","company":{"id":679825,"name":"Judgment Labs","domain":"judgmentlabs.ai","url":"https://alion.io/company/judgmentlabs","size_band":"51-200","is_staffing_agency":false,"employer_type":"direct","is_intermediary":false,"listed_via":null,"ats_vendor":"Ashby","truth_index":null},"role":"Product","role_family":"Product","seniority":null,"employment_type":"full_time","work_mode":"on_site","remote_scope":null,"remote_scope_basis":null,"remote_working_hours":null,"hiring_geo_confidence":"structured","locations":["San Francisco, United States"],"countries":["US"],"hiring_countries":[],"hiring_countries_total":0,"salary":null,"salary_estimate":{"min_usd":152000,"max_usd":316000,"period":"year","method":"role_country_seniority_unknown","sample_n":4538},"experience_years_min":null,"visa_sponsorship":false,"relocation_package":false,"has_equity":false,"technologies":[{"name":"AI Agents","optional":false},{"name":"Claude Code","optional":false},{"name":"OpenAI Codex","optional":false}],"status":"live","first_seen_at":"2026-08-13T07:25:41Z","employer_posted_date":"2026-08-13","last_verified_at":"2026-09-29T16:44:17Z","board_verified":true,"closed_at":null,"days_open":47,"trust":{"level":"ok","repost_count":null,"flags":[],"days_open":47},"description":"Product Engineer - Agents Job Description\nThe Role\nJudgment is the learning infrastructure for AI agents. Agents in production don't improve from prompts alone. They improve from experience: the tasks they attempt, the mistakes they make, the edge cases they hit. Here's how it works:\nWe ingest everything your agents do in production: traces, tool calls, decisions, outcomes\n\nJudgment turns that raw experience into structured signals: failure modes, behaviors, rubrics, evals\n\nTeams close the loop, shipping agent improvements validated against real production evidence\n\nYou'll build the product experiences that make this loop legible, and you'll build the agents that run it. This is not a role where you implement specs handed down. You'll own problems end-to-end: talking to customers, defining what to build, building it, and iterating until it's great.\nWhat You Will Accomplish\nJudgment Agent: Shape how the Judgment Agent runs large-scale investigations: parallel investigators working across thousands of production traces, each covering a different dimension (failure modes, tool errors, regressions, drift), merging results into one answer.\n\nVerification: Build the platform for verifying agent changes: hosted simulated environments for stateful agent evals, trajectory replay against changed agents, and monitors for unintended behavior changes.\n\nAgent investigation interfaces: Design how engineers understand what their agents did and why. Long traces, tool calls, decisions, failures. What does debugging look like when the \"program\" is a reasoning loop? How do you make a thousand-step trajectory legible in minutes?\n\nSwarm UX: A hundred parallel investigations is useless if engineers can't follow them. Design how humans watch a swarm work, redirect investigators that go down the wrong path, and consume findings without reading a hundred reports.\n\nThe improvement loop: Build the workflows that turn production trajectories into datasets, judges, and regression checks, so the path from \"found a problem\" to \"verified a fix\" feels like one motion.\n\nThe platform underneath: Workspaces, roles, permissions, billing, usage, and limits for teams running many agents across many environments.\n\nJudgment everywhere agents are built: An SDK and terminal-first experience so Claude Code, Codex, and OpenCode sessions can summon Judgment as a subagent mid-development.\n\nWhat You'll Bring\nExperience building and scaling end-to-end production systems, from data layer to UI\n\nStrong technical problem-solving skills, especially in fast-changing, ambiguous environments\n\nA builder and tinkerer's mindset with high agency - you find creative ways to overcome obstacles and ship\n\nHands-on experience building with LLMs or agents, or the drive to get there fast\n\nComfort working directly with customers to understand their needs and solve real-world problems\n\nExcellent communication skills - clear, direct, and persuasive across technical and non-technical audiences","description_format":"text","description_chars":2975,"description_truncated":false,"requirements":{"experience_years_min":null,"management_years_min":null,"team_size_min":null,"manages_managers":false,"education":null,"security_clearance":false,"languages":[]},"benefits":[],"hiring_locations":[],"hiring_excludes":[],"relocation_offered":false,"industries":["AI Evaluation & Observability"],"lifecycle":[{"event":"open","at":"2026-09-12T16:44:25Z"}],"liveness":{"score":24,"band":"cold","label":"Long shot","p_open":1,"p_active":0.428,"p_room":0.55,"age_days":46,"expected_fill_days":38,"reasons":["conf:3","win:tail"],"computed_at":"2026-09-29T05:45:00Z"},"pay":null,"html_url":"https://alion.io/job/judgmentlabs-product-engineer-agent","json_url":"https://alion.io/job/judgmentlabs-product-engineer-agent.json","meta":{"generated_at":"2026-09-30T02:21:00Z","cache_seconds":300,"methodology":"https://alion.io/methodology","terms":"https://alion.io/terms","contact":"https://alion.io/contact","api":"https://alion.io/developers","usage":{"tier":"crawler","counted_by":"address","units_charged":1,"used_today":1616,"day_limit":5000,"remaining_today":3384,"minute_limit":60,"resets_at":"2026-10-01T00:00:00Z"}}}