{"id":1741263,"url":"https://alion.io/job/uselemma-ai-aiml-engineer","title":"AI/ML Engineer","company":{"id":2539416,"name":"uselemma_ai","domain":"uselemma.ai","url":"https://alion.io/company/uselemma","size_band":"51-200","is_staffing_agency":false,"employer_type":"direct","is_intermediary":false,"listed_via":null,"ats_vendor":"Work at a Startup","truth_index":null},"role":"AI/ML","role_family":"AI/ML","seniority":null,"employment_type":"full_time","work_mode":"on_site","remote_scope":null,"remote_scope_basis":null,"remote_working_hours":null,"hiring_geo_confidence":"structured","locations":["San Francisco, United States"],"countries":["US"],"hiring_countries":[],"hiring_countries_total":0,"salary":{"min":120000,"max":200000,"currency":"USD","period":"year","gross":null,"usd_annual":200000},"salary_estimate":null,"experience_years_min":null,"visa_sponsorship":false,"relocation_package":false,"has_equity":true,"technologies":[{"name":"AI Agents","optional":false},{"name":"Embeddings","optional":false},{"name":"LLM","optional":false}],"status":"live","first_seen_at":"2026-10-03T03:24:54Z","employer_posted_date":"2026-10-03","last_verified_at":"2026-10-06T22:07:18Z","board_verified":true,"closed_at":null,"days_open":3,"trust":{"level":"ok","repost_count":null,"flags":[],"days_open":3},"description":"Lemma is production monitoring for AI agents. We catch the silent failures your observability tools and evals miss (think bad tool calls, lost context, and infinite loops) before your users find them.\nWhy this role exists\nAgents break silently. They call the wrong tool, forget what the user said three turns ago, and loop until someone pulls the plug. Soon they'll be responsible for the majority of the world’s economic work, and most teams won't even know when they fail.\nMaking agents reliable is the problem Lemma exists to solve. That means catching the unknown unknowns, the failures nobody thought to write an eval for, and closing the loop end to end so they get fixed, not just flagged. It's the foundation for building agents people can actually trust and the future of self-improving systems.\nThe hardest part of our product is deciding what counts as a failure.\nThere is no ground truth here, and no benchmark to climb. Every customer's agent is different, what \"wrong\" means changes from one to the next, and we have to get it right across production without anyone telling us what to look for. Being confidently wrong often costs us more trust than being right fifty times earns.\nThis role owns the intelligence in the loop: what we flag, how sure we are, and whether the fix we propose actually fixes it.\nWhat you’ll do\nOwn detection quality. Find the failures that don't look like failures: compliant but wrong, omissions, and patterns that only show up across thousands of traces\nTurn implicit signals into evidence. Rephrasing, abandonment, retries, and the other ways users tell you something broke without saying so\nBuild the evals for our own system. If we can't measure precision on a problem with no labels, we can't improve it\nMake patch generation trustworthy. Reproduce the failure, verify the fix, and know when to stay quiet instead of opening a bad PR\nKeep it affordable. LLM-as-judge on every event is easy. Doing it at a cost per event that doesn't eat the business is the actual job\nRead real customer traces every week. The best ideas here come from staring at production, not papers\nWhat we’re looking for\nHigh slope over years of experience. New grads and dropouts welcome\nA track record of shipping things people actually use\nHands-on with LLMs in production: evals, LLM-as-judge, embeddings, and knowing when a smaller model or no model is the right call\nReal research taste. You can tell a real improvement from noise, even when there's no ground truth to check against\nBonus: you were the customer once. You ran agents in production and got burned\nWho you’ll work with\nYou'll join a team of dropout founders and engineers from Amazon, Together, and Zoom. We've been founding operators at unicorns and at startups that went on to be acquired.\nOnsite in San Francisco. We don’t sponsor visas.\nWe keep this short on purpose. Target is an offer within two weeks of first contact.\nIntro call with a founder (30 minutes). What we're building, what you've built, whether the shape of the role actually fits what you want next.\nTake-home project and deep dive. We give you a real problem from Lemma. You work on it on your own time, then we walk through it together. The conversation matters more to us than the artifact.\nPaid work trial (onsite in-person). Real problem, real codebase, sitting with the team. You find out what working here actually feels like before you commit, which tells you more than anything we could say about it.\nMay ask for additional references.","description_format":"text","description_chars":3504,"description_truncated":false,"requirements":{"experience_years_min":null,"management_years_min":null,"team_size_min":null,"manages_managers":false,"education":null,"security_clearance":false,"languages":[]},"benefits":[],"hiring_locations":[],"hiring_excludes":[],"relocation_offered":false,"industries":["Artificial Intelligence","AI Agents"],"lifecycle":[{"event":"open","at":"2026-10-03T03:24:54Z"}],"visa":[],"liveness":{"score":86,"band":"hot","label":"Hiring now","p_open":1,"p_active":0.86,"p_room":1,"age_days":3,"expected_fill_days":34,"reasons":["conf:0","win:early"],"computed_at":"2026-10-06T05:45:30Z"},"pay":{"stated_usd_annual":200000,"is_top_pay":false},"html_url":"https://alion.io/job/uselemma-ai-aiml-engineer","json_url":"https://alion.io/job/uselemma-ai-aiml-engineer.json","meta":{"generated_at":"2026-10-06T22:14:14Z","cache_seconds":300,"methodology":"https://alion.io/methodology","terms":"https://alion.io/terms","contact":"https://alion.io/contact","api":"https://alion.io/developers","about":"Alion is a live layer of people, companies and AI agents: who they are, whether they are real and active right now, what they do and how to work with them, readable by people and by agents and paid per call.","catalog":"https://alion.io/catalog.json","usage":{"tier":"crawler","counted_by":"address","units_charged":1,"used_today":2455,"day_limit":5000,"remaining_today":2545,"minute_limit":60,"resets_at":"2026-10-07T00:00:00Z"}},"offers":[{"id":"company.slices","title":"One company in depth, by slice","status":"live","price":{"credits":0.02,"usd":0.002,"plus_per_slice":{"credits":0.05,"usd":0.005}},"unit":"per company, plus each slice with data","note":"the employer in depth","call":{"mcp_tool":"get_company","arguments":{"id":2539416},"rest":"https://alion.io/mcp/rest/get_company?id=2539416"},"human":"https://alion.io/catalog?offer=company.slices&for=job%2Fuselemma-ai-aiml-engineer"},{"id":"market.stats","title":"A market slice: pay, demand and time to fill","status":"live","price":{"credits":1,"usd":0.1},"unit":"per slice","note":"pay, demand and time to fill for this role and place","call":{"mcp_tool":"market_stats"},"human":"https://alion.io/catalog?offer=market.stats&for=job%2Fuselemma-ai-aiml-engineer"},{"id":"job.search","title":"Open jobs by role, technology, place, pay and visa","status":"live","price":{"credits":0.02,"usd":0.002},"unit":"per posting in a list","note":"similar open postings","call":{"mcp_tool":"search_jobs"},"human":"https://alion.io/catalog?offer=job.search&for=job%2Fuselemma-ai-aiml-engineer"},{"id":"company.verify","title":"Is this company real and active right now","status":"pilot","price":null,"unit":"per company","request":{"url":"https://alion.io/catalog/request","method":"POST","body":"{\"offer\": \"company.verify\", \"for\": \"job/uselemma-ai-aiml-engineer\", \"note\": \"what you need it for\"}"},"human":"https://alion.io/catalog?offer=company.verify&for=job%2Fuselemma-ai-aiml-engineer"}]}