{"id":847830,"url":"https://alion.io/job/metaforms-senior-ai-engineer","title":"Senior AI Engineer","company":{"id":680192,"name":"Metaforms","domain":"metaforms.ai","url":"https://alion.io/company/metaforms","size_band":"51-200","is_staffing_agency":false,"employer_type":"direct","is_intermediary":false,"listed_via":null,"ats_vendor":"Ashby","truth_index":null},"role":"AI/ML","role_family":"AI/ML","seniority":"senior","employment_type":"full_time","work_mode":"on_site","remote_scope":null,"remote_scope_basis":null,"remote_working_hours":null,"hiring_geo_confidence":"structured","locations":["Bengaluru, India"],"countries":["IN"],"hiring_countries":[],"hiring_countries_total":0,"salary":null,"salary_estimate":{"min_usd":23000,"max_usd":48000,"period":"year","method":"role_seniority_country_remote_cell","sample_n":29},"experience_years_min":4,"visa_sponsorship":false,"relocation_package":false,"has_equity":false,"technologies":[{"name":"AI Agents","optional":false},{"name":"Anthropic","optional":false},{"name":"Computer Use","optional":false},{"name":"Context Engineering","optional":false},{"name":"Function Calling","optional":false},{"name":"Gemini","optional":false},{"name":"Human-in-the-Loop","optional":false},{"name":"LLM","optional":false},{"name":"OpenAI","optional":false},{"name":"Python","optional":false},{"name":"Tool Use","optional":false},{"name":"TypeScript","optional":false},{"name":"Braintrust","optional":true},{"name":"Browser Agents","optional":true},{"name":"Langfuse","optional":true},{"name":"LangSmith","optional":true},{"name":"Promptfoo","optional":true}],"status":"live","first_seen_at":"2026-06-24T04:09:49Z","employer_posted_date":"2026-06-24","last_verified_at":"2026-09-29T01:28:14Z","board_verified":true,"closed_at":null,"days_open":97,"trust":{"level":"stale","repost_count":0,"flags":["stale"],"days_open":97},"description":"About Metaforms\nMarket research runs on 30-year-old survey platforms and armies of specialists hand-coding questionnaires in proprietary languages. Metaforms is the agent layer that does that work. Every survey is a program - full of skip logic, piping, quotas, and loops - and a single wrong number in a client report is unrecoverable. Our AI agents write production survey code, QA live deployments, process and clean large structured datasets, configure analysis, and generate client-ready reports, so agencies like Dynata, Savanta, and Borderless Access ship more projects with far less friction.\n1,000+ surveys processed monthly\n\nServing Fortune 500 companies across the globe\n\nRapid month-over-month growth\n\nWe’re Series A funded and scaling fast, aggressively growing our AI engineering team to build the next generation of production-grade AI agent systems.\nThe Role\nWe’re hiring a Senior AI Engineer to own the design, development, and continuous improvement of the AI agent systems that power modern research operations.\nThis is a high-ownership, high-impact role at the intersection of applied AI and systems engineering. You’ll work on genuinely hard problems: agent reliability at scale, long-context handling, cascading error mitigation, and evaluation infrastructure - like codegen agents that write in proprietary DSLs, computer-use agents that QA live deployments, data agents that clean tabular exports and configure multi-step analysis, and evals for outputs where “correct” is genuinely ambiguous. And you’ll do it on a team that ships fast and treats quality as non-negotiable.\nWhat You’ll Own\nAgent Harness and Architecture\nOwn the agent harness our production agents run on - the loop where agents plan, use tools, check their work, and recover from failures\n\nLead research and implementation for long-context handling and cascading-error challenges in multi-step agent pipelines\n\nDrive context engineering strategy and experimentation frameworks across the team\n\nEvaluation and Production Monitoring\nDefine structured rubrics for evaluating AI outputs on nuanced, ambiguous research tasks\n\nBuild continuous monitoring, tracing, and failure-mode analysis for agents in production - including the loop that turns production failures into test cases\n\nCreate tooling that lets domain experts refine and evolve the skill files, eval sets, and knowledge bases our agents consume\n\nReliability for High-Stakes Outputs\nBuild eval suites - regression sets, golden datasets, LLM-as-judge pipelines - that catch regressions before deploy\n\nDevelop evaluation datasets for DSLs, structured data transforms, and computed outputs to systematically find and close model weaknesses\n\nDesign human-in-the-loop and review workflows for outputs where a single wrong number in a client report is unrecoverable\n\nWhat We’re Looking For\nMust-Have\nBuilt and operated agentic systems in production - multi-step pipelines, tool use, codegen, computer-use, or data and reporting agents - not just prototypes\n\n4+ years of engineering experience, with at least 1 year focused on LLM/agent systems in production\n\nDeep hands-on experience with frontier model APIs (Anthropic, OpenAI, Gemini), evaluation frameworks, and AI system optimization\n\nStrong Python skills; Go or TypeScript a plus\n\nSolid grasp of context engineering and evaluation methodology\n\nStrong instincts for debugging complex, non-deterministic system failures\n\nHigh ownership: you drive problems to resolution independently and pull others in when it matters\n\nNice to Have\nExperience with LLM observability and eval tooling (Braintrust, Langfuse, LangSmith, Weave, promptfoo, or in-house equivalents)\n\nBackground in semantic parsing, DSLs, or structured-output generation\n\nPrior work on computer-use or browser agents\n\nExperience with human-in-the-loop agent workflows where proposals are reviewed before apply, or agents over large structured datasets\n\nWhy Metaforms\nWork at the frontier of production AI: systems handling 1,000+ research projects a month, with the reliability bar that implies\n\nA small, senior team where your decisions carry real architectural weight\n\nZero-bureaucracy culture: high autonomy, fast feedback loops, direct access to leadership\n\nWell-funded and financially stable, with a clear roadmap and the runway to execute on it\n\nBenefits\nFull family health insurance\n\n$1,000 USD annual learning and development budget\n\nDedicated mentor and coaching support\n\nFree snacks and dinner at the office","description_format":"text","description_chars":4476,"description_truncated":false,"requirements":{"experience_years_min":4,"management_years_min":null,"team_size_min":null,"manages_managers":false,"education":null,"security_clearance":false,"languages":[]},"benefits":["Health insurance"],"hiring_locations":[],"hiring_excludes":[],"relocation_offered":false,"industries":["Market Research","AI Agents"],"lifecycle":[{"event":"open","at":"2026-09-13T01:44:18Z"}],"liveness":{"score":11,"band":"cold","label":"Long shot","p_open":1,"p_active":0.396,"p_room":0.28,"age_days":96,"expected_fill_days":42,"reasons":["conf:1","win:tail","crowd:"],"computed_at":"2026-09-28T05:45:00Z"},"pay":null,"html_url":"https://alion.io/job/metaforms-senior-ai-engineer","json_url":"https://alion.io/job/metaforms-senior-ai-engineer.json","meta":{"generated_at":"2026-09-29T05:08:37Z","cache_seconds":300,"methodology":"https://alion.io/methodology","terms":"https://alion.io/terms","contact":"https://alion.io/contact","api":"https://alion.io/developers","usage":{"tier":"crawler","counted_by":"address","units_charged":1,"used_today":4910,"day_limit":5000,"remaining_today":90,"minute_limit":60,"resets_at":"2026-09-30T00:00:00Z"}}}