About the Role
This Staff Engineer role sits at the core of an applied agentic AI product that automates complex, multi-step workflows across desktop engineering tools used by hardware engineers every day. Reporting directly to the CTO, you will own the agent intelligence layer end-to-end and lead a small, focused team of AI engineers, a user researcher, and domain expert contractors. The work you do here directly determines how much real-world value the product delivers to enterprise customers.
What You'll Do
Lead development of the core agent intelligence layer that executes multi-step workflows across complex desktop engineering software.
Own the full product loop: define agent capabilities from user stories, build implementations, and benchmark against real workflows.
Drive agent task success rate by defining evaluation frameworks, establishing baselines, and iterating on completion metrics.
Set and enforce per-task token budgets and track cost per completed workflow to ensure commercial viability.
Build rigorous, reproducible evaluation infrastructure grounded in validated user stories.
Lead user story mapping and validation through engineer interviews and close collaboration with domain experts.
Translate validated user stories into testable evals, closing the loop between user research and agent benchmarking.
Own agent architecture decisions including tool-calling strategies, state management, error recovery, model routing, and context management.
Act as a player-coach: write production code, review designs, unblock the team, and raise engineering standards.
Collaborate cross-functionally with integrations, product, and customers during POCs to align agent behavior with real-world usage.
What We're Looking For
7+ years of software engineering experience, including at least 2 years building LLM-based agents that take real-world actions.
Deep experience designing LLM application architectures: model selection, context and window management, retrieval, and orchestration patterns.
Proven ability to build evaluation and benchmarking frameworks measuring task completion, cost efficiency, and failure modes.
Strong Python skills and hands-on familiarity with LLM tooling including function calling, tool APIs, observability and tracing, and evaluation frameworks.
Experience shipping AI or LLM tooling on top of proprietary engineering data or desktop engineering software, such as agents or MCP servers over CAD, PLM, or simulation platforms.
Technical leadership experience setting direction for small teams of 3 to 6 engineers while continuing to write and review production code.
Experience with desktop automation or programmatic control of applications such as COM or similar interfaces.
Domain familiarity with mechanical engineering, CAD, CAE, PLM, or adjacent engineering software industries.
Understanding of enterprise deployment constraints on locked-down corporate workstations.
Comfort operating in a fast-paced, high-intensity early-stage environment with significant customer demand.
Compensation & Benefits
Base salary range: $160,000 to $250,000 USD annually, plus equity. Visa sponsorship is not available for this role.
Location
On-site in San Francisco, California, United States.

