We are looking for an AI Systems Engineer to design, build, and ship the Platform Agent and the agentic AI infrastructure that powers our AI-first productivity governance platform. This role goes well beyond prompt engineering. You will architect multi-step reasoning workflows, design the retrieval and grounding systems that give our AI accurate context over enterprise workforce data, build the evaluation harnesses that ensure our AI is reliable and trustworthy at scale, and own the LLM operations layer that keeps production AI systems healthy.
The Platform Agent is a strategically important product initiative. It needs to reason over complex, heterogeneous workforce telemetry, surface actionable insights to operations and finance leaders, and operate with the transparency and governance standards that regulated enterprise environments demand. You will be the person who makes that real from architecture to production.
The core responsibilities for the job include the following:
Platform Agent Architecture and Development:
- Lead the design and development of the Platform Agent, an AI system capable of reasoning over workforce telemetry, benchmarking data, and workflow signals to surface insights and recommendations for enterprise operations leaders.
- Architect multi-agent and agentic workflow systems using orchestration frameworks (LangGraph, CrewAI, or equivalents), including tool use, memory management, task decomposition, and multi-step reasoning chains.
- Design human-in-the-loop checkpoints and escalation patterns that ensure the Platform Agent operates with appropriate governance and auditability in enterprise environments.
- Own the end-to-end architecture of the agent from the prompt layer through retrieval, reasoning, tool execution, and output formatting, ensuring coherence, reliability, and latency standards are met.
Prompt Architecture and Context Engineering:
- Design and maintain the prompt architecture for the platform AI features system prompts, task prompts, tool-calling schemas, and structured output specifications across multiple LLM providers.
- Implement advanced prompting techniques, including Chain-of-Thought, ReAct, and structured decomposition, to handle the non-trivial reasoning tasks involved in productivity analytics and workforce intelligence.
- Engineer context pipelines that give the AI accurate, relevant, and well-scoped information from workforce telemetry, including token budget management, context prioritization, and grounding strategies.
- Build and maintain PromptOps practices for version control, change management, and regression monitoring for prompts in production.
RAG, Retrieval and Knowledge Architecture:
- Design and implement Retrieval-Augmented Generation (RAG) systems that allow the platform AI to reason accurately over enterprise-scale workforce data, benchmarks, process definitions, and historical patterns.
- Own the embedding, indexing, and retrieval pipeline, including vector database selection and optimization (Pinecone, Weaviate, Milvus, or equivalent), chunking strategies, and retrieval quality.
- Build domain-specific retrieval strategies suited to the nature of ProHance data, time-series telemetry, aggregated workforce metrics, and multi-level benchmarking datasets, not generic document retrieval.
- Ensure retrieved context is accurate, traceable, and grounded, enabling AI outputs that enterprise buyers can interrogate and trust.
Evaluation, Quality, and LLM Operations:
- Build and maintain evaluation harnesses' golden datasets, automated eval pipelines, and regression suites to measure AI output quality across hallucination, groundedness, relevance, and reasoning accuracy.
- Develop adversarial test cases and red-team scenarios appropriate to the ProHance domain: incorrect benchmark claims, misleading productivity narratives, and misattributed anomalies.
- Own LLM operations in production: token cost tracking, latency monitoring, model drift detection, and observability tooling (LangSmith, Helicone, or equivalent).
- Establish the quality bar and release criteria for AI features, ensuring that what ships to enterprise customers meets the trust and reliability standards the market demands.
Cross-functional Collaboration:
- Partner with the Senior Data Scientist to consume intelligence layer outputs, ensuring that model-produced signals, benchmarks, and anomalies are correctly retrieved and reasoned over by the agent.
- Collaborate with the AI product team to translate product requirements into agentic workflows and AI feature specifications.
- Work with engineering to ensure AI systems are integrated cleanly into the platform with appropriate security, access controls, and privacy handling for enterprise workforce data.
- Contribute to external and internal documentation that explains how Platform AI works, supporting the trust architecture that enterprise buyers require.
Requirements:
- Education: B. Tech / M. Tech in Computer Science or a related engineering field.
- Experience: 3-5 years in software engineering or AI product development, with at least 2 years of hands-on production LLM application development.
- Agentic systems: Demonstrable experience designing and shipping agentic or multi-step AI systems, not just single-turn prompt-response applications.
- LLM frameworks: Strong proficiency with LangChain, LangGraph, LlamaIndex, or equivalent orchestration frameworks; ability to go beyond framework defaults when needed.
- RAG and retrieval: Hands-on experience with RAG architecture, vector databases, embedding pipelines, and retrieval quality optimization.
- Evaluation discipline: Experience building eval harnesses, golden datasets, and automated quality pipelines, not just manual spot-checking.
- Python: Expert-level Python for API orchestration, data processing, and AI pipeline development.
- Production mindset: Track record of shipping AI systems that work reliably in production, not just demos or prototypes.
Preferred Qualifications:
- Experience with enterprise-grade AI governance requirements, auditability, explainability, access control, and data privacy in regulated environments.
- Familiarity with workforce analytics, operational telemetry, or time-series data as a retrieval and reasoning domain.
- Experience fine-tuning or adapting open-source LLMs (LoRA, QLoRA) for domain-specific tasks.
- Contributions to open-source LLM tooling, published research on LLM evaluation, or demonstrated thought leadership in AI systems engineering.
- Experience with multimodal reasoning or structured data (tables, metrics) as first-class context for LLMs.

