This role sets the standard for how Gen AI gets built into Zenskar's product. You're accountable for whether the Gen AI features Zenskar ships are reliable, safe, and cost-sane in production, not for whether any single feature gets built. Concretely, this role exists so that no team ever ships an LLM integration that hallucinates a wrong invoice amount or burns unbounded API costs, because there was no standard pattern to build against.
This is an IC role with org-wide influence, not team ownership. You'll code on the critical paths where it matters most, but your leverage comes from setting patterns other teams build on top of, not from personally shipping every Gen AI feature. Please think 1-3 years for how Gen AI gets integrated across the product, not one team's roadmap. What this role is not: foundation model training, deep RAG/retrieval specialization as a primary mandate, fine-tuning specialization, or MLOps infrastructure ownership. Those are adjacent lanes. This is LLM application engineering: building real product features on top of frontier model APIs and making them production-reliable.
Responsibilities:
- Turn ambiguous product asks into working LLM-based features: prompt and context design, tool calling, orchestration across model calls, and structured output handling.
- Design for the failure modes non-deterministic systems actually have: hallucination, drift, and silent wrongness. Build guardrails, and know the difference between demo-grade and production-grade because you've filled the gap between them.
- Make sure an LLM is never the unverified source of truth for a financial decision. Design the verification or human-in-the-loop layer wherever money is actually at stake.
- Pick the right model or vendor for the job (cost, quality, latency, data policy) and avoid lock-in. Once chosen, run it efficiently: token economics, caching, batching, and per-request cost tracking.
- Set the technical standard other teams build their Gen AI features against, and get multiple teams actually to adopt it without having direct authority over them.
- Build and own the eval harness that catches quality regressions before they ship, not ad hoc spot-checks.
- Know when not to use an LLM at all, where a deterministic system would be cheaper, more reliable, or simply correct. Name and quantify technical debt in Gen AI systems (prompt sprawl, eval gaps, untracked cost, brittle integrations) and negotiate time to address it.
- Shape who joins the team by holding a technical bar in interview loops, without owning headcount decisions.
Requirements:
- 8+ years of software engineering experience, including meaningful experience shipping AI-powered products to production.
- A real, shipped-at-scale system in your track record doesn't have to be Gen AI. If it's ML work, it shipped to real users, not fine-tuning or research that stayed in a notebook.
- Deep understanding of LLM application architecture: tool use, structured outputs, retrieval, orchestration, and where these actually break in production.
- Experience building agentic systems that do multi-step reasoning and interact with tools or business systems.
- A track record of treating prompts as versioned, testable, deployable artifacts, not strings scattered through the codebase.
- Strong RAG fundamentals: retrieval quality, chunking, embeddings, evaluation, and knowledge system design.
- Experience designing evaluation frameworks and regression pipelines for AI systems.
- Strong grasp of AI observability: tracing, monitoring, feedback loops, and production debugging.
- A track record of getting more than one team to adopt a pattern or standard you set, without owning those teams.
- Judgment for when not to use an LLM at all.
- Strong backend engineering skills and enough frontend ability to own an AI experience end to end.
- Can clearly walk through an AI system you've built two levels below the pitch: what failed, what you learned, and how you improved reliability over time.
Good-to-Have:
- Memory architectures: episodic memory, procedural memory, retrieval systems, and knowledge stores.
- Agent orchestration frameworks (LangGraph, Pydantic AI, OpenAI Agents SDK, CrewAI, or a custom runtime).
- Long-running autonomous workflows and event-driven agent systems.
- Voice, multimodal, or real-time AI systems.
- Fine-tuning experience and understanding of model internals beyond API consumption.
- Open-source model deployment and inference infrastructure.
- Experience with financial systems, billing platforms, revenue operations, accounting, or fintech.
- Familiarity with MCP, tool ecosystems, and AI platform architecture.
- Startup experience and comfort operating with high ownership and no formal authority.

