We are seeking a Senior AI Engineer to take end-to-end ownership of the core AI systems that power Avaamo's voice and generative AI products. This is a high-ownership, AI-native role: you will be the directly responsible owner of our Automatic Speech Recognition (ASR) improvement roadmap, the owner of our Retrieval-Augmented Generation (RAG) pipelines, and the lead for building our OpenRouter-style LLM gateway that routes traffic intelligently across model providers. You will own these systems from roadmap to production design, implementation, evaluation, cost, and quality. Just as important as what you build is how you build: LLMs are your default building material, and AI tools are your default way of working, from AI-assisted coding and agentic automation to eval-driven development where every change ships behind a measurable quality gate.
Responsibilities:
- Own Avaamo's ASR improvement roadmap end-to-end accuracy (WER), latency, and robustness across accents, domains, and noisy real-world audio, including 8 kHz telephony.
- You are the single point of accountability for speech recognition quality.
- Lead fine-tuning and domain adaptation of state-of-the-art speech models (e. g., Whisper, Qwen3.8 Qwen-Audio/Qwen-Omni, Conformer/RNN-T, wav2vec 2.0) for enterprise vocabulary product names, alphanumeric IDs, and industry-specific terms.
- Own recognition pipeline capabilities: contextual biasing, custom vocabulary, hotword boosting, punctuation, inverse text normalisation, and entity formatting.
- Own streaming, real-time recognition behaviour endpointing, partial hypotheses, and barge-in within tight latency budgets for live voice conversations.
- Build and own the ASR evaluation framework: curated test sets, WER and entity-level metrics, regression tracking, and A/B comparisons and make the data-backed call on engine selection between in-house and third-party ASR.
- Own Avaamo's RAG pipelines end-to-end: ingestion, chunking, embeddings, indexing, hybrid retrieval, reranking, and grounded generation with direct accountability for answer quality in production.
- Drive retrieval quality improvements: chunking strategy, embedding model selection and fine-tuning, metadata filtering, and rerankers.
- Own hallucination reduction: grounding, citations, answerability detection, and guardrails, with faithfulness and relevance metrics that you define and track.
- Build and own automated, AI-native RAG evaluation for retrieval precision/recall, LLM-as-judge harnesses and close the loop with production feedback.
- Own the cost and latency envelope of the RAG stack: caching, index tuning, prompt and context optimisation, and model right-sizing.
- Lead the design and build-out of Avaamo's unified LLM gateway (OpenRouter-style) fronting multiple providers: OpenAI, Anthropic, Google, Azure, and open-weight models served via vLLM or similar behind one consistent API. This is your project to architect and deliver.
- Own the routing policy: capability-, cost-, latency-, and availability-aware routing with automatic failover, retries, and load balancing across providers.
- Own gateway platform features: request/response normalisation, streaming, token accounting and cost attribution, rate limiting, and semantic caching.
- Own LLM observability: tracing, logging, quality monitoring, and drift/regression detection across model versions.
- Own the model strategy: continuously evaluate new frontier and open-weight models and drive routing and model-switching decisions with data.
- Own ML deployment end-to-end: containerised packaging, CI/CD for models and services, model registry and versioning, canary/shadow rollouts, and safe rollback across the ASR, RAG, and LLM gateway stacks.
- Own inference at scale: GPU serving and optimisation (dynamic batching, quantisation, KV-cache management), autoscaling policies, and throughput/latency SLOs for real-time voice and chat traffic.
- Own scaling and reliability: capacity planning, load and soak testing, high-availability and multi-region deployment, monitoring and alerting, and cost-per-conversation optimisation; you are on the hook for how your systems behave in production.
- Build internal AI leverage: create the agents, eval harnesses, and AI-assisted tooling that multiply the whole team's output and champion AI-native practices across the engineering org.
- Mentor engineers and raise the team's bar on AI-native engineering, evaluation rigour, and ownership culture.
Requirements:
- Bachelor's or Master's degree in Computer Science, Machine Learning, or a related field.
- 5-8 years of experience in machine learning engineering or applied ML roles.
- Proven experience building and operating LLM-based systems in production (RAG, agents, or LLM serving/routing); this is a must-have for the role.
- Strong programming skills in Python and hands-on experience with PyTorch for model fine-tuning and inference.
- Experience with cloud platforms such as AWS, GCP, or Azure, and with containerised services (Docker, Kubernetes).
- Demonstrated end-to-end ownership of a production AI/ML system from design through deployment, scaling, and live operations.
- An AI-native way of working: daily use of AI coding assistants, agentic workflows, and eval-driven development is expected; be prepared to show us how you build with AI.

