We are hiring an AI Engineer to serve as a technical contributor across multiple customer engagements. On each engagement you will set the technical direction, make the key modeling decisions, and stay hands-on throughout, acting as a senior technical point of contact with the customer, explaining trade-offs, managing expectations, and turning results into clear business recommendations and outcomes. Across engagements you will help build the reusable assets, patterns, and technical standards that raise the bar for a scaling AI engineering organization.
What You’ll Do
- Own the technical strategy on customer engagements, making the architecture and modeling decisions and being accountable for the results.
- Ramp quickly into unfamiliar domains and problem types, scoping the right approach for each customer's data, constraints, and timeline.
- Stay hands-on: build the pipelines, train and evaluate the models, run the experiments, and write the critical code.
- Set the technical bar and support other engineers through design reviews, mentorship,
and pairing.
- Act as a senior technical point of contact with customers, communicating progress,
risks, and results to both engineers and senior stakeholders, and managing expectations
through ambiguity.
- Engineer features and datasets from large-scale customer data, and integrate signals into ML model training and runs.
- Design, build, and evaluate LLM and agentic solutions including prompt and context design, retrieval, tool use, multi-step agent workflows, and orchestration.
- Design and run structured, parallel experiments that measure the gains from GenAI approaches over strong conventional ML methods.
- Own model and agent development end to end, including feature integration, hyperparameter optimization, and error analysis.
- Define and run the evaluation framework including task-level quality metrics, LLM and agent evaluation harnesses, and ablation studies.
- Establish the path to production: model and agent serving, latency and cost management, shadow-mode testing, A/B framework readiness, and guardrail metrics.
- Build reusable accelerators, reference architectures, and internal standards that carry from one engagement to the next.
- Deliver clear technical documentation and lead knowledge-transfer sessions so each customer's teams can operate and iterate independently after handoff.
Required Qualifications
- 10+ years in applied machine learning / data science, with deep hands-on experience in
building and shipping production ML systems across multiple problem domains.
- Hands-on experience with LLMs in production: prompt and context engineering, retrieval-augmented generation, embeddings, and reasoning about evaluation, latency, and cost.
- Hands-on experience building agentic systems including tool use and function calling, multi-step workflows, orchestration frameworks, and agent evaluation and guardrails.
- Experience with Amazon Bedrock or comparable managed LLM platforms.
- Strong communication skills, able to explain modeling decisions, trade-offs, and results
clearly to engineers, data scientists, and senior business stakeholders, and to manage
expectations through ambiguity.
- Customer-facing or stakeholder-facing experience: building trust, navigating competing priorities, and serving as a senior technical voice in high-stakes conversations.
- Comfort working across several engagements and customer contexts at once, switching domains without losing technical rigor.
- A track record of technical leadership through mentoring engineers, driving design
decisions, and setting standards.
- Strong track record taking models from experimentation to production, owning the offline-to-online validation story (evaluation metrics, ablations, shadow testing, A/B readiness).
- Deep, hands-on expertise in deep learning and embedding-based architectures with a major framework (PyTorch or TensorFlow).
- Strong feature engineering on large datasets using the modern data stack (Spark, SQL, distributed data lakes).
- Rigorous experimental methodology including hyperparameter optimization and a disciplined, hypothesis-driven approach to measuring true lift.
- Hands-on AWS experience across the ML lifecycle, and strong proficiency in Python.
- Experience in MLOps and LLMOps including model and prompt versioning, monitoring, evaluation pipelines, and reproducible training.
Preferred Qualifications
- Prior experience in a client-facing consulting or professional-services delivery environment.
- Advanced degree in Computer Science, Machine Learning, Statistics, or a related quantitative field.

