This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Senior ML / AI Engineer based in United States.
This role is focused on building the AI/ML platform that enables high-impact, production-grade AI initiatives in healthcare.
You will design and operate agentic systems capable of handling sensitive, financially consequential decisions with strong governance and reliability.
Your work will span LLM applications, RAG pipelines, self-hosted model inference, MLOps, observability, and cloud infrastructure.
You will help build systems that are not only capable, but also grounded, auditable, secure, and resilient in production.
The role combines deep AI engineering with hands-on software engineering and infrastructure ownership across AWS and GPU-based environments.
You will also help extend the shared platform to new AI use cases while mentoring engineers and establishing strong development practices.
This is an opportunity to work on genuinely challenging AI problems where correctness, transparency, and responsible deployment are essential.
Accountabilities:
- Design, build, and deploy production-grade agentic AI systems using dynamic tool calling, structured outputs, multi-agent orchestration, verification, grounding checks, LLM-based evaluation, and human-in-the-loop workflows.
- Develop and maintain RAG pipelines covering document ingestion, chunking, embeddings, hybrid retrieval, reranking, retrieval-quality evaluation, and security controls.
- Implement strict access controls across member and client data while protecting AI systems against prompt injection, data poisoning, and other retrieval or model-level risks.
- Operate self-hosted LLM and VLM inference services on AWS GPU infrastructure, optimizing throughput, latency, concurrency, and cost through techniques such as continuous batching, prefix caching, quantization, and KV-cache tuning.
- Own the MLOps and AI governance layer, including offline evaluation frameworks, golden datasets, quality scoring, canary and gray releases, automated rollback, budget controls, and production monitoring.
- Build reliable AI services with durable state, circuit breakers, retries, fail-closed defaults, PHI-safe logging and redaction, RBAC, and horizontal scalability.
- Develop and maintain AWS infrastructure using services such as EKS, RDS/pgvector, SQS, S3/KMS, and Cognito, with infrastructure managed through Terraform.
- Establish comprehensive observability using OpenTelemetry, metrics, traces, and immutable audit records.
- Take new AI initiatives from initial problem definition through architecture, evaluation, deployment, governance, and ongoing production operation.
- Mentor engineers and promote strong standards for typed, tested, maintainable, and production-ready code.
- Participate in design reviews and make pragmatic decisions balancing model autonomy, determinism, reliability, and safety for high-stakes AI applications.
- Investigate and resolve production incidents while continuously improving system resilience, performance, and operational maturity.
- 5+ years of experience building and deploying ML/AI systems in production environments, with hands-on LLM application development within the last 1-2 years.
- Strong Python expertise, including typed, tested, production-grade development and solid software engineering fundamentals.
- Comfortable working across asynchronous web services, data layers, cloud infrastructure, and production systems.
- Practical expertise with LLM application patterns including prompting, structured/function calling, RAG, embeddings, vector search, retrieval optimization, agent and tool-use loops, and AI evaluation.
- Strong understanding of AI quality and risk management, including grounding, hallucination mitigation, evaluation datasets, guardrails, and reliability strategies.
- Hands-on AWS or equivalent cloud and MLOps experience, including Kubernetes, containers, Terraform, CI/CD, observability, and model-serving performance optimization.
- Proven track record of owning production reliability, including durable state, failure handling, scalability, incident response, and debugging.
- Experience self-hosting and optimizing open-weight models using technologies such as vLLM, TGI, or comparable inference frameworks on GPU infrastructure.
- Experience serving embeddings and reranking models, including technologies such as BGE or Infinity.
- Experience with LangGraph, LangChain, or comparable agent orchestration frameworks and multi-agent architectures.
- Experience working with regulated data and security frameworks such as HIPAA, SOC 2, or ISO 27001, including PHI/PII handling, RBAC, and audit trails.
- Healthcare, insurance, or claims experience is strongly advantageous, particularly familiarity with X12 835, EOBs, benefits, or related financial and healthcare data.
- Experience with PostgreSQL/pgvector, MongoDB or DocumentDB, SQS, and Cognito/OIDC.
- Familiarity with OpenTelemetry, Prometheus, and Grafana.
- Frontend development experience with React and TypeScript to support internal AI builders, consoles, or tooling.
- Experience with Amazon Bedrock, multi-provider model abstractions, or on-premises/air-gapped AI deployments is a plus.
- Strong written communication, sound technical judgment, and the ability to make practical decisions in ambiguous situations.
- Ability to work collaboratively while taking significant ownership of complex technical initiatives.
- Must be authorized to work for any employer in the United States; visa sponsorship is not available.
- Base salary range of $150,000-$210,000, depending on experience.
- 100% employer-paid medical, dental, and vision coverage for employees and dependents.
- 401(k) retirement plan with company matching contributions of up to 4%.
- Opportunity to work on high-impact AI systems in healthcare where reliability, governance, and correctness are critical.
- Hands-on exposure to advanced technologies including LLMs, VLMs, agentic AI, RAG, GPU inference, AWS, Kubernetes, and MLOps.
- Opportunity to build foundational AI infrastructure used across multiple initiatives rather than working on isolated prototypes.
- Strong emphasis on AI governance, evaluation, observability, security, and responsible deployment.
- Collaborative, mission-oriented environment with opportunities to influence architecture, engineering practices, and future AI initiatives.
- Opportunity to contribute to technology designed to improve healthcare access, affordability, and outcomes.
Requirements:
Benefits:

