368,746open jobs
9,444companies
47,506added this week
Browse all
Salary
$77k – $138k per year (Estimated)
Location
Remote (Bulgaria, Cyprus, Czech Republic, Hungary, Italy, Poland, Romania, Serbia, Spain, Ukraine, Albania, Estonia, Greece, Lithuania, Moldova, North Macedonia, Portugal, Slovakia, Slovenia)
Seniority
Architect · 6+ years exp
Employment
Contractor
Overview
Company
Impact
Profile match
Neurons Lab helps financial institutions move from AI-curious to AI-enabled. We deliver AI training programs and custom AI agents designed for regulated environments — from pilot to production, at scale.

About the project (description, duration, stage)

The client is the largest US network of in-home veterinary hospice and end-of-life care. A major US private-equity sponsor drives the AI program and plans more projects across its portfolio.

We built a real-time voice copilot for their Veterinary Care Coordinators (VCCs). The copilot listens to live calls with pet families. It extracts appointment and clinical fields while the call runs. It fills the client's scheduling system through a Chrome extension. A second workstream, the Vet Visit Copilot, sends each vet an AI pre-visit briefing by email (Amazon SES).

Next is the production phase.

Stage: production SOW in executive alignment; start expected September 2026.

Duration: multi-month, with strong extension probability. 0.5 FTE minimum; ramp toward 1.0 FTE as production scales.

Why the role is open: the current architect moves to another strategic build. He stays at 0.15-0.2 FTE for supervision and knowledge transfer during ramp-up, so the new architect gets a structured handover.

Objective

  • Own the technical architecture and delivery of the voice copilot from validated PoC to production

  • Hit the bar this client tests against: latency, accuracy, concurrency, and cost

  • Keep expectations aligned: production polish is in scope now; protect the team from silent scope creep

  • Transfer knowledge continuously to the client's team and Neurons Lab engineers

Areas of Responsibility

Technical architecture & hands-on implementation

  • Own the full pipeline: streaming speech-to-text, LLM field extraction, Chrome-extension delivery, and AWS infrastructure

  • Drive latency work: cut P95 from ~6s toward ~2s; remove post-processing corner cases (occasional ~1min lag on one field type)

  • Run model A/B tests (current pair: Claude Haiku vs GPT Luna) with golden-set evaluation for phonetic name and email accuracy

  • Own evaluation and cost: Langfuse traces, accuracy dashboards, real per-call cost from live calls, and an optimization plan

  • Harden for production: 5-10+ concurrent calls, strict data isolation between users, monitoring, alerting, and safe rollback

  • Ship epics end to end (example: the SES email briefing service); always keep a demo fallback so a live session never fails

Working with client stakeholders

  • Front technical discussions with a meticulous client; VCCs test edge cases and expect production quality

  • Present concrete system behavior, with numbers - this account rewards evidence, not slides

  • Hold the scope line: tie every feedback item to the SOW; route roadmap items (learning loop, persistent memory) to future phases

  • Keep internal discussions internal; all client-facing materials pass ADM review before sending

Team & knowledge

  • Lead the AI Engineer and the pod: set tasks, review output, unblock fast

  • Absorb the handover from the outgoing architect (0.15-0.2 FTE supervision window) and become independent fast

  • Run knowledge-transfer sessions; the project must have no single point of failure

  • Support the production SOW with estimates and architecture options when the account team asks

Skills

  • Real-time voice pipelines: streaming STT, turn handling, low-latency LLM inference - hands-on

  • LLM engineering: prompt engineering, structured extraction, guardrails, model A/B evaluation

  • Observability and evals: Langfuse or similar; golden datasets; latency, accuracy, and cost dashboards

  • AWS: Bedrock, serverless patterns, SES; token economics and per-call cost engineering

  • Full-stack pragmatism: strong Python; enough TypeScript / Chrome-extension knowledge to own the integration

  • Clear spoken and written English for demanding US executives

Knowledge

  • Contact-center / agent-assist patterns and metrics (handle time, cost per call, concurrency)

  • Production LLM operations: load testing, data isolation, incident handling

  • Nice to have: empathy-sensitive domains (healthcare, veterinary, insurance) and PE-sponsored rollouts

Experience

Key characteristics (screen for all four):

  • Voice AI in production - mandatory. Shipped at least one real-time voice or speech product to real users (agent assist, voice bot, live transcription copilot). Candidates will demo real artifacts at the interview.

  • 6+ years hands-on AI/ML engineering, with strong recent LLM production practice

  • Latency and reliability record. Can show measured P95 reductions and concurrency fixes on a live system

  • Consulting / client-facing seniority. Calm and precise under detailed UAT scrutiny; manages expectations well

Nice to have:

  • Chrome extension delivery; telephony / streaming stacks (Amazon Connect, Twilio, LiveKit)

  • Langfuse in production

  • US client experience with Eastern-time overlap

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
368,746 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
In your city
$110k – $131k per year • Remote • Full-Time • 10+ years exp • Bachelor's Degree
DevOps
AWS
Incident Management
VMWare
Apply
In office • Part-Time • 2+ years exp • Bachelor's Degree • Ness Ziona
C#
Java
AI/ML
Copilot
Cursor
DevOps
CI/CD
GitHub
Apply
$105k – $252k per year • Remote • Full-Time • 18+ years exp • Bachelor's Degree
Python
Java
Java
Gradle
DevOps
Ansible
AWS
CI/CD
CloudFormation
Configuration Management
Docker
GitHub Actions
GitLab CI
Helm
Jenkins
Kubernetes
Platform Engineering
Terraform
GitHub
GitLab
Cybersecurity
Sonatype Nexus IQ
Management
Confluence
Jira
Apply
$54k – $175k per year (Estimated) • Remote • Full-Time
JavaScript
Node JS
TypeScript
Frontend
React.js
Sass
DevOps
AWS
CI/CD
Docker
Kubernetes
Rest API
Apply
$35k – $113k per year (Estimated) • Remote • Full-Time
JavaScript
Node JS
TypeScript
Frontend
React.js
Sass
DevOps
AWS
CI/CD
Docker
Kubernetes
Rest API
Apply
$67k – $129k per year (Estimated) • Remote/Hybrid • Contractor • 6+ years exp
Python
AI/ML
Langfuse
LLM
LLM Evaluation
LLM Guardrails
Structured Outputs
DevOps
AWS
Robotics
Imitation Learning
Apply
$52k – $92k per year (Estimated) • Remote • Full-Time • 4+ years exp
Python
SQL
Databases
OpenSearch
pgvector
Pinecone
PostgreSQL
AI/ML
Composio
LLM
RAG
Unstructured.io
Model Context Protocol
DevOps
AWS
GCP
AWS Step Functions
Cybersecurity
GDPR
Management
Google Workspace
Slack
Apply
$39k – $85k per year (Estimated) • Remote • Full-Time
SQL
AI/ML
Knowledge Distillation
LLM
Human-in-the-Loop
Knowledge Graph
Cybersecurity
GDPR
Management
Slack
Apply
$74k – $129k per year (Estimated) • Remote • Full-Time • Madrid
AI/ML
AWS Bedrock
Claude
Claude Code
Cursor
LiteLLM
LLM
OpenRouter
LLM Guardrails
OpenAI Codex
AI Agents
Model Context Protocol
DevOps
AWS
Apply
$62k – $106k per year (Estimated) • Remote • Part-Time • 5+ years exp • Madrid
Databases
Neo4j
AI/ML
AutoGen
LangChain
LLM
Prompt Engineering
AI Agents
EU AI Act
Knowledge Graph
Apply
See all jobs
This is one of many
368,746 more open roles from verified company boards, updated every day.