368,634open jobs
9,437companies
50,578added this week
Browse all
Salary
$124k – $259k per year (Estimated)
Location
In office (Palo Alto)
Seniority
Senior · 5+ years exp
Employment
Full-Time
Overview
Company
Impact
Profile match
BrightAI deploys autonomous AI devices, edge sensors, and foundation models to transform critical infrastructure from reactive to proactive. Maximize asset life and eliminate catastrophic downtime.

Senior AI Engineer - Edge Dialog Systems

BrightAI is a high-growth Physical AI company transforming how businesses interact with the physical world through intelligent automation. Our AI platform processes visual, spatial, and temporal data from billions of real-world events-captured across edge devices, mobile sensors, and cloud infrastructure-to enable intelligent decision-making at scale.

We are now hiring a Sr. AI Engineer - Edge Dialog Systems to own and evolve the on-device conversational AI that powers our industrial safety wearable. The assistant guides field technicians through safety-critical procedures by voice, and it runs on the device itself, under hard latency, memory, and thermal budgets, with deterministic safeguards that take precedence over model output.

This is a systems role rather than a prompt-and-retrieve role. The competencies of a strong LLM and RAG engineer are needed for the position, but the work itself is dialog systems engineering at the edge, where most conversational turns are resolved by deterministic and embedding-based methods, and the language model is the last resort rather than the first move.

You will work at the intersection of natural language understanding (NLU), small language models (SLM), and embedded software, building a system in which a wrong answer is a safety concern and not merely a quality issue.

Responsibilities

  • Own the on-device dialog pipeline end to end: intent routing, hybrid intent classification (pattern matching combined with embedding similarity and out-of-domain detection), text normalization for noisy speech input, and the multi-step guided-procedure engine.
  • Maintain and extend the deterministic safety layer that wraps the language model-confirmation and echo-back gating, criticality tagging, negation handling-so that a misheard answer on a safety-critical step cannot pass silently.
  • Run SLM inference on-device under memory, computational complexity, and latency budgets, and reduce the per-turn inference cost through model selection, quantization, and runtime optimization.
  • Preserve and extend the zero-shot configuration model, in which new device commands and customer procedures are authored as data rather than released as code, so that a new customer can be onboarded in hours rather than weeks.
  • Coordinate the device deployment pipeline with the edge team 
  • Maintain the API contract with on-device voice pipeline & its speech-to-text (STT) stack.
  • Define and run on-device benchmarks: latency, accuracy, and false-accept/reject rates on safety-critical steps; use measurements to drive engineering decisions.
  • Build and maintain golden datasets and a non-regression suite, and use them as the release gate as the command and procedure catalogs grow.
  • Lead the migration from zero-shot to fine-tuned on-device models in order to reduce latency, without reintroducing a per-customer retraining burden.
  • Collaborate with product, firmware, and cloud teams, and bring new capabilities online, including additional languages, device commands, and  guided workflows.

Educational Background

  • Trained in AI, Machine Learning, Electrical/Computer Engineering, or a related field, with specialization in NLP, speech, or deep learning; or equivalent industry experience delivering production conversational AI systems.
  • Applied background in NLU, dialog systems, or on-device machine learning.

Required Skills & Expertise

LLM and retrieval foundation - the baseline for this role

  • 5+ years of experience in ML/AI, with a strong focus on NLP, LLMs, or conversational AI.
  • Strong applied experience with LLMs: prompting, structured output, tool and function calling, evaluation, and retrieval-augmented generation (RAG) - together with the judgment to recognize when a model should not be used at all.
  • Solid command of embeddings and semantic similarity (e.g., cosine similarity, centroid versus maximum-similarity strategies, threshold tuning, and out-of-domain detection).
  • Strong Python with the ability to write clean, tested, reviewable code. Fluent with pytest, and Git and has CI discipline.

Edge dialog systems engineering - what this role additionally requires

  • Strong experience building edge conversational systems, including multi-turn dialog/state management and efficient intent/NLU pipelines using local-first, cheap-to-expensive inference strategies.
  • Disambiguation and repair: resolving ambiguous intent and noisy spoken references, and asking a clarifying question or re-prompting rather than committing to a confident wrong answer.
  • Comfort placing deterministic guardrails around a probabilistic model, including safety floors, confirmation gating, and negation handling, and experience with state-machine or workflow engines covering branching, variable capture, and resumability.
  • Experience running models on constrained hardware such as NPUs, mobile, or embedded targets, under real latency, memory budgets. This includes ONNX and onnxruntime, model quantization, and cross-architecture packaging for aarch64.
  • Practical embedded development workflow: Linux, Docker, adb, systemd services, and the ability to diagnose problems from device logs.
  • The ability to take ownership of an existing, non-trivial codebase and keep it healthy.
  • A safety-first instinct, a misheard answer on a safety-critical step is treated as a hazard, not as a metric regression.
  • Measurement discipline: benchmarking under realistic device conditions, curated golden datasets used as regression gates, and attention to embedding collisions and centroid drift as the command catalog grows.
  • Excellent problem-solving skills and strong written and verbal communication, with the ability to collaborate across engineering, product, and domain experts.

Growth axis - desirable rather than required

  • SLM fine-tuning, including LoRA and QLoRA, instruction and format tuning, and distillation of a larger evaluator model into a smaller on-device model.
  • Latency and footprint optimization: INT8 and INT4 quantization, ONNX export and graph optimization, hardware-aware model selection, and profiling to reduce inference costs.
  • A pragmatic view of the boundary between configuration-driven adaptation and fine-tuning.
  • Building the data flywheel that supports this work, turning on-device session logs into evaluation sets and training data.

 

Bonus Qualifications

  • Speech recognition experience and comfort working downstream of noisy transcription.
  • Familiarity with LLM-as-a-judge evaluation.
  • Multilingual NLU; the system's data layer is already language-partitioned.
  • Industrial, field-service, or safety-critical product experience, for example in utilities, energy, or manufacturing.
  • GO familiarity, for integration with the on-device voice pipeline agent.
  • Exposure to MCP or agentic tooling.
  • Prior experience in a startup or a fast-paced team, building products from ground up.

 

Location & Type

  • Full-time, on-site or hybrid, based in Palo Alto, CA.
Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
368,634 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
Palo Alto
$41k – $89k per year (Estimated) • Remote/Hybrid • Full-Time • 8+ years exp • Bengaluru
C#
TypeScript
JavaScript
C#
.NET
Databases
Apache Kafka
AI/ML
Copilot
LLM
OpenAI
Frontend
Angular
GraphQL
DevOps
Azure
Azure AKS
Azure DevOps
CI/CD
Docker
GitHub
GitHub Actions
Grafana
Kubernetes
Prometheus
Rest API
Apply
$19k – $49k per year (Estimated) • In office • Full-Time • 5+ years exp • Bachelor's Degree • Hyderabad
SQL
C#
TypeScript
JavaScript
C#
.NET
Databases
MS SQL
Frontend
Angular
DevOps
CI/CD
Git
GitHub
Jenkins
Apply
$72k – $119k per year • In office • Full-Time • Bachelor's Degree • Boca Raton • Alpharetta • Dayton
Java
DevOps
CI/CD
Git
Apply
$80k – $175k per year • In office • Full-Time • Toronto
Python
AI/ML
AWS Bedrock
Claude
Copilot
LLM
Prompt Engineering
RAG
Context Engineering
AI Agents
DevOps
AWS
CI/CD
Splunk
GitHub
Apply
$79k – $159k per year (Estimated) • In office • Full-Time • 6+ years exp • Lincoln
C#
C#
.NET
DevOps
AWS
Azure
CI/CD
Docker
Dynatrace
GCP
GitHub
GitHub Actions
Grafana
Jenkins
Kubernetes
Splunk
QA
Cypress
JMeter
k6
Pact
Playwright
Postman
Rest-Assured
Selenium
Supertest
WebDriverIO
Apply
$78k – $196k per year (Estimated) • In office • 2+ years exp • Palo Alto
Python
AI/ML
Anomaly Detection
DVC
Keras
LightGBM
MLFlow
Multimodal AI
ONNX
PyTorch
SHAP
TensorFlow
TFLite
XGBoost
Time Series Forecasting
DevOps
CI/CD
Git
Apply
$125k – $261k per year (Estimated) • In office • 5+ years exp • Master's Degree • Palo Alto
Python
Databases
FAISS
Pinecone
Weaviate
AI/ML
Claude
Falcon
Fine-tuning
LangChain
LlamaIndex
LLM
Mistral
Multimodal AI
NLP
Prompt Engineering
PyTorch
RAG
Reinforcement Learning
Semantic Search
Transformers
Hugging Face
Semantic Search
AI Agents
Time Series Forecasting
DevOps
Vector
Apply
$146k – $235k per year (Estimated) • In office • 6+ years exp • Bachelor's Degree • Palo Alto
Bash
C++
Python
AI/ML
Computer Vision
OpenCV
DevOps
AWS
Azure
CI/CD
Docker
Git
RTOS
Robotics
Inverse Kinematics
Model Predictive Control
Motion Planning
ROS
ROS2
SLAM
IoT
MQTT
Apply
Staff Product Manager 3 months ago
$158k – $288k per year (Estimated) • In office • 8+ years exp • Palo Alto
AI/ML
Edge AI
Apply
Chief of Staff 8 hours ago
$120k – $150k per year • Equity 0.4–0.7% • In office • Full-Time • 3+ years exp • Palo Alto
AI/ML
AI Agents
Apply
$140k – $310k per year (Estimated) • Remote/Hybrid • Bachelor's Degree • Palo Alto
Databases
Apache Kafka
NATS
DevOps
AWS
Azure
CI/CD
Docker
GCP
Grafana
gRPC
Kubernetes
OpenTelemetry
Platform Engineering
Prometheus
Robotics
EtherCAT
IoT
MQTT
OPC UA
Apply
$139k – $294k per year (Estimated) • In office • Palo Alto
Python
AI/ML
Fine-tuning
Hybrid Search
LLM
Prompt Engineering
RAG
Human-in-the-Loop
Knowledge Graph
AI Agents
Function Calling
DevOps
AWS
Apply
$137k – $292k per year (Estimated) • In office • Palo Alto
Python
AI/ML
Hybrid Search
LLM
Prompt Engineering
RAG
Human-in-the-Loop
AI Agents
Function Calling
DevOps
AWS
Apply
$150k – $271k per year (Estimated) • Remote/Hybrid • Full-Time • 5+ years exp • Bachelor's Degree • Palo Alto
C++
Java
Python
Rust
AI/ML
LLM
Apply
See all jobs
This is one of many
368,634 more open roles from verified company boards, updated every day.