604,099open jobs
31,594companies
86,696added this week
Browse all
Salary
$180k – $400k per year
Location
In office (Santa Clara)
Seniority
Senior
Employment
Full-Time
Overview
Company
Impact
Profile match

About Boson AI: At Boson AI, we are not just building AI solutions; we are pioneering the future of enterprise AI. Driven by a passion for cutting-edge AI research, particularly in the transformative areas of large language models and agentic systems, our mission is to tackle the most complex real-world problems for businesses and unlock significant value. We are a dynamic and collaborative team of researchers and engineers who thrive on pushing the boundaries of what's possible, dedicated to delivering high-quality, reliable products that seamlessly integrate into the fabric of enterprise workflows and set new industry standards.

About the Role: Build and operate the core platform behind Boson's model APIs and agentic products. You'll own the infrastructure that every Boson agent runs on - API serving, state management, data pipelines, context retrieval, and execution runtime - and make it fast, reliable, and easy for product teams to build on.

Responsibilities

  • Own and evolve the core platform infrastructure: API serving layer, state management, policy enforcement engine, and execution runtime for agentic workflows.
  • Design and operate high-throughput, low-latency distributed services that back our model API products - including request routing, load management, rate limiting, and multi-tenant isolation.
  • Build and maintain downstream data pipelines (ETL/ELT) for API logs, usage analytics, and billing - ensuring data correctness, freshness, and queryability at scale.
  • Develop production-grade internal SDKs and libraries with clean APIs, strong type safety, and clear contracts that product teams can build on confidently.
  • Architect context and memory systems for conversational workloads - low-latency retrieval, caching, and integration with vector stores and retrieval pipelines.
  • Instrument end-to-end observability: define SLIs/SLOs, build structured logging and tracing, and drive reliability improvements across the platform.
  • Collaborate closely with ML and product teams to integrate model serving, voice runtime, and tooling infrastructure under tight latency and quality constraints.

Qualifications

  • 3+ years building and operating backend systems at scale - you've owned services that other teams depend on in production.
  • Strong distributed systems fundamentals: concurrency, fault tolerance, consistency tradeoffs, capacity planning.
  • Hands-on experience with data pipeline infrastructure (Kafka/Kinesis, Spark/Flink, Airflow, or similar) for log processing, analytics, or ETL workloads.
  • Track record of designing APIs and frameworks adopted by other engineering teams - you care about developer experience and long-term maintainability.
  • Proficiency in at least one systems language (Go, Rust, Java, C++) or Python in a performance-sensitive context.
  • Comfortable working across the stack: cloud infrastructure (AWS/GCP), containerized deployments (K8s), CI/CD, and production oncall.

Bonus point

  • Experience with LLM serving, agentic orchestration patterns (ReAct, planner-executor), or RAG pipelines.
  • Familiarity with emerging agent integration protocols (MCP, A2A) or orchestration frameworks (LangChain, LlamaIndex).
  • Background in real-time media systems (audio/video streaming, low-latency signaling).
  • Experience building high-stakes platform services (payments, identity, core data) where correctness and auditability are non-negotiable.
Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
604,099 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
Santa Clara
Software Engineer 4 hours ago
$130k – $220k per year • In office • Full-Time • 3+ years exp • San Francisco
Python
Go
Java
Rust
AI/ML
Prompt Engineering
AI Agents
Semantic Search
Semantic Search
Tool Use
Apply
$56k – $147k per year (Estimated) • Remote/Hybrid • Internship • Master's Degree • Ireland
Python
JavaScript
SQL
C#
Analytics
Power BI
Apply
$37k – $61k per year (Estimated) • Remote/Hybrid • Internship • Bachelor's Degree • Dublin
Python
MATLAB
Analytics
Power BI
Apply
Cloud Engineer 1 hour ago
$179k – $318k per year • In office • TS/SCI • 6+ years exp • Bachelor's Degree
Python
DevOps
Helm
CI/CD
GitOps
AWS
Docker
Kubernetes
Apply
$128k – $214k per year • In office • TS/SCI • 14+ years exp • Bachelor's Degree
Python
DevOps
Splunk
Ansible
CloudFormation
Packer
GitLab CI
CI/CD
Windows Server
AWS
Configuration Management
GitLab
Cybersecurity
Nessus
Analytics
Apache NiFi
Apply
Datacenter Technician 15 days ago
$36k – $72k per year • In office • Full-Time • Barrie
AI/ML
NLP
Apply
$150k – $270k per year • In office • Full-Time • Santa Clara
Python
Go
Java
Rust
C++
Databases
Apache Kafka
AI/ML
Spark
AI Agents
Flink
LLM
Edge AI
Agentic Workflows
DevOps
GCP
CI/CD
AWS
Docker
Kubernetes
Vector
Amazon Kinesis
Analytics
ETL/ELT
Apply
$90k – $180k per year • In office • Full-Time • 4+ years exp • Toronto
AI/ML
CUDA Toolkit
CUDA
NCCL
InfiniBand
DevOps
Terraform
Ansible
GCP
Prometheus
SLURM
Azure
AWS
Kubernetes
Grafana
SRE
Configuration Management
HPC
Apply
Frontend Engineer 1 year ago
$150k – $400k per year • In office • Full-Time • 3+ years exp • Bachelor's Degree • Santa Clara
JavaScript
TypeScript
Databases
Supabase
AI/ML
Multimodal AI
AI Agents
LLM
Edge AI
Frontend
Tailwind CSS
Next.js
D3.js
Chart.js
React.js
Apache ECharts
Mobile
Firebase
DevOps
WebRTC
WebSockets
CI/CD
GitHub
Management
Agile
Apply
$108k – $289k per year • In office • Full-Time • Bachelor's Degree • Toronto
Python
Rust
TypeScript
AI/ML
Fine-tuning
JAX
Multimodal AI
AI Agents
PyTorch
LLM
RAG
Edge AI
DevOps
GCP
WebRTC
Azure
AWS
GitHub
Apply
$140k – $311k per year (Estimated) • Remote/Hybrid • Full-Time • 10+ years exp • PhD • Santa Clara
JavaScript
Java
AI/ML
AI Agents
Frontend
Svelte
Web3
Layer 2
Robotics
Digital Twin
Apply
$145k – $314k per year (Estimated) • Remote/Hybrid • Contractor • Master's Degree • Santa Clara
Python
AI/ML
llama.cpp
LoRA
vLLM
Fine-tuning
Quantization
Multimodal AI
Knowledge Distillation
AI Agents
VLM
SGLang
GGUF
TensorRT
PEFT
TensorRT-LLM
Transformers
PyTorch
LLM
Mixture of Experts
DPO
Post-training
Edge AI
Speculative Decoding
KV Cache
Multi-Agent Systems
Model Distillation
Analytics
ETL/ELT
Management
Freshdesk
Apply
$201k – $352k per year • Equity • In office • Full-Time • 10+ years exp • Bachelor's Degree • Santa Clara
Python
Java
TypeScript
AI/ML
Cursor
Windsurf
Claude Code
Embeddings
AI Agents
LLM
RAG
OpenAI Codex
LLM Guardrails
Mobile
Clean Architecture
DevOps
Vector
Management
ServiceNow
Apply
$240k – $420k per year • Equity • In office • Full-Time • 10+ years exp • Bachelor's Degree • Santa Clara
Python
Java
AI/ML
Cursor
Windsurf
Claude Code
Embeddings
Function Calling
AI Agents
RAG
OpenAI Codex
Human-in-the-Loop
LLM Guardrails
Multi-Agent Systems
Tool Use
Mobile
Clean Architecture
DevOps
Vector
Management
ServiceNow
Apply
$240k – $420k per year • Equity • In office • Full-Time • 15+ years exp • Bachelor's Degree • Santa Clara
Python
Go
Java
AI/ML
Cursor
Windsurf
Claude Code
AI Agents
RAG
OpenAI Codex
LLM Guardrails
Mobile
Clean Architecture
DevOps
Vector
Management
ServiceNow
Apply
See all jobs
This is one of many
604,099 more open roles from verified company boards, updated every day.