368,530open jobs
9,432companies
50,439added this week
Browse all
Salary
$175k – $200k per year
Location
Remote (United States)
Seniority
Staff · 8+ years exp
Employment
Full-Time
Overview
Company
Impact
Profile match
Arcadia is a climate technology company headquartered in Washington, District of Columbia, and founded in 2014. The company runs Arc, a platform that pulls utility bill and meter data from thousands of providers so other companies can build energy products on top of it, and enrolls households in community solar projects. It works with clean energy developers, software vendors, and consumer brands that need standardized access to fragmented utility data.

Arcadia is the most trusted healthcare platform powering outcomes. We transform complex healthcare data into trusted intelligence, helping providers, payers, and life sciences organizations act with clarity, make confident decisions, and achieve measurable clinical, operational, and financial outcomes.

Built on a comprehensive data foundation spanning tens of millions of patient lives, Arcadia combines advanced analytics and responsible AI to surface meaningful insights, coordinate action, and improve performance at scale. Our approach to AI and automation is governed and transparent - designed to strengthen human expertise, not replace judgment or obscure responsibility.

Hundreds of organizations rely on Arcadia to improve cost, quality, and outcomes. Backed by Nordic Capital, we continue to invest in our platform, AI capabilities, and people as we pursue our purpose: helping healthcare deliver better outcomes for every person, every community, and every generation.

Why This Role Is Important to Arcadia

Arcadia’s data and analytics platform is used by hundreds of health systems, ACOs, payers, and life sciences organizations, touching tens of millions of patient lives. This role owns how our agentic capabilities perform at that same scale: accurate, transparent about their own confidence, and safe for the clinicians, care teams, and patients who depend on them.

As a staff-level individual contributor, you will own the product-layer decisions that shape agent behavior, including prompting, retrieval and context, memory and state, evaluation, and escalation, while partnering with Product and Engineering on the systems that support them. Your work will help Arcadia make evidence-based launch decisions and scale responsible AI that is steerable, trustworthy, and ready for real healthcare workflows.

What Success Looks Like

In 3 months

  • You have establisheda production-grounded baseline for priority agentic workflows, with documented failure modes, severity-weighted evaluationrubrics, and a clear measurement plan
  • You have mapped the current retrieval, context, memory, and escalation patterns and identifiedthe highest-value opportunities to improve reliability, calibration, and cost
  • You have earned trust across Product and Engineering by turning production evidence into clear, actionable recommendations

In 6 months

  • Production-representative evaluation suites and regression checks inform model-change decisions for priority agentic workflows
  • You have delivered measurable improvements in accuracy, reliability, steerability, latency, or cost for one or more priority workflows
  • Human-review and escalation behavior hasbeen validatedunder adversarial and edge-case conditions, with decision criteria and ownership boundaries clearly documented

In 12 months

  • Arcadia has a repeatable product-layer AI performance practice that moves from production failure to diagnosis, experiment, evaluation, and release decision
  • High-severity regressions are caught earlier, and agentbehavior is more transparent, calibrated, and trustworthy at scale
  • Model cards, intended-use guidance, limitations, and performance documentation are current and useful to product and customer-facing teams

What You'll Be Doing

  • Design and iterate onagent behavior across real, live workflows, including long-horizon, multi-turn agentic tasks
  • Design retrieval and context architecture so the right source data reaches a model in the right structureand agents remaingrounded in real data rather than filling gaps with assumptions
  • Design memory and state handling across multi-turn and multi-agent flows, determiningwhat is carried forward, summarized, or dropped and why
  • Create context and prompt templates that combine few-shot examples, structured formatting, and reasoning scaffolding for consistent agent behavior
  • Improve performance through prompting, tool-use strategy, and context construction, validatedthrough direct experimentation rather than guesswork
  • Build and run evaluations against real production conditions to measure performance, regressions, failure modes, and edge cases
  • Author evaluation rubrics, quality heuristics, and thresholds that weight failures by severity and cost, not just frequency, and monitorthose measures against production behavior
  • Design and validateescalation paths that route agents to human review based on confidence and uncertainty while preserving safety and consistency under adversarial and edge-case conditions
  • Design for cost-aware performance alongside latency, reliability, and accuracy through efficient context construction and tool-call economy
  • Evaluate and sign off on model changes by baselining current behavior, running comparative evaluations, and making the go/no-go call before a change reaches a customer
  • Maintain product-level AI documentation, including model cards, intended use, limitations, and known failure modes, so customer-facing teams work from actual agent behavior
  • Partner closely with Product and product managersto ensure agents are not just capable, but steerable, trustworthy, and ready to scale

What You'll Bring

  • We value equivalent practical experience that demonstratesthe depth requiredfor this staff-level role
  • 8+ years of production software engineering experience, including 3+ years of hands-on ownership of ML, LLM, or agentic systems in production, with direct experience in healthcare, finance, or another regulated industry
  • Demonstrated ability to diagnose why an agent failed, correctly attribute the fix to instruction, retrieval, context, or memory design, and weigh failures by severity and cost rather than frequency alone
  • Hands-on experience with RAG architecture, production-grounded evaluation frameworks, and fallback or human-in-the-loop logic for automated systems
  • Working familiarity with AWS AI/ML services, including Bedrock and SageMaker, sufficient to build and evaluate effectively in Arcadia’s environment
  • Evidence-led judgment and the credibility to push back on launch decisions, paired with a builder’s instinct to run the experiment and move from a production failure to a fix

Would Love for You to Have

  • Experience applying AI to healthcare data or workflows where safety, transparency, and calibrated uncertainty directly affect care teams or patients
  • Experience with long-horizon, multi-turn or multi-agent workflows and product-level AI documentation such as model cards

What You'll Get

  • The opportunity to define how agent performance, safety, and readiness are measured for productionhealthcare workflows
  • Meaningful ownership across prompts, context, memory, evaluations, and escalation patterns at product scale
  • A cross-functional role translating production evidence into AI improvements used across Arcadia’s platform
  • A mission-driven company working to improve how patients receive care
  • A flexible, remote-friendly culture with personality and heart
  • Employee-driven programs and initiatives for personal and professional development
  • Membership in the talented, energized, diverse, and purpose-driven Arcadian community
Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
368,530 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
In your city
$110k – $131k per year • Remote • Full-Time • 10+ years exp • Bachelor's Degree
DevOps
AWS
Incident Management
VMWare
Apply
$169k – $321k per year (Estimated) • In office • 8+ years exp • Bachelor's Degree • Phoenix
AI/ML
AI Agents
Anomaly Detection
LLM Guardrails
DevOps
AWS
Kong
Amazon S3
API Gateway
Cybersecurity
Zero Trust
Apply
$152k – $239k per year • Remote • Full-Time • 8+ years exp
SQL
Databases
Snowflake
AI/ML
LLM
Model Context Protocol
Analytics
A/B Testing
Apply
$105k – $252k per year • Remote • Full-Time • 18+ years exp • Bachelor's Degree
Python
Java
Java
Gradle
DevOps
Ansible
AWS
CI/CD
CloudFormation
Configuration Management
Docker
GitHub Actions
GitLab CI
Helm
Jenkins
Kubernetes
Platform Engineering
Terraform
GitHub
GitLab
Cybersecurity
Sonatype Nexus IQ
Management
Confluence
Jira
Apply
$69k – $171k per year (Estimated) • Remote/Hybrid • Full-Time • 3+ years exp • Bachelor's Degree • Bogotá
SQL
Databases
Databricks
AI/ML
dbt
LLM
NLP
Context Engineering
Analytics
Power BI
Tableau
Apply
$125k – $145k per year • Remote • Full-Time
Python
SQL
AI/ML
dbt
DevOps
AWS
Management
Confluence
Jira
Apply
$250k – $350k per year • Remote • Full-Time
Java
Kotlin
Python
TypeScript
Databases
Apache Kafka
ElasticSearch
OpenSearch
AI/ML
dbt
Spark
AI Agents
DevOps
AWS
Kubernetes
Platform Engineering
Cybersecurity
HIPAA
Apply
$141k – $281k per year (Estimated) • Remote • Full-Time • United States
Java
Kotlin
Python
TypeScript
Databases
Apache Kafka
ElasticSearch
OpenSearch
AI/ML
dbt
Spark
AI Agents
DevOps
AWS
Kubernetes
Platform Engineering
Cybersecurity
HIPAA
Apply
Director, Analytics 2 months ago
$180k – $200k per year • Remote • Full-Time • 10+ years exp
SQL
DevOps
AWS
Analytics
Power BI
Tableau
Apply
$180k – $220k per year • Remote • Full-Time • 12+ years exp • Bachelor's Degree
AI/ML
RAG
AI Agents
Apply
See all jobs
This is one of many
368,530 more open roles from verified company boards, updated every day.