368,910open jobs
9,449companies
47,822added this week
Browse all
Salary
$180k – $220k per year
Location
Remote (United States)
Seniority
Staff · 10+ years exp
Overview
Company
Impact
Profile match
Ursa Space Systems aggregates synthetic aperture radar imagery from many satellite operators and turns it into analytic products. Founded in 2014 in Ithaca, New York, it is best known for measuring global oil storage from radar shadows. Its data reaches commodity traders, insurers and government users.

AI TEVV Engineer (Test, Evaluation, Verification & Validation)

About Ursa Space Systems

Ursa Space Systems is building an AI-native geospatial insights platform that guides the acquisition, analysis, and integration of satellite and geospatial data into customer workflows, giving decision makers an edge. Leveraging hundreds of data sources, AI agents, and proprietary analytics, Ursa Space provides fast, actionable information to a range of industries, including finance, energy, and defense. Our customers receive contextual, comprehensive reporting that goes beyond surface-level observations.

Job Summary

Ursa Space is looking for an AI TEVV Engineer to define how we prove our AI-native geospatial platform works and to whom. This is an evaluation-science role at its core, not a test-automation role. The central skill is measurement under uncertainty: where classical QA asks "does this function return the correct output" (a deterministic pass/fail question), AI evaluation asks "what is the error rate, on what distribution of inputs, under what operating conditions, and is that rate acceptable for this mission" (a measurement question with confidence intervals). That work is closer to experimental design and psychometrics than to writing test suites. You will own the strategy, methodology, and evidence that let us and our customers' risk officers trust the platform's outputs and demonstrate where, and how well, they hold.

This position reports to the Director of System Requirements, with a functional reporting line and charter that preserve evaluation independence from the teams whose outputs are under test. This position is fully remote and exempt.

Responsibilities

  • Design, develop, and plan the TEVV strategy across the platform's algorithms, AI/ML models, agentic workflows, data pipelines, and analytic products, aligned to the NIST AI RMF Measure function
  • Design statistically defensible evaluations: error metrics and acceptance criteria on representative input distributions, with explicit confidence intervals
  • Define and document the platform's context of use - the validated operating envelope (modalities, geographies, resolutions, conditions, target classes) within which accuracy claims hold
  • Evaluate ground-truth and "golden" datasets, including annotation and adjudication protocols, inter-rater reliability, and quantified uncertainty in the reference data itself
  • Implement a layered evaluation posture: a verifiable core (accuracy, groundedness, format), a rubric-scored middle layer with documented inter-rater reliability, and an honest residual of expert holistic review
  • Evaluate generative and natural-language outputs for claim-level groundedness whether each assertion is traceable to a citable source alongside rubric-based, human-adjudicated assessment
  • Stand up continuous monitoring and re-validation certification gates plus ongoing surveillance watching for model, prompt, retrieval, and agent-behavior drift
  • Author and maintain the TEVV evidence set: test plans, traceability matrices, metrics, acceptance criteria, and credibility-assessment documentation
  • Support DoD AI test-and-evaluation expectations (including DoD Directive 3000.09), contractual milestones, acceptance testing, and demonstrations to government stakeholders.
  • Distinguish internal TEVV from organizationally independent IV&V, and partner with external IV&V agents where required
  • Partner with Engineering teams to embed evaluability, observability, and traceability from design onward.
  • Contribute to emerging standards (NIST AI TEVV consortium, ISO/IEC SC 42 / 42001), aligning our methodology so evidence packages map to customers' compliance frameworks.
  • 30% travel.
  • Perform all other duties as assigned.

Requirements

  • B.S. in Computer Science, Statistics, or Systems Engineering, or a related quantitative discipline (M.S./Ph.D. a plus)
  • 10+ years of relevant experience, centered on evaluation, measurement, or test-and-evaluation of AI/ML or data-driven systems - not solely software QA or test automation
  • Demonstrated ability to design statistically defensible evaluations: input-distribution design, error-rate estimation, confidence intervals, and context-tied acceptance criteria
  • Hands-on experience building ground-truth/golden datasets - adjudication protocols, inter-rater reliability, and reference-data uncertainty
  • Experience supporting U.S. government contracts (aerospace, defense, or intelligence preferred), including requirements traceability and compliance documentation
  • Working knowledge of the NIST AI RMF and how TEVV evidence maps to customer compliance regimes
  • Strong quantitative skills and Python proficiency for analysis and evaluation (Pandas/Polars, NumPy, ML evaluation libraries)
  • Comfort using AI-assisted tools for rapid development and testing
  • Organized and self motivated, able to work successfully with a remote team
  • A creative, flexible mindset for complex problems
  • A fast, reliable internet connection if working remotely

Preferred Skills

  • Aligning evaluation methodology to the NIST AI RMF, ISO/IEC 42001, and emerging NIST AI TEVV consortium / ISO/IEC SC 42 work; standards participation a plus
  • Defining credibility-assessment frameworks tied to context of use rather than fixed, context-free thresholds
  • Evaluating image and signal processing outputs across modalities (SAR, electro-optical, RF), including the proxy nature of geospatial reference data
  • Evaluating generative and agentic systems: rubric design, human-adjudicated evaluation, and claim-level groundedness
  • Heritage V&V/assurance standards (IEEE 1012, DO-178C, ISO/IEC 25010, CMMI) and formal IV&V experience
  • GIS tools and libraries; SpatioTemporal Asset Catalog (STAC) experience
  • NoSQL and/or SQL databases (Mongo, MySQL, Postgres)
  • Test automation and CI/CD: Python and/or JavaScript, frameworks (e.g., pytest, Jest), and regression/monitoring suites in pipelines
  • Common AWS services (e.g. S3, Lambda, ECS, ECR, DynamoDB) and microservice-based architectures
  • Software tooling (e.g. Git, Docker, Anaconda, virtual environments)
  • Data quality, observability, and monitoring tooling
  • Experience with customer-facing software products

Compensation

  • Ranges: $180,000 - $220,000
  • Compensation range includes base salary and is eligible for an annual bonus.
  • New hires salaries are typically between the range minimum and the salary range midpoint. Actual placement in the range will depend on a candidate’s job-related skills, experience, and expertise, as evaluated during the interview process.

Inclusion Statement

We are dedicated to the belief that all lives have equal value. We strive for a global and cultural workplace that supports ever greater diversity, equity, and inclusion - of voices, ideas, and approaches - and we support this diversity through all our employment practices.

All applicants and employees who are drawn to serve our mission will enjoy equality of opportunity and fair treatment without regard to race, color, age, religion, pregnancy, sex, sexual orientation, disability, gender identity, gender expression, national origin, genetic information, veteran status, marital status, and prior protected activity.

Location

  • We are headquartered in Ithaca, NY and have a remote workforce in other locations throughout the United States.

Please note: applications without a relevant cover letter will not be considered. In your cover letter, we would like to hear your personal voice and learn about your sincere interest in Ursa Space Systems.

Benefits and Perks

  • Competitive Compensation
  • Discretionary PTO & Flexible Scheduling
  • Stock Options
  • 401(k) Match
  • Medical, Dental and Vision Coverage for you and your dependents
  • FSA & HSA Plans
  • Employer-paid Life Insurance
  • Employer-paid LTD and STD for Parental and Family Care
  • 11 Paid Holidays
  • Employee Resource Groups
  • Educational Assistance Program
  • Professional Development Opportunities
  • And more…

Company Values

  • Use the team
  • Figure it out and own it
  • Aim for elegant simplicity
  • Empower diversity & inclusivity
  • Do the right thing
  • Be scrappy
Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
368,910 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
In your city
$35k per year • In office • Internship • Bachelor's Degree • Freiburg im Breisgau
C++
Java
Node JS
Python
C#
JavaScript
C#
.NET
AI/ML
AI Agents
DevOps
CI/CD
GitLab
Apply
$80k – $175k per year • In office • Full-Time • Toronto
Python
AI/ML
AWS Bedrock
Claude
Copilot
LLM
Prompt Engineering
RAG
Context Engineering
AI Agents
DevOps
AWS
CI/CD
Splunk
GitHub
Apply
$35k per year • In office • Internship • Bachelor's Degree • Freiburg im Breisgau
JavaScript
Python
TypeScript
Apply
$126k – $227k per year (Estimated) • In office • Contractor • 6+ years exp • Bachelor's Degree • Fremont
C#
Python
Apply
$158k – $288k per year (Estimated) • In office • 8+ years exp • Bachelor's Degree • Chicago
Python
SQL
Python
pySpark
Databases
Databricks
Snowflake
AI/ML
Spark
DevOps
AWS
Apply
$185k – $199k per year • Equity • Remote/Hybrid • 4+ years exp • Bachelor's Degree
Python
SQL
AI/ML
Claude
Computer Vision
Fine-tuning
Function Calling
Image Segmentation
LangGraph
LLM
LoRA
Multimodal AI
PEFT
PyTorch
LangChain
Transformers
OpenAI Agents SDK
Structured Outputs
AI Agents
Model Context Protocol
DevOps
AWS
CI/CD
Docker
Git
Vector
Management
Confluence
Jira
Miro
SpaceTech
GDAL
Apply
See all jobs
This is one of many
368,910 more open roles from verified company boards, updated every day.