404,711open jobs
14,073companies
78,231added this week
Browse all
Location
In office (San Francisco)
Employment
Full-Time
Overview
Company
Impact
Profile match
Specter tracks growth signals across millions of private companies to help investors find opportunities early. Founded in 2021, it combines web, hiring, app and social data into company momentum scores. Venture funds and corporate development teams use it for sourcing.

Company Background

Specter's mission is to help automate the physical world.

Today, we build video sensors with state-of-the-art AI agents that answer any question, anywhere in their environments. Our systems can automatically detect and reason about any physical activity captured on camera, from security incidents (e.g. perimeter intrusion, theft, LPR), to safety monitoring (e.g. PPE detection, injured people), to operational efficiency (e.g. material tracking, congestion monitoring). We offer both long range wireless (1km range) and wired sensor variants to suit any deployment.

Our co-founders Xerxes and Philip are passionate about empowering our partners in the fast approaching world of physical AI and robotics. We are a small, fast growing team who hail from Anduril, Tesla, Uber, and the U.S. Special Forces.

The Role

We’re hiring a Fleet Reliability Engineer to keep our sensor fleet running in the field by building the data, analytics, and recovery mechanisms that prevent failures from becoming incidents. As we scale toward thousands of sensors, fleet health becomes a data-and-systems problem.

This is the proactive, highest-leverage side of reliability: own the telemetry and data pipeline, verify that fixes hold fleet-wide, and turn field signal into cost-weighted decisions about what to fix first. Much of today’s operational load is addressable through better instrumentation, alert hygiene, and recovery verification, at little to no field cost.

Responsibilities:

Reliability Data Platform - Primary

  • Own the fleet’s reliability data pipeline end to end: telemetry aggregation, storage, and instrumentation.

  • Drive down observability cost - own the tooling spend and cut what we pay for but don’t use.

  • Instrument the fleet and own the health metrics that measure reliability.

Proof-of-Recovery & Alert Hygiene

  • Verify that fixes hold fleet-wide, not just on the device that paged.

  • Cut alert noise at the source - separate real failures from self-resolving ones.

  • Turn repeat failure patterns into automated detection and recovery.

Fleet Health & Failure-Mode Analytics

  • Turn fleet telemetry into a live picture of which cohorts, hardware revisions, and firmware versions are trending toward failure, and why.

  • Build the failure-mode analysis that tells engineering what to fix at the source.

  • Own fleet-wide trend and forecasting work, including power and solar planning.

Reliability Economics & Prioritization

  • Score reliability work in dollars - field-trip cost, hardware-return cost, observability spend - and prioritize the most expensive problems first.

  • Set and track the fleet’s reliability targets: uptime, offline rate, truck-rolls per sensor-year.

  • Give the team the data to make reliability-versus-cost tradeoffs.

Qualifications:

  • Strong data and software skills - Python (or Go) and SQL - and the ability to own a data pipeline end to end.

  • Hands-on building and tuning observability stacks (OpenTelemetry, Grafana, Prometheus, Datadog, or similar), including their cost.

  • Experience operating physical or embedded device fleets at scale, and reasoning about how hardware fails in the field.

  • Comfortable turning messy field telemetry into trends, failure modes, and forecasts.

  • Fluency with databases and data modeling (PostgreSQL or equivalent); infrastructure-as-code familiarity (Terraform or similar) a plus.

  • Bias toward building mechanisms over doing manual work.

  • Nice to have: reliability/SRE fundamentals (SLOs, error budgets, proof-of-recovery) applied to a physical fleet.

  • Nice to have: experience across the hardware-software boundary - power, connectivity, and physical failure modes.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
404,711 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
San Francisco
Remote • Bachelor's Degree
PowerShell
Python
DevOps
Ansible
Azure
CI/CD
Configuration Management
Datadog
GCP
GitLab
GitLab CI
IAM
Incident Management
OpenTelemetry
SigNoz
Terraform
VMWare
Windows Server
Cybersecurity
Crowdstrike
GDPR
HashiCorp Vault
ISO 27001
Okta
Qualys Cloud Platform
Tines
Cryptography
Vault
Management
Power Automate
Apply
$13k – $31k per year (Estimated) • Remote/Hybrid • Full-Time • Bachelor's Degree • Chennai
AI/ML
AI Agents
Edge AI
Apply
$13k – $30k per year (Estimated) • Remote/Hybrid • Full-Time • Bachelor's Degree • Mumbai
AI/ML
AI Agents
Edge AI
Apply
$13k – $29k per year (Estimated) • Remote/Hybrid • Full-Time • Bachelor's Degree • Mumbai
SQL
Databases
MS SQL
AI/ML
AI Agents
Edge AI
DevOps
Azure
Analytics
Power BI
Tableau
Apply
$18k – $43k per year (Estimated) • Remote/Hybrid • Full-Time • Bachelor's Degree • Mumbai
AI/ML
AI Agents
Edge AI
DevOps
IAM
Incident Management
Apply
$151k – $308k per year (Estimated) • In office • Full-Time • San Francisco
C++
Go
Python
Rust
AI/ML
AI Agents
Time Series Forecasting
Physical AI
Cybersecurity
Wireshark
Apply
$149k – $305k per year (Estimated) • In office • Full-Time • San Francisco
Python
Databases
Apache Kafka
Redpanda
Kafka
AI/ML
Computer Vision
Multimodal AI
ONNX
PyTorch
TensorRT
Physical AI
DevOps
CI/CD
Docker
Kubernetes
Apply
$130k – $257k per year (Estimated) • In office • Full-Time • 5+ years exp • Bachelor's Degree • San Francisco
MATLAB
Python
AI/ML
AI Agents
Physical AI
Apply
$150k – $306k per year (Estimated) • In office • Full-Time • San Francisco
C++
Python
Rust
TypeScript
C
C
Embedded C
AI/ML
AI Agents
Physical AI
DevOps
Buildkite
CI/CD
GitHub Actions
NixOS
Self-Healing
GitHub
Apply
$125k – $280k per year (Estimated) • In office • Full-Time • San Francisco
Bash
Go
Python
Rust
AI/ML
AI Agents
Physical AI
DevOps
Amazon EKS
AWS
CI/CD
GitOps
Kubernetes
NixOS
Platform Engineering
Terraform
IAM
Apply
$85k – $105k per year • Equity 0–0.1% • In office • Full-Time • San Francisco
Management
Slack
Marketing
HubSpot
LinkedIn
Apply
$260k – $310k per year • Equity 0.1–0.4% • In office • Full-Time • 3+ years exp • San Francisco
Management
Slack
Marketing
HubSpot
Apply
In office • Internship • San Francisco
AI/ML
LLM
Management
Slack
Apply
$65k – $100k per year • Remote • Contractor • San Francisco
AI/ML
Claude
Apply
Coordinator: Docket 2 hours ago
$71k – $95k per year • In office • 2+ years exp • Bachelor's Degree • San Francisco
Apply
See all jobs
This is one of many
404,711 more open roles from verified company boards, updated every day.