Salary
≈ $110k – $194k per year (Estimated)
Location
Remote (United States)
Seniority
Senior · 5+ years exp
Overview
Company
Impact
Profile match
SimpliSafe is a home security company headquartered in Boston, Massachusetts, and founded in 2006. The company designs and manufactures wireless security systems, including motion sensors, entry sensors, indoor and outdoor cameras, and smoke detectors, all of which are integrated with professional monitoring services. It serves over six million customers across the United States and the United Kingdom, offering DIY installation and no-contract subscription plans for residential and small business security.
We are looking for an experienced Data Engineer to join our machine learning team and serve as the vital link between raw data and high-performing ML models. In this role, you will drive the design, implementation, and management of an integrated system dedicated to data curation, annotation workflows, quality assurance, and dataset composition. In this role, you will drive the design and implementation of an integrated data system dedicated to advanced data curation, annotation workflows, quality assurance, and dataset composition. If you have a passion for data infrastructure, automation, and optimizing human-in-the-loop annotation processes at scale, we'd love to hear from you.
Responsibilities:
- Data Lake Management: Design, build, and maintain our AWS data lake architecture (focusing on S3 Athena, and Glue) to serve as the highly accessible foundation for all ML data workloads.
- ETL and Pipeline Development: Architect and build scalable, robust data pipelines to ingest, transform, and deliver massive datasets seamlessly.
- Build Specialized Datasets: Partner closely with ML engineers and modelers to curate raw data and prepare high-quality baseline training sets, experimental datasets for testing new ML hypotheses, and targeted validation sets specifically designed to evaluate "hard cases" and edge cases.
- Advanced Data Curation: Collaborate with the modeling team to leverage data embeddings to cluster, filter, and surface the most informative data points for
- annotation, ensuring we are spending our annotation budget on the highest-value data.
- Platform Integration: Build secure, reliable integrations between our internal data ecosystem and specialized third-party data annotation platforms.
- Manage the Annotation Lifecycle: Oversee end-to-end annotation workflows, including job creation, platform configuration, and syncing/maintaining large annotated datasets.
- Ensure Data Quality: Partner with ML engineers and modelers to conduct spot annotation quality assurance (QA), establishing a shared workflow to ensure labeled data meets strict accuracy and consistency standards at scale.
- Data Governance and Schema Design: Design optimized database schemas for large-scale data exploration and implement best practices in data cataloging, dataset version control, and data stewardship.
- Best Practices: Advocate for and implement best practices in data stewardship, dataset version control, and system monitoring.
Requirements:
- Experience: 5+ years of experience in data engineering, with at least 3 years explicitly focused on building data pipelines, managing data lakes, and curating large-scale datasets.
- Cloud Data Ecosystems: Strong, hands-on experience with the AWS Data stack, specifically Amazon S3 Athena, and Glue (or deep expertise in an equivalent ecosystem like GCP BigQuery).
- Programming and Databases: High proficiency in Python and advanced SQL, with a proven track record of designing robust database schemas and complex ETL processes.
- Navigating Ambiguity: Proven skill in developing organized, reliable data structures and pipelines within highly unstructured data settings.
- Data Lifecycle: Solid understanding of the ML data lifecycle, including data tracking, dataset versioning, and the specific data structures required by modeling teams.
- Communication: Exceptional communication skills, specifically the ability to translate complex data requirements into easily understandable instructions for annotation teams.
Nice to Have:
- Embedding and Vector Tech: Experience working with embeddings for data curation as well as familiarity with vector databases or similarity search libraries (e. g., FAISS, Pinecone, and Milvus).
- Big Data Technologies: Experience with distributed computing and large-scale data processing frameworks (e. g., Spark, Hadoop, Ray, or AWS EMR).
- Annotation Platforms: Hands-on experience with specialized data labeling and annotation platforms (e. g., Scale AI, Labelbox, Snorkel, Toloka).
- Data Tracking Tools: Familiarity with tools such as MLflow or Weights and Biases for dataset and artifact tracking.
Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
368,291 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Free forever. No card. Under a minute.
Your match
How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.
Recommended for you based on this role
Similar stack
Same company
In your city
Application Security Engineer
1 day ago
$185k – $260k per year • Remote • Full-Time • 8+ years exp • Bachelor's Degree
DevOps
AWS
CI/CD
GCP
Kubernetes
GitHub
Cybersecurity
Clair
Dependabot
OWASP Top 10
OWASP ZAP
Snyk
Trivy
Apply
Cloud Engineer (AWS)
4 hours ago
In office • Full-Time • 5+ years exp • Bachelor's Degree • Dalian
Python
DevOps
AWS
CI/CD
CloudFormation
Docker
FinOps
GitHub Actions
Kubernetes
Terraform
GitHub
Apply
≈ $20k – $48k per year (Estimated) • In office • Full-Time • Moscow
Python
SQL
Databases
Oracle
AI/ML
Hadoop
Management
Confluence
Jira
Apply
Full-Stack Software Engineer
1 day ago
Remote/Hybrid • Full-Time • 4+ years exp • Hong Kong
Node JS
JavaScript
Node JS
Nest.JS
Databases
Apache Kafka
Frontend
Vue.js
DevOps
AWS
Docker
Terraform
Amazon ECS
IoT
MQTT
Apply
$21k per year • In office • Contractor • Yekaterinburg
Python
SQL
Python
FastAPI
Flask
AI/ML
Claude
Claude Code
Embeddings
Function Calling
LLM
RAG
OpenAI
OpenAI Codex
Structured Outputs
DevOps
Docker
Git
Apply
Sr Business System Analyst
3 days ago
$89k – $119k per year • Remote/Hybrid • 3+ years exp • Bachelor's Degree • Boston
SQL
Analytics
Tableau
Management
Jira
Apply
Sr. Manager, Data Platform Engineering
17 days ago
$171k – $228k per year • Remote/Hybrid • 8+ years exp • Bachelor's Degree • Boston
Databases
Apache Kafka
Databricks
DevOps
AWS
CI/CD
Platform Engineering
Amazon S3
Apply
Customer Experience Analyst
18 days ago
$65k – $88k per year • Remote/Hybrid • 3+ years exp • Master's Degree • Richmond
Python
SQL
Analytics
Tableau
Apply
Customer Experience Analyst
18 days ago
$80k – $105k per year • Remote/Hybrid • 3+ years exp • Master's Degree • Boston
Python
SQL
Analytics
Tableau
Apply
Senior Software Engineer, CLV Monitoring
20 days ago
$117k – $156k per year • Remote/Hybrid • Boston
C#
Node JS
TypeScript
JavaScript
Node JS
Nest.JS
Databases
Apache Kafka
DynamoDB
Redis
AI/ML
Claude
Claude Code
Cursor
DevOps
CI/CD
Apply
This is one of many
368,291 more open roles from verified company boards, updated every day.

