1,456,966open jobs
88,711companies
228,521added this week
Browse all
Salary
≈ $20k – $41k per year (Estimated)
Location
Hybrid (Pune, India)
Seniority
Senior · 4+ years exp
Employment
Full-Time

Confirmed on the employer's own hiring board on Oct 11, 2026. First seen by Alion on Oct 6, 2026. Pattern scores B on the Alion truth index.

Overview
Company
Impact
Profile match
Pattern is an e-commerce accelerator that helps consumer brands sell on Amazon, Walmart, TikTok Shop and other online marketplaces worldwide. Founded in 2013, it buys and resells brand inventory or manages marketplace channels on brands' behalf, combining marketplace advertising, content, pricing, fulfillment and brand protection with its own data and software. The company is headquartered in Lehi, Utah, United States.

What makes this role different

Most data engineering ends at a table. Pattern's ad-tech output leaves the warehouse and spends a client's advertising budget within the hour. A silently wrong join or an unguarded backfill is a customer-facing incident, not a dashboard discrepancy - so correctness, idempotency and data-quality gating are the job, not paperwork after the job.

The system you'll work on

    Destiny is Pattern's automated Ads optimizer. Once a day it discovers the keywords worth buying for every eligible product, assembles a wide feature store from performance, bid-history and search-results data, runs 5 machine-learning models, and picks the bid level that hits each product group's return on ad spend (ROAS) and budget target. A second pipeline then pushes those campaign, keyword and budget edits to the marketplace Ads API every 15 minutes.

    It is a large, opinionated data system: a roughly 17,000-line orchestrated SQL codebase, a feature and label store several hundred columns wide, 5 model training and batch-scoring jobs, and blocking data-quality gates in front of every outward write. You would be one of the engineers who owns it end to end.

Roles and Responsibilities

    • Develop, deploy, and support automated, scalable batch data pipelines from a variety of sources into the lakehouse.

    • Own and extend Airflow orchestration for a multi-DAG, cross-triggered daily pipeline and a 15-minute action pipeline - including branching, parallel task groups, cross-DAG triggers, backfill and full-refresh paths, and safe reruns.

    • Write and tune large analytical SQL: multi-hundred-column joins, window functions, incremental merges, and the warehouse-sizing and query-profile work needed to keep a daily run inside its window and its budget.

    • Extend the feature store - add new features and labels, wire them through the join layer, and preserve the leakage and data-completeness conventions that make the models trainable.

    • Orchestrate model training and batch inference on SageMaker from Airflow: build training and scoring datasets, manage S3 and Parquet round-trips, containerized training images, instance sizing, and loading predictions and metrics back into the warehouse.

    • Develop and implement data auditing strategies and processes to ensure data quality - including blocking data-quality checks in front of outward writes - and set thresholds that catch bad data without needlessly halting live bidding.

    • Identify and resolve problems in large-scale data processing workflows; maintain pipeline processes and troubleshoot failures, including on-call triage when a run breaks before market open.

    • Guard the safety properties of an outward-writing system: idempotency, new-data detection, action validation and invalidation, and audit trails for every change pushed to marketplace.

    • Collaborate with data scientists, advertising strategists, and platform teams to specify data requirements and provide access to data.

    • Translate business and analytics requirements - ROAS targets, budget pacing, playbook rules, branded versus non-branded strategy - into a comprehensive data model and pipelines.

    • Foster data expertise and own data quality for assigned areas of ownership; work with data infrastructure to triage issues and drive to resolution.

    • Mentor and provide technical direction to other data engineers, and review their SQL and DAG changes.

What "basics of machine learning" means here

    You are not expected to invent model architectures - data scientists own the modeling. You are expected to be a competent, unsupervised partner to them, which means being able to:

    • Build training and evaluation datasets correctly - train/test splits over time, holdout windows, and a working instinct for target leakage in rolling-window features.

    • Reason about class imbalance and resampling (many keyword-hours have no clicks), and about clamping or bounding predictions before they drive a bid.

    • Read regression metrics - MAE, RMSE, MAPE, WMAPE - plus feature importances, and tell “the model got worse” apart from “the upstream data got worse”.

    • Operate the model lifecycle: retraining cadence, hyperparameters as configuration, prediction and metric persistence, validation tables, and drift monitoring.

    • Understand how model outputs compose into a decision - here, predicted clicks, conversion rate, cost per click and basket revenue combining into an expected ROAS per bid, net of cannibalization.

Required qualifications

    • Bachelor's degree in Data Science, Data Analytics, Information Management, Computer Science, Information Technology, a related field, or equivalent professional experience.

    • 4+ years of overall professional experience.

    • 4+ years of hands-on experience with SQL and Python, including advanced SQL - window and analytic functions, complex joins, incremental merges, and query tuning.

    • 3+ years building production data pipelines on modern data architectures, with real ownership of scheduling, dependencies, retries and backfills, at scale and across many source systems.

    • 2+ years working with cloud data warehouses such as Snowflake, Redshift or BigQuery.

    • Production experience with a workflow orchestrator - Airflow strongly preferred - including debugging failed runs in a live system.

    • Experience orchestrating ML training and batch inference from a scheduler, on SageMaker or an equivalent platform.

    • Working knowledge of applied machine learning fundamentals as described above: dataset construction, leakage, evaluation metrics, and model lifecycle operations.

    • Comfort with AWS - at minimum S3 and IAM - and with columnar file formats.

    • Demonstrated ownership of data quality: testing, monitoring, alerting, and root-cause analysis on pipelines other people depend on.

    • Excellent software engineering and scripting practice - version control, code review, modular and reviewable changes.

    • Strong communication skills, in both presentation and comprehension, with the aptitude for cross-collaboration across data management, data science and analytics domains.

    • Ability to lead and mentor a team of data engineers.

Preferred Qualification

    • Experience with digital advertising, bidding or auction systems - Amazon Ads, Google Ads, or a demand-side platform.

    • Advanced Snowflake - streams and tasks, stored procedures, UDFs, clustering, cost and performance tuning.

    • Experience with time-series data and forecasting, and with hourly or day-parted grains.

    • Background in big data, non-relational databases, machine learning or data mining.

    • Experience with data-quality frameworks such as Soda, Great Expectations or dbt tests.

    • Experience with open-source and distributed data platforms: Spark, Hive, Trino/Presto, Cassandra, DynamoDB or Elasticsearch.

    • Broader cloud experience: SNS, SQS, SES, Lambda, Glue, ECR and containerized workloads.

    • Expertise in data governance.

    • Experience working productively with AI coding agents on a large existing codebase.

Your First 90 days

    • Days 1-30 - Read the pipeline end to end and shadow a daily run. Ship small SQL and DAG fixes, take your first on-call triage with support, and be able to explain how a bid becomes an edit on marketplace.

    • Days 31-60 - Own a stage. Add features to the feature store and wire them through, tune a slow task that threatens the run window, and add or re-threshold a data-quality check that catches something real.

    • Days 61-90 - Lead a change that spans the pipeline and the action layer - a new signal, a new playbook rule, or a reliability improvement - with the tests, monitoring and rollback story that make it safe to leave running.

    • Why Pattern?

      • The company is a rocket ship experiencing phenomenal growth

      • We have tailwinds and a long runway; we're barely scratching the surface

      • We have big opportunities that will get you energized and excited

      • Great benefits including time off, insurance, competitive pay

      • Pattern provides equal employment opportunities to all employees and applicants for employment and prohibits discrimination and harassment of any type without regard to race, color, religion, age, sex, national origin, disability, genetic information, protected veteran status, sexual orientation, gender identity or expression, or any other characteristic protected by federal, state, or local laws.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
1,456,966 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Data Science
Similar stack
Same company
Pune
Data Scientist-3 1 day ago
≈ $17k – $41k per year (Estimated) • In office • Full-Time • 5+ years exp • Bengaluru
Python
Python
pySpark
AI/ML
Spark
Machine Learning
DevOps
AWS
Apply
≈ $27k – $53k per year (Estimated) • In office • 10+ years exp • Bachelor's Degree • Pune
SQL
Analytics
Power BI
Master Data Management
Management
Agile
Apply
≈ $21k – $43k per year (Estimated) • Hybrid • Full-Time • 5+ years exp • India
Python
SQL
Databases
Apache Iceberg
Apache Kafka
Azure Cosmos DB
AI/ML
Spark
Feature Store
DevOps
Rest API
Helm
Azure
GitOps
ArgoCD
Docker
Kubernetes
Platform Engineering
Bicep
Cybersecurity
Polaris
Analytics
Dimensional Modeling
Apply
AWS Data Architect 3 months ago
≈ $28k – $55k per year (Estimated) • In office • Full-Time • 10+ years exp • Kochi
Python
SQL
Python
pySpark
Databases
Databricks
Apache Kafka
Amazon Redshift
AI/ML
Spark
Amazon SageMaker
Machine Learning
DevOps
Terraform
CloudFormation
Azure
AWS
Amazon S3
Amazon Kinesis
Analytics
ETL/ELT
Data Vault
Collibra
Alation
Apply
≈ $22k – $45k per year (Estimated) • In office • Full-Time • 8+ years exp • Bachelor's Degree • Bengaluru
Databases
Databricks
AI/ML
Airflow
Analytics
Azure Data Factory
Apply
$176k – $235k per year • In office • Full-Time • 10+ years exp • High School Diploma • United States
Python
SQL
AI/ML
AI Agents
Machine Learning
Apply
≈ $53k – $123k per year (Estimated) • In office • Full-Time • 3+ years exp • Bachelor's Degree • Veldhoven
Python
SystemVerilog
DevOps
Git
Linux
Apply
≈ $33k – $52k per year (Estimated) • In office • Internship • Master's Degree • Veldhoven
Python
C++
Apply
≈ $62k – $113k per year (Estimated) • Hybrid • Full-Time • 2+ years exp • Bachelor's Degree • Wilton
Python
MATLAB
Apply
≈ $20k – $35k per year (Estimated) • In office • 8+ years exp • Mumbai
Python
Java
C#
C#
.NET
DevOps
Ansible
GCP
Azure
Jenkins
Git
AWS
Management
UiPath
Agile
Waterfall
Apply
Senior Data Engineer 5 months ago
≈ $111k – $212k per year (Estimated) • Hybrid • Full-Time • 7+ years exp • Bachelor's Degree • Lehi
Python
SQL
Databases
Snowflake
Apache Iceberg
Delta Lake
Apache Kafka
Google BigQuery
Amazon Redshift
Trino
BigQuery
AI/ML
Spark
Airflow
DevOps
Terraform
AWS
Kubernetes
Amazon EKS
AWS Lambda
Amazon Kinesis
Analytics
ETL/ELT
Dimensional Modeling
Apply
In office • Full-Time • Pune
Management
Asana
Apply
≈ $61k – $114k per year (Estimated) • In office • Full-Time • 3+ years exp • Las Vegas
Management
Microsoft Office
Apply
≈ $88k – $179k per year (Estimated) • Hybrid • Full-Time • 5+ years exp • Bachelor's Degree • Lehi
AI/ML
ChatGPT
Perplexity
Analytics
A/B Testing
Apply
≈ $65k – $140k per year (Estimated) • Hybrid • Full-Time • Melbourne
Marketing
Salesforce
HubSpot
Apply
≈ $20k – $41k per year (Estimated) • In office • 5+ years exp • Pune
SQL
Perl
Databases
Oracle
DevOps
Datadog
AWS
Unix
Apply
Data Engineer 1 day ago
≈ $15k – $36k per year (Estimated) • In office • 4+ years exp • Bachelor's Degree • Pune
DevOps
Azure
CI/CD
GitHub
Analytics
Azure Data Factory
Apply
≈ $19k – $39k per year (Estimated) • In office • 5+ years exp • Pune
SQL
Databases
MS SQL
DevOps
Datadog
Linux
Windows
Apply
Engineer 7 hours ago
≈ $13k – $32k per year (Estimated) • In office • 6+ years exp • Bachelor's Degree • Pune
Apply
≈ $24k – $54k per year (Estimated) • In office • Full-Time • 8+ years exp • Bachelor's Degree • Pune
JavaScript
TypeScript
SQL
C#
C#
.NET
Frontend
Angular
DevOps
Rest API
Azure
IAM
Cybersecurity
Active Directory
Management
Scrum
Apply
See all jobs
This is one of many
1,456,966 more open roles from verified company boards, updated every day.