{"id":1945611,"url":"https://alion.io/job/pattern-senior-data-engineer-data-and-analytics","title":"Senior Data Engineer - Data and Analytics","company":{"id":9124,"name":"Pattern","domain":"pattern.com","url":"https://alion.io/company/pattern","size_band":"1001-5000","is_staffing_agency":false,"employer_type":"direct","is_intermediary":false,"listed_via":null,"ats_vendor":"Lever","truth_index":{"grade":"B","score":79,"open_postings":78,"ghost_share":0.269,"stale_share":0,"repost_share":0,"time_to_fill_p50_days":90,"computed_at":"2026-10-10T05:45:15Z"}},"role":"Data Science","role_family":"Data Science","seniority":"senior","employment_type":"full_time","work_mode":"hybrid","remote_scope":null,"remote_scope_basis":null,"remote_working_hours":null,"hiring_geo_confidence":"structured","locations":["Pune, India"],"countries":["IN"],"hiring_countries":[],"hiring_countries_total":0,"salary":null,"salary_estimate":{"min_usd":20000,"max_usd":41000,"period":"year","method":"role_seniority_country_remote_cell","sample_n":56},"experience_years_min":4,"visa_sponsorship":false,"relocation_package":false,"has_equity":false,"technologies":[{"name":"Amazon Redshift","optional":false},{"name":"Amazon S3","optional":false},{"name":"Amazon SageMaker","optional":false},{"name":"AWS","optional":false},{"name":"AWS Lambda","optional":false},{"name":"BigQuery","optional":false},{"name":"Cassandra","optional":false},{"name":"dbt","optional":false},{"name":"DynamoDB","optional":false},{"name":"ElasticSearch","optional":false},{"name":"Feature Store","optional":false},{"name":"Google BigQuery","optional":false},{"name":"Great Expectations","optional":false},{"name":"IAM","optional":false},{"name":"Machine Learning","optional":false},{"name":"Presto","optional":false},{"name":"Python","optional":false},{"name":"Snowflake","optional":false},{"name":"Spark","optional":false},{"name":"SQL","optional":false},{"name":"Time Series Forecasting","optional":false},{"name":"Trino","optional":false}],"status":"live","first_seen_at":"2026-10-06T05:12:48Z","employer_posted_date":"2026-10-06","last_verified_at":"2026-10-11T22:18:12Z","board_verified":true,"closed_at":null,"days_open":5,"trust":{"level":"ok","repost_count":0,"flags":["company_stale"],"days_open":5},"description":"What makes this role different\nMost data engineering ends at a table. Pattern's ad-tech output leaves the warehouse and spends a client's advertising budget within the hour. A silently wrong join or an unguarded backfill is a customer-facing incident, not a dashboard discrepancy - so correctness, idempotency and data-quality gating are the job, not paperwork after the job.\nThe system you'll work on\nDestiny is Pattern's automated Ads optimizer. Once a day it discovers the keywords worth buying for every eligible product, assembles a wide feature store from performance, bid-history and search-results data, runs 5 machine-learning models, and picks the bid level that hits each product group's return on ad spend (ROAS) and budget target. A second pipeline then pushes those campaign, keyword and budget edits to the marketplace Ads API every 15 minutes.\nIt is a large, opinionated data system: a roughly 17,000-line orchestrated SQL codebase, a feature and label store several hundred columns wide, 5 model training and batch-scoring jobs, and blocking data-quality gates in front of every outward write. You would be one of the engineers who owns it end to end.\nRoles and Responsibilities\nDevelop, deploy, and support automated, scalable batch data pipelines from a variety of sources into the lakehouse.\n\nOwn and extend Airflow orchestration for a multi-DAG, cross-triggered daily pipeline and a 15-minute action pipeline - including branching, parallel task groups, cross-DAG triggers, backfill and full-refresh paths, and safe reruns.\n\nWrite and tune large analytical SQL: multi-hundred-column joins, window functions, incremental merges, and the warehouse-sizing and query-profile work needed to keep a daily run inside its window and its budget.\n\nExtend the feature store - add new features and labels, wire them through the join layer, and preserve the leakage and data-completeness conventions that make the models trainable.\n\nOrchestrate model training and batch inference on SageMaker from Airflow: build training and scoring datasets, manage S3 and Parquet round-trips, containerized training images, instance sizing, and loading predictions and metrics back into the warehouse.\n\nDevelop and implement data auditing strategies and processes to ensure data quality - including blocking data-quality checks in front of outward writes - and set thresholds that catch bad data without needlessly halting live bidding.\n\nIdentify and resolve problems in large-scale data processing workflows; maintain pipeline processes and troubleshoot failures, including on-call triage when a run breaks before market open.\n\nGuard the safety properties of an outward-writing system: idempotency, new-data detection, action validation and invalidation, and audit trails for every change pushed to marketplace.\n\nCollaborate with data scientists, advertising strategists, and platform teams to specify data requirements and provide access to data.\n\nTranslate business and analytics requirements - ROAS targets, budget pacing, playbook rules, branded versus non-branded strategy - into a comprehensive data model and pipelines.\n\nFoster data expertise and own data quality for assigned areas of ownership; work with data infrastructure to triage issues and drive to resolution.\n\nMentor and provide technical direction to other data engineers, and review their SQL and DAG changes.\n\nWhat \"basics of machine learning\" means here\nYou are not expected to invent model architectures - data scientists own the modeling. You are expected to be a competent, unsupervised partner to them, which means being able to:\nBuild training and evaluation datasets correctly - train/test splits over time, holdout windows, and a working instinct for target leakage in rolling-window features.\n\nReason about class imbalance and resampling (many keyword-hours have no clicks), and about clamping or bounding predictions before they drive a bid.\n\nRead regression metrics - MAE, RMSE, MAPE, WMAPE - plus feature importances, and tell “the model got worse” apart from “the upstream data got worse”.\n\nOperate the model lifecycle: retraining cadence, hyperparameters as configuration, prediction and metric persistence, validation tables, and drift monitoring.\n\nUnderstand how model outputs compose into a decision - here, predicted clicks, conversion rate, cost per click and basket revenue combining into an expected ROAS per bid, net of cannibalization.\n\nRequired qualifications\nBachelor's degree in Data Science, Data Analytics, Information Management, Computer Science, Information Technology, a related field, or equivalent professional experience.\n\n4+ years of overall professional experience.\n\n4+ years of hands-on experience with SQL and Python, including advanced SQL - window and analytic functions, complex joins, incremental merges, and query tuning.\n\n3+ years building production data pipelines on modern data architectures, with real ownership of scheduling, dependencies, retries and backfills, at scale and across many source systems.\n\n2+ years working with cloud data warehouses such as Snowflake, Redshift or BigQuery.\n\nProduction experience with a workflow orchestrator - Airflow strongly preferred - including debugging failed runs in a live system.\n\nExperience orchestrating ML training and batch inference from a scheduler, on SageMaker or an equivalent platform.\n\nWorking knowledge of applied machine learning fundamentals as described above: dataset construction, leakage, evaluation metrics, and model lifecycle operations.\n\nComfort with AWS - at minimum S3 and IAM - and with columnar file formats.\n\nDemonstrated ownership of data quality: testing, monitoring, alerting, and root-cause analysis on pipelines other people depend on.\n\nExcellent software engineering and scripting practice - version control, code review, modular and reviewable changes.\n\nStrong communication skills, in both presentation and comprehension, with the aptitude for cross-collaboration across data management, data science and analytics domains.\n\nAbility to lead and mentor a team of data engineers.\n\nPreferred Qualification\nExperience with digital advertising, bidding or auction systems - Amazon Ads, Google Ads, or a demand-side platform.\n\nAdvanced Snowflake - streams and tasks, stored procedures, UDFs, clustering, cost and performance tuning.\n\nExperience with time-series data and forecasting, and with hourly or day-parted grains.\n\nBackground in big data, non-relational databases, machine learning or data mining.\n\nExperience with data-quality frameworks such as Soda, Great Expectations or dbt tests.\n\nExperience with open-source and distributed data platforms: Spark, Hive, Trino/Presto, Cassandra, DynamoDB or Elasticsearch.\n\nBroader cloud experience: SNS, SQS, SES, Lambda, Glue, ECR and containerized workloads.\n\nExpertise in data governance.\n\nExperience working productively with AI coding agents on a large existing codebase.\n\nYour First 90 days\nDays 1-30 - Read the pipeline end to end and shadow a daily run. Ship small SQL and DAG fixes, take your first on-call triage with support, and be able to explain how a bid becomes an edit on marketplace.\n\nDays 31-60 - Own a stage. Add features to the feature store and wire them through, tune a slow task that threatens the run window, and add or re-threshold a data-quality check that catches something real.\n\nDays 61-90 - Lead a change that spans the pipeline and the action layer - a new signal, a new playbook rule, or a reliability improvement - with the tests, monitoring and rollback story that make it safe to leave running.\n\nWhy Pattern?\nThe company is a rocket ship experiencing phenomenal growth\n\nWe have tailwinds and a long runway; we're barely scratching the surface\n\nWe have big opportunities that will get you energized and excited\n\nGreat benefits including time off, insurance, competitive pay\n\nPattern provides equal employment opportunities to all employees and applicants for employment and prohibits discrimination and harassment of any type without regard to race, color, religion, age, sex, national origin, disability, genetic information, protected veteran status, sexual orientation, gender identity or expression, or any other characteristic protected by federal, state, or local laws.","description_format":"text","description_chars":8245,"description_truncated":false,"requirements":{"experience_years_min":4,"management_years_min":null,"team_size_min":null,"manages_managers":false,"education":{"level":"bachelor","optional":false},"security_clearance":false,"languages":[]},"benefits":[],"hiring_locations":[{"name":"India","iso":"IN","kind":"country"}],"hiring_excludes":[],"relocation_offered":false,"industries":["Artificial Intelligence","Commerce","Marketplaces","Industrial Real Estate"],"lifecycle":[{"event":"open","at":"2026-10-06T06:38:05Z"}],"visa":[],"liveness":{"score":86,"band":"hot","label":"Hiring now","p_open":1,"p_active":0.86,"p_room":1,"age_days":4,"expected_fill_days":90,"reasons":["conf:1","win:early"],"computed_at":"2026-10-10T05:45:15Z"},"pay":null,"html_url":"https://alion.io/job/pattern-senior-data-engineer-data-and-analytics","json_url":"https://alion.io/job/pattern-senior-data-engineer-data-and-analytics.json","meta":{"generated_at":"2026-10-11T23:44:30Z","cache_seconds":300,"methodology":"https://alion.io/methodology","terms":"https://alion.io/terms","contact":"https://alion.io/contact","api":"https://alion.io/developers","about":"Alion is a live layer of people, companies and AI agents: who they are, whether they are real and active right now, what they do and how to work with them, readable by people and by agents and paid per call.","catalog":"https://alion.io/catalog.json","usage":{"tier":"crawler_verified","counted_by":"address","units_charged":1,"used_today":14926,"day_limit":null,"remaining_today":null,"minute_limit":300,"resets_at":"2026-10-12T00:00:00Z"}},"offers":[{"id":"company.slices","title":"One company in depth, by slice","status":"live","price":{"credits":0.02,"usd":0.002,"plus_per_slice":{"credits":0.05,"usd":0.005}},"unit":"per company, plus each slice with data","note":"the employer in depth","call":{"mcp_tool":"get_company","arguments":{"id":9124},"rest":"https://alion.io/mcp/rest/get_company?id=9124"},"human":"https://alion.io/catalog?offer=company.slices&for=job%2Fpattern-senior-data-engineer-data-and-analytics"},{"id":"market.stats","title":"A market slice: pay, demand and time to fill","status":"live","price":{"credits":1,"usd":0.1},"unit":"per slice","note":"pay, demand and time to fill for this role and place","call":{"mcp_tool":"market_stats"},"human":"https://alion.io/catalog?offer=market.stats&for=job%2Fpattern-senior-data-engineer-data-and-analytics"},{"id":"job.search","title":"Open jobs by role, technology, place, pay and visa","status":"live","price":{"credits":0.02,"usd":0.002},"unit":"per posting in a list","note":"similar open postings","call":{"mcp_tool":"search_jobs"},"human":"https://alion.io/catalog?offer=job.search&for=job%2Fpattern-senior-data-engineer-data-and-analytics"},{"id":"company.verify","title":"Is this company real and active right now","status":"pilot","price":null,"unit":"per company","request":{"url":"https://alion.io/catalog/request","method":"POST","body":"{\"offer\": \"company.verify\", \"for\": \"job/pattern-senior-data-engineer-data-and-analytics\", \"note\": \"what you need it for\"}"},"human":"https://alion.io/catalog?offer=company.verify&for=job%2Fpattern-senior-data-engineer-data-and-analytics"}]}