Salary
≈ $18k – $48k per year (Estimated)
Location
In office (Gurgaon)
Employment
Full-Time
First seen by Alion on Sep 28, 2026.
Overview
Company
Impact
Profile match
Hire top-tier tech talent across India with AI-powered precision. CubicAI delivers a 98% match rate, 4-day shortlisting, and verified profiles from a 1M+ talent pool. Post a job free.
Job Description
About the job team: product (fashion, taste, personalization) why this role exists polopan is building consumer ai where taste, context, and judgment matter more than raw scale. our models are only as good as the truthfulness of the data beneath them. this role exists to make sure our catalog and product data pipelines are: clean explainable trustworthy hard to lie to we care less about how fast things move, and more about whether they ever need to be questioned again. what you'll be responsible for 1. building catalog data pipelines design and maintain pipelines that ingest, normalize, enrich, and version product/catalog data define schemas that age well as the product evolves handle messy, incomplete, and inconsistent data without hiding the mess make catalog data usable for downstream systems (search, recommendations, personalization) 2. owning data clarity end-to-end decide what should be logged and what should not ensure every dataset has a clear purpose and owner detect and debug silent failures, drift, and data pollution make pipelines observable, debuggable, and boring in the best way 3. making decisions irreversible build systems that allow the team to confidently: trust metrics kill features iterate without second-guessing the data reduce ambiguity for product and machine learning decisions, not add to it 4. setting engineering standards early establish patterns for data hygiene, versioning, and validation write documentation that explains why something exists, not just how push back on over-engineering and under-thinking equally what we care about (more than speed) we don't measure this role by: number of tickets closed lines of code written how fast you ship we measure it by: how much confusion disappears after your work exists how rarely your systems need revisiting how confidently others can build on top of what you've built sometimes deadlines will exist — not to rush you, but to force clarity on what truly matters. what we're looking for you'll likely resonate if you: enjoy turning messy reality into clean, minimal systems think deeply about schemas, contracts, and downstream consequences prefer deleting data to hoarding it care about correctness, not cleverness are calm under constraint and decisive under deadlines experience that helps (not all required): building data pipelines (etl / elt) in production environments working with catalog, marketplace, or content-heavy datasets designing event schemas and data contracts debugging data quality issues that don't throw errors familiarity with batch + near-real-time systems tech stack specifics matter less than your judgment. python would be nice to have. what this role is not not a ship fast, break things role not a model-training or research-heavy ML role not a growth or analytics-only role this is a foundational engineering role. What you build early will shape everything that comes after. how success looks (first 90 days) we trust our catalog data without caveats product and ML teams stop asking is this data right at least one major product decision becomes irreversible because of your work parts of the system become confidently deletable if that sounds like a good problem to work on, we'd like to talk. final note we're building this company deliberately. if you care more about clarity than velocity, and about doing things once, properly, you'll feel at home here.
Required Skills
[-]
Additional Information
NA
About the job team: product (fashion, taste, personalization) why this role exists polopan is building consumer ai where taste, context, and judgment matter more than raw scale. our models are only as good as the truthfulness of the data beneath them. this role exists to make sure our catalog and product data pipelines are: clean explainable trustworthy hard to lie to we care less about how fast things move, and more about whether they ever need to be questioned again. what you'll be responsible for 1. building catalog data pipelines design and maintain pipelines that ingest, normalize, enrich, and version product/catalog data define schemas that age well as the product evolves handle messy, incomplete, and inconsistent data without hiding the mess make catalog data usable for downstream systems (search, recommendations, personalization) 2. owning data clarity end-to-end decide what should be logged and what should not ensure every dataset has a clear purpose and owner detect and debug silent failures, drift, and data pollution make pipelines observable, debuggable, and boring in the best way 3. making decisions irreversible build systems that allow the team to confidently: trust metrics kill features iterate without second-guessing the data reduce ambiguity for product and machine learning decisions, not add to it 4. setting engineering standards early establish patterns for data hygiene, versioning, and validation write documentation that explains why something exists, not just how push back on over-engineering and under-thinking equally what we care about (more than speed) we don't measure this role by: number of tickets closed lines of code written how fast you ship we measure it by: how much confusion disappears after your work exists how rarely your systems need revisiting how confidently others can build on top of what you've built sometimes deadlines will exist — not to rush you, but to force clarity on what truly matters. what we're looking for you'll likely resonate if you: enjoy turning messy reality into clean, minimal systems think deeply about schemas, contracts, and downstream consequences prefer deleting data to hoarding it care about correctness, not cleverness are calm under constraint and decisive under deadlines experience that helps (not all required): building data pipelines (etl / elt) in production environments working with catalog, marketplace, or content-heavy datasets designing event schemas and data contracts debugging data quality issues that don't throw errors familiarity with batch + near-real-time systems tech stack specifics matter less than your judgment. python would be nice to have. what this role is not not a ship fast, break things role not a model-training or research-heavy ML role not a growth or analytics-only role this is a foundational engineering role. What you build early will shape everything that comes after. how success looks (first 90 days) we trust our catalog data without caveats product and ML teams stop asking is this data right at least one major product decision becomes irreversible because of your work parts of the system become confidently deletable if that sounds like a good problem to work on, we'd like to talk. final note we're building this company deliberately. if you care more about clarity than velocity, and about doing things once, properly, you'll feel at home here.
Required Skills
[-]
Additional Information
NA
Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
961,875 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Free forever. No card. Under a minute.
Your match
How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.
Recommended for you based on this role
Data Science
Similar stack
Same company
Gurgaon
Data Steward -RSD
1 hour ago
≈ $32k – $67k per year (Estimated) • In office • 5+ years exp • Bachelor's Degree • Pretoria
Analytics
Master Data Management
Apply
Senior Data Engineer
1 month ago
$106k – $180k per year • In office • Full-Time • 5+ years exp • Bachelor's Degree • Milwaukee
Python
SQL
Databases
Snowflake
AI/ML
Airflow
dbt
DevOps
Prometheus
CI/CD
Git
AWS
AWS Lambda
Cortex
Amazon S3
IAM
Analytics
ETL/ELT
Fivetran
Management
Agile
Apply
Data Mapper/Engineer
1 hour ago
≈ $21k – $42k per year (Estimated) • Remote (India) • 5+ years exp • Bengaluru
Databases
Oracle
Analytics
ETL/ELT
Apply
Data Architect
1 hour ago
$52k – $73k per year (gross) • Remote (India) • 12+ years exp • India
Python
SQL
Python
pySpark
Databases
Snowflake
AI/ML
Copilot
Cursor
Spark
ChatGPT
Claude Code
dbt
LLM
RAG
LLMOps
DevOps
AWS
SLI/SLO/SLA
Amazon S3
IAM
Analytics
Power BI
Data Vault
Dimensional Modeling
Apply
SQL-Entwickler (m/w/d)
1 month ago
$52k per year • In office • Full-Time • Vienna
SQL
PowerShell
Analytics
Power BI
SSIS
SSAS
Apply
Data & AI Engineer
1 day ago
In office • Full-Time • Phnom Penh
AI/ML
Prompt Engineering
LLM
RAG
Machine Learning
DevOps
Platform Engineering
Analytics
ETL/ELT
Apply
IT Integration Developer
1 day ago
≈ $119k – $219k per year (Estimated) • Hybrid • 7+ years exp • Bachelor's Degree • Carrollton
Analytics
ETL/ELT
Informatica
Apply
Solutions Architect
1 day ago
In office • 5+ years exp • Bachelor's Degree
Python
SQL
PowerShell
DevOps
Rest API
Azure
Cybersecurity
Microsoft Entra ID
Analytics
Power BI
ETL/ELT
Management
Power Automate
Power Apps
Agile
Scrum
Service Desk
Apply
Senior Scala Engineer
2 days ago
≈ $97k – $225k per year (Estimated) • Equity • Remote (likely Poland)
Python
Java
Scala
Databases
Apache Kafka
Google BigQuery
BigQuery
AI/ML
Spark
dbt
Machine Learning
DevOps
GCP
GitHub Actions
Consul
Azure
CI/CD
AWS
Docker
Kubernetes
Service Mesh
GitHub
Linux
Windows
Unix
Apply
Senior Scale Engineer
2 days ago
≈ $87k – $175k per year (Estimated) • Remote (likely Poland)
Python
Java
Scala
Databases
Apache Kafka
Google BigQuery
BigQuery
AI/ML
Spark
dbt
Machine Learning
DevOps
GCP
GitHub Actions
Consul
Azure
CI/CD
AWS
Docker
Kubernetes
Service Mesh
GitHub
Linux
Windows
Unix
Apply
Lead Data / Staff Data Engineer
2 days ago
≈ $28k – $50k per year (Estimated) • In office • Full-Time • 7+ years exp • Bengaluru
Python
DevOps
AWS
Apply
Manager, Data Engineer
2 days ago
≈ $29k – $51k per year (Estimated) • In office • Full-Time • 11+ years exp • Bengaluru
Python
SQL
Databases
Snowflake
Databricks
ClickHouse
AI/ML
Machine Learning
DevOps
GCP
Azure
AWS
Analytics
ETL/ELT
Apply
≈ $13k – $31k per year (Estimated) • In office • Full-Time • 5+ years exp • Gurgaon
SQL
Databases
PostgreSQL
AI/ML
Copilot
Claude
ChatGPT
AI Agents
Tabnine
DevOps
Git
GitHub
Management
Jira
Agile
Scrum
QA
Postman
Insomnia
Apply
In office • Full-Time • Gurgaon
Apply
Apply
≈ $13k – $30k per year (Estimated) • In office • 5+ years exp • Gurgaon
SQL
Databases
MS SQL
Azure SQL Database
DevOps
Terraform
Azure DevOps
GitHub Actions
Azure
CI/CD
Bicep
Cybersecurity
Microsoft Entra ID
Analytics
Power BI
Cryptography
Vault
Apply
≈ $41k – $86k per year (Estimated) • Remote (India) • Full-Time • Master's Degree • Bengaluru • Gurgaon • Hyderabad • Pune • Mumbai
Cybersecurity
ISO 27001
GDPR
Management
Agile
Apply
Sustainability Consultant II
1 day ago
≈ $15k – $37k per year (Estimated) • Hybrid • Full-Time • 6+ years exp • Master's Degree • Noida • Mumbai • Gurgaon
Design
AutoCAD
Management
Microsoft Office
Apply
Microsoft Fabric and Power BI Lead
1 day ago
Hybrid • Full-Time • 7+ years exp • Noida • Hyderabad • Gurgaon
Python
SQL
PowerShell
Databases
Microsoft Fabric
DevOps
Rest API
Azure
CI/CD
Git
Platform Engineering
Analytics
Power BI
Azure Data Factory
Management
ServiceNow
Apply
Senior Data Analyst
1 day ago
≈ $14k – $29k per year (Estimated) • Hybrid • Full-Time • Bachelor's Degree • Gurgaon
Python
SQL
Databases
Snowflake
AI/ML
Copilot
DevOps
AWS
Analytics
Tableau
Power BI
Apply
This is one of many
961,875 more open roles from verified company boards, updated every day.

