As an SDE-2 Data Engineer, you will play a key role in building and scaling the foundational data systems that power Glance's next generation of AI products.
We are looking for a highly skilled and hands-on Data Engineer to own the design, development, and operation of large-scale data platforms and capabilities.
You will work closely with Applied Scientists, ML Engineers, Product Managers, Analytics teams, and Platform Engineers to build reliable data pipelines, feature stores, identity systems, catalog infrastructure, and self-service capabilities that accelerate experimentation and machine learning development.
You will be expected to independently drive complex technical initiatives from design through production while maintaining high standards of quality, scalability, and operational excellence.
The candidate will have responsibilities across the following functions:
Data Platform Development:
- Design and build scalable batch and real-time data pipelines using Spark, Flink, Kafka, and Airflow.
- Develop data products that support analytics, experimentation, recommendation systems, personalization, and AI applications.
- Build and maintain highly reliable ETL/ELT frameworks processing billions of events and catalog updates.
User Data Platform:
- Develop systems for user identity resolution and cross-surface signal aggregation across Mobile, TV, OEM, and Commerce ecosystems.
- Build datasets and services that support user profiling, audience creation, segmentation, and personalization.
- Contribute to deterministic and probabilistic identity stitching frameworks.
Commerce Catalog Platform:
- Build ingestion and enrichment pipelines for affiliate feeds, merchant catalogs, D2C integrations, and product metadata.
- Design scalable schemas and taxonomy frameworks for large and evolving commerce catalogs.
- Develop catalog quality, deduplication, normalization, and enrichment systems.
Feature Store and ML Enablement:
- Build reusable feature generation frameworks for ML and recommendation systems.
- Create low-latency feature pipelines serving training and online inference workloads.
- Partner with Applied Scientists to improve feature discoverability, governance, and reusability.
AI-Powered Engineering Capabilities:
- Develop internal AI-powered tools, agents, and self-service platforms that improve developer productivity.
- Build solutions for: Pipeline debugging, Data quality triage, SQL generation and optimization, Metadata discovery, Schema change analysis, Cost optimization recommendations.
Reliability and Operational Excellence:
- Own production services and pipelines with strong SLAs.
- Build observability into every layer through monitoring, lineage, alerting, reconciliation, and quality checks.
- Participate in incident response, root-cause analysis, and operational reviews.
- Continuously improve platform reliability, performance, and cost efficiency.
Expectations:
- Accelerate AI InnovationYou will enable faster experimentation and model deployment by building trusted, reusable data assets and feature pipelines.
- Power Personalized ExperiencesYour systems will help create a unified understanding of users across multiple surfaces, enabling highly personalized commerce experiences.
- Improve Platform Reliability - You will build observability-first infrastructure that ensures data quality, lineage, and trust across the ecosystem.
- Scale Commerce Intelligence - Your work will transform fragmented commerce and engagement signals into a strategic advantage for Glance's AI-powered commerce platform.
- Increase Engineering VelocityThrough automation, self-service capabilities, and AI-assisted workflows, you will reduce operational overhead and accelerate development cycles.
Requirements:
- 3 - 5.5 years of experience in Data Engineering, Distributed Systems, or Data Platform development.
- Strong experience owning large-scale production systems end-to-end.
- Data Engineering Expertise: Strong hands-on experience with Apache Spark, Kafka, Flink, Airflow, Distributed Data Processing, Batch and Streaming Architectures.
Data Modeling:
- Strong understanding of dimensional modeling, data warehousing, and large-scale schema design.
- Experience managing complex datasets and evolving schemas.
Data Quality and Observability:
- Experience with data validation frameworks, Lineage systems, Monitoring and alerting, Reconciliation pipelines, and CI/CD for data systems.
Cloud and Platform Engineering:
- Experience with GCP, Databricks, BigQuery, Infrastructure as Code, Cluster management, and Performance tuning and cost optimization.
Software Engineering:
- Strong programming skills in Python, Scala or Java, SQL
- Strong understanding of system design, Distributed systems, Performance optimization, and reliability engineering
Commerce Domain Experience:
- Experience working with: Product catalogs, Affiliate commerce platforms, Merchant feeds, Search and recommendation systems.
Identity and Personalization:
- Experience with: Identity resolution, Audience platforms, Customer 360 systems, and User profiling and segmentation.
Feature Stores and ML Platforms:
- Experience building: Feature stores, Training data pipelines, Real-time inference data systems, MLOps infrastructure.
AI-Assisted Engineering:
- Exposure to LLM-powered developer tools, Agents and copilots, Metadata intelligence systems.
- Automated debugging and remediation workflows.
What Success Looks Like in 12 Months:
- Built and scaled multiple production-grade pipelines powering personalization and commerce intelligence.
- Reduced data quality incidents through automated observability and reconciliation frameworks.
- Delivered reusable feature generation capabilities adopted by Applied Science teams.
- Improved platform efficiency through workload optimization and infrastructure cost reduction.
- Developed self-service capabilities that significantly improve productivity for data consumers and ML teams.
What This Role Is Not:
- Not a pure ETL developer role focused only on pipeline implementation.
- Not a pure platform operations role.
- Not a people-management role.

