{"id":1197406,"url":"https://alion.io/job/shopee-ai-data-engineer-marketplace-intelligence-data-2027-graduate","title":"AI Data Engineer , Marketplace Intelligence & Data (2027 Graduate)","company":{"id":49909,"name":"Shopee","domain":"shopee.com","url":"https://alion.io/company/shopee","size_band":null,"is_staffing_agency":false,"is_intermediary":false,"listed_via":null,"ats_vendor":"Career site","truth_index":null},"role":"Data Science","role_family":"Data Science","seniority":"junior","employment_type":null,"work_mode":"on_site","remote_scope":null,"remote_scope_basis":null,"remote_working_hours":null,"hiring_geo_confidence":"structured","locations":["Singapore"],"countries":["SG"],"hiring_countries":[],"hiring_countries_total":0,"salary":null,"salary_estimate":{"min_usd":52000,"max_usd":148000,"period":"year","method":"global_role_cell_scaled_by_country","sample_n":276},"experience_years_min":null,"visa_sponsorship":false,"relocation_package":false,"has_equity":false,"technologies":[{"name":"CLIP","optional":false},{"name":"Data Augmentation","optional":false},{"name":"Flink","optional":false},{"name":"Function Calling","optional":false},{"name":"Java","optional":false},{"name":"Knowledge Distillation","optional":false},{"name":"LLM","optional":false},{"name":"Model Distillation","optional":false},{"name":"Multimodal AI","optional":false},{"name":"Python","optional":false},{"name":"Ray","optional":false},{"name":"Spark","optional":false},{"name":"SQL","optional":false},{"name":"Tool Use","optional":false},{"name":"VLM","optional":false},{"name":"Dimensional Modeling","optional":true},{"name":"Langfuse","optional":true}],"status":"live","first_seen_at":"2026-09-24T19:12:45Z","employer_posted_date":"2026-09-24","last_verified_at":"2026-09-24T23:13:44Z","board_verified":true,"closed_at":null,"days_open":0,"trust":{"level":"ok","repost_count":null,"flags":[],"days_open":0},"description":"In a data-driven and evaluation-driven manner, build an efficient closed loop for data iteration and establish an end-to-end data system spanning data sourcing, labeling, processing, synthesis, and evaluation. Continuously build high-quality datasets and evaluation sets to keep improving foundation model capabilities and to drive the development of AI models and applications.\nResponsibilities cover one or more of the following directions:\nDesign and implement high-performance, scalable, and distributed data infrastructure covering the full lifecycle - data storage, ingestion, cleaning, labeling, management, and analysis - and continuously improve data engineering efficiency.\nDesign audio-visual multimodal training data strategies; develop efficient data processing, synthesis, and optimization operators and pipelines, and build a multimodal data asset repository to meet the data needs of large model development.\nBuild a \"data-model-evaluation\" closed loop together with Agents, using data to drive rapid iteration of large models.\nTrack cutting-edge techniques and methods in the large-model data domain, explore innovative approaches such as data augmentation, data distillation, and high-quality data filtering, and land them in real business scenarios to increase data value.\nBachelor's degree or above in Computer Science, Software Engineering, Data Science, Statistics, or a related field.\nSolid programming fundamentals; proficient in Python and Java, competent in SQL; strong command of common data structures and algorithms.\nFamiliar with at least one big data processing framework (any of Spark / Flink / Ray), or strong self-learning ability backed by relevant coursework / projects.\nFamiliar with the fundamentals of large models / multimodal / AIGC (LLM, Diffusion, T2V, CLIP / VLM, etc.), or having relevant side projects.\nFamiliar with the fundamentals of Agents (Tool Use, trajectory, Memory, multi-turn interaction), or having worked on Agent-related projects.\nGood \"data sense\": able to identify data quality issues and willing to be rigorous about data accuracy.\nGood communication skills and a collaborative mindset.\nGood To Have:\n Experience with large-scale data processing (from coursework / internships / competitions), having handled TB-PB scale data. Familiarity with data warehouse dimensional modeling (Kimball / OneData approach).\nExperience with model evaluation / benchmarking (exposure to VBench, ELO, GSB, Langfuse, etc. is a plus).","description_format":"text","description_chars":2477,"description_truncated":false,"requirements":{"experience_years_min":null,"management_years_min":null,"team_size_min":null,"manages_managers":false,"education":{"level":"bachelor","optional":false},"security_clearance":false,"languages":[]},"benefits":[],"hiring_locations":[],"hiring_excludes":[],"relocation_offered":false,"industries":["Marketplaces","Commerce"],"lifecycle":[{"event":"open","at":"2026-09-24T19:12:45Z"}],"liveness":{"score":86,"band":"hot","label":"Hiring now","p_open":1,"p_active":0.86,"p_room":1,"age_days":0,"expected_fill_days":14,"reasons":["conf:1","win:early","comp:junior,brand"],"computed_at":"2026-09-25T00:19:54Z"},"pay":null,"html_url":"https://alion.io/job/shopee-ai-data-engineer-marketplace-intelligence-data-2027-graduate","json_url":"https://alion.io/job/shopee-ai-data-engineer-marketplace-intelligence-data-2027-graduate.json","meta":{"generated_at":"2026-09-25T00:19:54Z","cache_seconds":300,"methodology":"https://alion.io/methodology","terms":"https://alion.io/terms","contact":"https://alion.io/contact","api":"https://alion.io/developers","usage":{"tier":"crawler","counted_by":"address","units_charged":1,"used_today":319,"day_limit":5000,"remaining_today":4681,"minute_limit":60,"resets_at":"2026-09-26T00:00:00Z"}}}