{"id":1438693,"url":"https://alion.io/job/allegro-data-scientist-consumer-data","title":"Data Scientist (Consumer Data)","company":{"id":1876195,"name":"Allegro","domain":"allegro.eu","url":"https://alion.io/company/allegro-eu","size_band":"1001-5000","is_staffing_agency":false,"employer_type":"direct","is_intermediary":false,"listed_via":null,"ats_vendor":"SuccessFactors","truth_index":null},"role":"Data Science","role_family":"Data Science","seniority":"middle","employment_type":null,"work_mode":"remote","remote_scope":"stated_countries","remote_scope_basis":"inferred_company_offices","remote_working_hours":null,"hiring_geo_confidence":"inferred","locations":[],"countries":[],"hiring_countries":["PL"],"hiring_countries_total":1,"salary":null,"salary_estimate":{"min_usd":77000,"max_usd":175000,"period":"year","method":"global_role_seniority_cell","sample_n":550},"experience_years_min":2,"visa_sponsorship":false,"relocation_package":false,"has_equity":false,"technologies":[{"name":"Consul","optional":false},{"name":"Docker","optional":false},{"name":"Fine-tuning","optional":false},{"name":"GCP","optional":false},{"name":"GitHub","optional":false},{"name":"GitHub Actions","optional":false},{"name":"Hybrid Search","optional":false},{"name":"Knowledge Graph","optional":false},{"name":"Kubernetes","optional":false},{"name":"Machine Learning","optional":false},{"name":"Python","optional":false},{"name":"Semantic Search","optional":false},{"name":"Service Mesh","optional":false},{"name":"Time Series Forecasting","optional":false},{"name":"Windows","optional":false}],"status":"live","first_seen_at":"2026-09-16T00:00:00Z","employer_posted_date":"2026-09-16","last_verified_at":"2026-09-29T09:58:52Z","board_verified":false,"closed_at":null,"days_open":15,"trust":{"level":"ok","repost_count":null,"flags":[],"days_open":15},"description":"Job Description\nIn the Consumer domain, we create and maintain customer-facing applications that help millions of clients each day to complete their purchases. We are looking for a mid-level Data Scientist to join our team and help move our search engine beyond simple keyword matching toward a Domain-Aware Hybrid Search (integrating LLMs, Semantic Search, and Knowledge Graphs). In this role, you will be responsible for supporting the development of the Central Query Understanding mechanism and implementing Predictive & Assistive Discovery to reduce the \"cost of search\" through automated intent recognition.\nIn your daily work, you will handle the following tasks:\nDesigning, developing, and deploying models that solve complex business problems, including predictive, segmentation, forecasting, and recommendation models.\nDeveloping query-to-category and query-to-product matching models to enhance the search experience.\nBuilding internal AI Engineering competence, including fine-tuning small, efficient models.\nTaking an active part in all Data Science project phases: from problem formulation and data exploration to modeling, automation, release, and monitoring.\nUsing a wide range of model types, including boosting, Bayesian methods, causal inference, optimization methods, deep learning, and forecasting models.\nCollaborating across teams, partnering with business stakeholders, analytics, and data engineering teams on experiment setup and deployment.\nEnsuring models reflect business processes' specificity and staying up-to-date with GenAI challenges.\nWe are looking for people with:\nHave a degree strongly related to statistical/mathematical modeling.\nPossess at least 2 years of experience in data analysis and building machine learning solutions that have been released to production.\nHave a strong understanding of statistical and machine learning methods, specifically for forecasting and decision tree-based algorithms.\nAre proficient in Python and efficient in using basic development tools.\n Can process massive datasets (terabytes of data) using Google Cloud Platform solutions, working with tabular, spatial, natural language, image, and time-series data.\nKnow English at a B2 level and Polish at a C1 level.\nDemonstrate core competencies in Analytical Thinking, Learning Agility, Cooperation, and Continuous Improvement & Innovation.\nWhat's in it for you:\nFlexible working hours in the hybrid model (4/1) - working hours start between 7:00 a.m. and 10:00 a.m. We also have 30 days of occasional remote work.\nAnnual bonus based on your annual performance and company results.\nWell-located offices (with e.g. fully equipped kitchens, bicycle parking, terraces full of greenery) and excellent work tools (e.g., raised desks, ergonomic chairs, interactive conference rooms).\nA 16\" or 14\" MacBook Pro or corresponding Dell with Windows (if you don't like Macs) and all the necessary accessories.\nA wide selection of fringe benefits in a cafeteria plan - you choose what you like (e.g., medical, sports or lunch packages, insurance, purchase vouchers).\nEnglish classes that we pay for related to the specific nature of your job.\nA training budget, inter-team tourism, hackathons, and an internal learning platform where you will find multiple trainings.\nAn additional day off for volunteering, which you can use alone, with a team, or with a larger group of people connected by a common goal.\nSocial events for Allegro people - Spin Kilometers, Family Day, Fat Thursday, Advent of Code, and many other occasions we enjoy.\nAnd that's just the beginning! You can read more about the benefits here.\n#goodtobehere means that:\nYou will join a team you can count on - we work with top-class specialists who have knowledge- and experience-sharing in their DNA.\nYou will love our level of autonomy in team organization, the space for continuous development, and the opportunity to try new things.\nYou get to choose which technology solves the problem and you are responsible for what you create.\nYou will value our Developer Experience and the full platform of tools and technologies that make creating software easier. We rely on an internal ecosystem based on self-service and widely used tools such as Kubernetes, Docker, Consul, GitHub, and GitHub Actions. Thanks to this, you can contribute to Allegro from your very first days on the job.\nYou will be equipped with modern AI tools to automate repetitive tasks, allowing you to focus on developing new services and refining existing ones (also leveraging AI support).\nYou will create solutions that will be used (and loved!) by your friends, family and millions of our customers.\nYou will meet the Allegro Scale, which starts with over 1000 microservices, an open-source data bus (Hermes) with 300K+ rps, a Service Mesh with 1M+ rps, tens of petabytes of data, and production-used machine learning.\nYou will become part of Allegro Tech - We speak at industry conferences, cooperate with tech communities, run our own blog (it's been over 10 years!), record podcasts, lead guilds, and we organize our own internal conference - the Allegro Tech Meeting. We create solutions we love (and can) to talk about!\nSend us your CV and... see you at Allegro!","description_format":"text","description_chars":5217,"description_truncated":false,"requirements":{"experience_years_min":2,"management_years_min":null,"team_size_min":null,"manages_managers":false,"education":null,"security_clearance":false,"languages":[{"language":"Polish","level":"All levels","optional":false},{"language":"English","level":"All levels","optional":false}]},"benefits":["Apple Macbook","Cafeteria","Flexible schedule"],"hiring_locations":[{"name":"Poland","iso":"PL","kind":"country"}],"hiring_excludes":[],"relocation_offered":false,"industries":["Commerce","Marketplaces"],"lifecycle":[{"event":"open","at":"2026-09-29T04:13:31Z"}],"liveness":{"score":69,"band":"ok","label":"Likely open","p_open":1,"p_active":0.767,"p_room":0.9,"age_days":15,"expected_fill_days":23,"reasons":["conf:43","velocity","win:mid","comp:brand"],"computed_at":"2026-10-01T05:45:00Z"},"pay":null,"html_url":"https://alion.io/job/allegro-data-scientist-consumer-data","json_url":"https://alion.io/job/allegro-data-scientist-consumer-data.json","meta":{"generated_at":"2026-10-01T12:18:39Z","cache_seconds":300,"methodology":"https://alion.io/methodology","terms":"https://alion.io/terms","contact":"https://alion.io/contact","api":"https://alion.io/developers","usage":{"tier":"crawler","counted_by":"address","units_charged":1,"used_today":3637,"day_limit":5000,"remaining_today":1363,"minute_limit":60,"resets_at":"2026-10-02T00:00:00Z"}}}