{"id":1035510,"url":"https://alion.io/job/sigma-software-group-senior-data-scientist-2","title":"Senior Data Scientist","company":{"id":6266,"name":"Sigma Software Group","domain":"sigma.software","url":"https://alion.io/company/sigma-software-group","size_band":"1001-5000","is_staffing_agency":false,"employer_type":"services","is_intermediary":false,"listed_via":null,"ats_vendor":"SmartRecruiters","truth_index":null},"role":"Data Science","role_family":"Data Science","seniority":"senior","employment_type":"full_time","work_mode":"remote","remote_scope":"stated_countries","remote_scope_basis":"board_field","remote_working_hours":null,"hiring_geo_confidence":"structured","locations":["Warsaw, Poland"],"countries":["PL"],"hiring_countries":["PL"],"hiring_countries_total":1,"salary":null,"salary_estimate":{"min_usd":66000,"max_usd":97000,"period":"year","method":"role_seniority_country_remote_cell","sample_n":60},"experience_years_min":5,"visa_sponsorship":false,"relocation_package":false,"has_equity":false,"technologies":[{"name":"A/B Testing","optional":false},{"name":"CatBoost","optional":false},{"name":"LightGBM","optional":false},{"name":"Machine Learning","optional":false},{"name":"NumPy","optional":false},{"name":"pySpark","optional":false},{"name":"Python","optional":false},{"name":"Scikit-learn","optional":false},{"name":"SHAP","optional":false},{"name":"Spark","optional":false},{"name":"SQL","optional":false},{"name":"Vertex AI","optional":false},{"name":"XGBoost","optional":false}],"status":"live","first_seen_at":"2026-09-11T17:43:08Z","employer_posted_date":"2026-09-11","last_verified_at":"2026-09-26T21:39:25Z","board_verified":true,"closed_at":null,"days_open":15,"trust":{"level":"ok","repost_count":null,"flags":[],"days_open":15},"description":"Join Sigma Software to help build advanced machine learning solutions for one of the large-scale players in the programmatic advertising ecosystem. We are looking for a Senior Machine Learning Engineer with strong production ML expertise and deep interest in real-time optimization systems, large-scale behavioral data, and AdTech challenges.\nIn this role, you will work with a dedicated Sigma Software team on a predictive modeling platform integrated with a live ad exchange processing hundreds of millions of auction requests daily. You will contribute to sophisticated ML solutions involving bid optimization, calibration, counterfactual evaluation, and constrained decision-making systems.\nWe as a company offer the opportunity to work on technically challenging products, collaborate with experienced engineers and data scientists, and make a direct impact on large-scale production systems.\nCUSTOMER\nOur Customer is a technology company operating supply-side infrastructure within the programmatic advertising ecosystem. The company manages a high-scale ad exchange platform and is investing in predictive decisioning capabilities to improve advertising performance, audience targeting, and campaign optimization through advanced machine learning technologies.\nPROJECT\nThe project focuses on building a predictive modeling and optimization platform on top of a live ad exchange environment. The platform evaluates and filters advertising supply in real time, predicts high-performing audience contexts, builds look-alike audiences from small seed datasets, and optimizes campaign performance across multiple business objectives and operational constraints.\nThe team works on complex machine learning challenges including censored bid-landscape modeling, sparse and delayed conversion attribution, calibration systems, counterfactual evaluation, and constrained optimization models. The solution is designed for large-scale production use and close collaboration with the Customer’s internal data science organization.\n Build and improve censored bid-landscape models to estimate clearing-price distributions from partially observed auction data\n Develop real-time win probability estimation models responsive to bid pricing dynamics\nDesign and implement hierarchical lift estimation models with confidence-bound-based selection strategies\nBuild conversion propensity models using sparse, delayed, and aggregate-only labels\nDevelop look-alike audience modeling approaches using positive-unlabeled learning and embedding-based nearest-neighbor techniques\nImplement advertiser-level calibration strategies while independently monitoring ranking and calibration quality\nDesign robust offline evaluation frameworks using inverse-propensity scoring, doubly-robust estimators, and importance reweighting\nDefine exploration strategies and propensity logging approaches to ensure reliable downstream correction and evaluation\nDevelop constrained optimization mechanisms for campaign objectives, pricing constraints, and volume targeting\nContribute to data diagnostics, capability assessments, and evidence-based model recommendations\nCollaborate with the Customer team during post-launch tuning and performance validation cycles\nPrepare technical documentation and knowledge transfer materials for the Customer’s internal data science team\nParticipate in architecture discussions and contribute to scalable ML platform design decisions\n 5+ years of experience in Machine Learning or Data Science with production-grade models measured against business KPIs\nStrong Python skills including numpy, pandas, and scikit-learn\nStrong SQL skills and experience working with large-scale datasets\nDeep practical experience with XGBoost, LightGBM, or CatBoost\nStrong understanding of regularization, calibration methods, and categorical feature handling\nStrong knowledge of probability, statistics, confidence intervals, and statistical power analysis\nExperience with feature engineering for structured and behavioral datasets\nHands-on experience with Spark or PySpark\nPractical knowledge of experimentation frameworks and A/B testing methodologies\nExperience with advanced validation approaches including temporal splits, leakage detection, drift analysis, and slice-based metrics\nUnderstanding of explainability techniques such as SHAP and permutation importance\nUpper-Intermediate English level or higher\nWILL BE A PLUS\nExperience in AdTech modeling including CTR/CVR prediction, bid-landscape modeling, audience segmentation, and RTB mechanics\nExperience working with sparse, delayed, or censored labels\nKnowledge of attribution modeling, survival analysis, and positive-unlabeled learning\nPractical experience with counterfactual and off-policy evaluation techniques\nUnderstanding of calibration methods including isotonic regression and Platt scaling\nExperience with hierarchical, empirical-Bayes, or partial-pooling models\nKnowledge of constrained or multi-objective optimization approaches\nExperience with uplift modeling and causal inference methods\nExperience with Vertex AI or similar managed ML training environments\nPublications, competitive modeling achievements, or open-source contributions related to Machine Learning or AdTech\n PERSONAL PROFILE\nStrong analytical and problem-solving skills\nAbility to work effectively in a highly data-driven environment\nStrong communication and stakeholder management abilities\nAbility to explain complex modeling decisions to technical and non-technical audiences\nProactive mindset with strong ownership mentality\nAttention to detail and scientific rigor in experimentation and evaluation","description_format":"text","description_chars":5627,"description_truncated":false,"requirements":{"experience_years_min":5,"management_years_min":null,"team_size_min":null,"manages_managers":false,"education":null,"security_clearance":false,"languages":[{"language":"English","level":"Upper-Intermediate (B2)","optional":false}]},"benefits":[],"hiring_locations":[{"name":"Poland","iso":"PL","kind":"country"}],"hiring_excludes":[],"relocation_offered":false,"industries":["IT Outsourcing & Dedicated Teams","Custom Software Development"],"lifecycle":[{"event":"open","at":"2026-09-18T16:14:44Z"}],"liveness":{"score":71,"band":"hot","label":"Hiring now","p_open":1,"p_active":0.791,"p_room":0.9,"age_days":14,"expected_fill_days":29,"reasons":["conf:5","velocity","win:mid"],"computed_at":"2026-09-26T05:45:00Z"},"pay":null,"html_url":"https://alion.io/job/sigma-software-group-senior-data-scientist-2","json_url":"https://alion.io/job/sigma-software-group-senior-data-scientist-2.json","meta":{"generated_at":"2026-09-27T02:38:00Z","cache_seconds":300,"methodology":"https://alion.io/methodology","terms":"https://alion.io/terms","contact":"https://alion.io/contact","api":"https://alion.io/developers","usage":{"tier":"crawler","counted_by":"address","units_charged":1,"used_today":2380,"day_limit":5000,"remaining_today":2620,"minute_limit":60,"resets_at":"2026-09-28T00:00:00Z"}}}