{"id":1252802,"url":"https://alion.io/job/zoop-one-senior-data-scientist","title":"Senior Data Scientist","company":{"id":3798387,"name":"Zoop.one","domain":"zoop.one","url":"https://alion.io/company/zoop-one","size_band":null,"is_staffing_agency":false,"employer_type":"direct","is_intermediary":false,"listed_via":null,"ats_vendor":null,"truth_index":null},"role":"Data Science","role_family":"Data Science","seniority":"senior","employment_type":null,"work_mode":"on_site","remote_scope":null,"remote_scope_basis":null,"remote_working_hours":null,"hiring_geo_confidence":"structured","locations":["Pune, India"],"countries":["IN"],"hiring_countries":[],"hiring_countries_total":0,"salary":null,"salary_estimate":{"min_usd":20000,"max_usd":41000,"period":"year","method":"role_seniority_country_remote_cell","sample_n":54},"experience_years_min":6,"visa_sponsorship":false,"relocation_package":false,"has_equity":false,"technologies":[{"name":"BigQuery","optional":false},{"name":"Computer Vision","optional":false},{"name":"Dask","optional":false},{"name":"Deck.gl","optional":false},{"name":"Feature Store","optional":false},{"name":"GDAL","optional":false},{"name":"Git","optional":false},{"name":"Google BigQuery","optional":false},{"name":"LightGBM","optional":false},{"name":"Machine Learning","optional":false},{"name":"NLP","optional":false},{"name":"Plotly","optional":false},{"name":"PostGIS","optional":false},{"name":"Python","optional":false},{"name":"Spark","optional":false},{"name":"SQL","optional":false},{"name":"Streamlit","optional":false},{"name":"Time Series Forecasting","optional":false},{"name":"XGBoost","optional":false},{"name":"JavaScript","optional":true},{"name":"PostgreSQL","optional":true}],"status":"live","first_seen_at":"2026-09-12T09:26:32Z","employer_posted_date":null,"last_verified_at":"2026-09-12T09:26:32Z","board_verified":false,"closed_at":null,"days_open":23,"trust":{"level":"not_scored","repost_count":null,"flags":[],"days_open":23},"description":"About the role :\n\nWe are scaling our geospatial data-science capability - turning multi-source data (mobility, points of interest, demographic datasets, transaction signals, satellite imagery) into validated location attributes at fine spatial granularity (grid- and address-level) and powering ML models that are served as real-time APIs. Think: uncovering \"the why behind the where.\"\n\nYou will own the core data-science work: engineering location features from heterogeneous sources, building geospatial ML models (site selection, sales forecasting, catchment and propensity), cleaning and fusing data, and working with our full-stack engineer to ship models and attributes into production APIs.\n\nThis is a build role, not a maintenance role. You will help define how our attribute layer, modelling approach, and feature store come together.\n\nWhat you'll do :\n\n- Engineer location attributes from heterogeneous sources (mobility/smartphone, POI, demographic, web listed, transactions, satellite) at grid and address-level granularity.\n\n- Build, validate, and productionize geospatial ML models: site scoring, demand/sales forecasting, trade-area and catchment analysis, consumer/segment propensity.\n\n- Design data-quality pipelines that detect and correct bias, anomalies, and missing data, and keep attributes fresh and accurate.\n\n- Establish spatially-aware validation (avoiding spatial leakage) so models generalise across cities and geographies.\n\n- Partner with the full-stack engineer to expose models and attributes as real-time, address-level APIs and contribute to a reusable feature store.\n\n- Translate ambiguous business questions (site selection, expansion, risk) into modellable problems and defensible insights for enterprise clients.\n\nPrimary skills :\n\n- Geospatial data science: Strong command of spatial concepts and workflows: coordinate systems/projections, spatial joins, grid/indexing systems (H3, geohash, S2), spatial statistics (spatial autocorrelation / Moran's I), and catchment/trade-area analysis.\n\n- Python geospatial stack: Hands-on with GeoPandas, Shapely, Rasterio, GDAL/OGR, and spatial SQL via PostGIS (or BigQuery GIS). Comfortable manipulating vector and raster data at scale.\n\n- Machine learning for tabular/spatial problems: Solid grounding in regression and classification, gradient boosted trees (XGBoost/LightGBM), and feature selection, applied to problems like site scoring, demand/sales forecasting, and propensity.\n\n- Large-scale feature engineering: Ability to design and generate location attributes from heterogeneous raw sources, and to reason about a feature store - versioning, reuse, freshness - as the backbone of the work.\n\n- Data fusion, hygiene, and geocoding: Integrating messy, heterogeneous datasets; imputation, anomaly/bias detection, deduplication and entity resolution; robust geocoding and address/lat-long normalisation.\n\n- Performance & spatial-query optimisation: Processing very large point/grid datasets efficiently: spatial indexing (R-tree / GiST), optimised spatial joins, partitioning, query-plan diagnosis, and geometry simplification - to control runtime and cost.\n\n- Big-data and pipeline fluency: Advanced SQL plus distributed processing for large spatial workloads (Spark or Dask), and building reliable, repeatable data pipelines.\n\n- Productionizing models: Experience turning models into deployable, real-time APIs in collaboration with engineering - clean, tested, well-documented code (Git) and an understanding of latency, monitoring, and reproducibility.\n\nSecondary skills :\n\n- Mobility & foot-traffic analytics: Working with smartphone/mobility data for catchment, footfall, and movement patterns.\n\n- NLP for unstructured/web-listed data: Extracting structure from text-based sources (listings, reviews, POI descriptions).\n\n- Geospatial visualisation: kepler.gl, deck.gl, Plotly, or Streamlit for interactive, map-based storytelling and internal tooling.\n\n- Statistics & causal inference: Econometrics, uplift/causal methods, forecasting (time series).\n\n- Remote sensing / satellite imagery: Computer vision on imagery (CNNs), Google Earth Engine, land-use classification, building-footprint extraction, NDVI/change detection.\n\n- Domain knowledge: Retail/CPG site selection, BFSI credit risk / NPA reduction, e-commerce, or insurance use cases.\n\n- Stakeholder communication & B2B product sense: Explaining models and trade-offs to non-technical enterprise clients and shaping the product.\n\nQualifications :\n\n- 6+ years of applied data-science experience, with at least one project involving geospatial or location data end-to-end.\n\n- Degree in Computer Science, Statistics, Geoinformatics/GIS, Physics, Engineering, or a related quantitative field - or equivalent demonstrable experience.\n\nSkills\nPython, SQL, Machine Learning, Spark, BigQuery, Data Science, Git, Geo Spatial","description_format":"text","description_chars":4861,"description_truncated":false,"requirements":{"experience_years_min":6,"management_years_min":null,"team_size_min":null,"manages_managers":false,"education":null,"security_clearance":false,"languages":[]},"benefits":[],"hiring_locations":[],"hiring_excludes":[],"relocation_offered":false,"industries":["Identity Management","Information Security","Human Resources"],"lifecycle":[{"event":"open","at":"2026-09-25T18:04:06Z"}],"visa":[],"liveness":{"score":43,"band":"fade","label":"Fading","p_open":0.85,"p_active":0.673,"p_room":0.75,"age_days":22,"expected_fill_days":23,"reasons":["seen:22","velocity","win:late"],"computed_at":"2026-10-05T05:45:15Z"},"pay":null,"html_url":"https://alion.io/job/zoop-one-senior-data-scientist","json_url":"https://alion.io/job/zoop-one-senior-data-scientist.json","meta":{"generated_at":"2026-10-06T00:20:35Z","cache_seconds":300,"methodology":"https://alion.io/methodology","terms":"https://alion.io/terms","contact":"https://alion.io/contact","api":"https://alion.io/developers","about":"Alion is a live layer of people, companies and AI agents: who they are, whether they are real and active right now, what they do and how to work with them, readable by people and by agents and paid per call.","catalog":"https://alion.io/catalog.json","usage":{"tier":"crawler","counted_by":"address","units_charged":1,"used_today":407,"day_limit":5000,"remaining_today":4593,"minute_limit":60,"resets_at":"2026-10-07T00:00:00Z"}},"offers":[{"id":"company.slices","title":"One company in depth, by slice","status":"live","price":{"credits":0.02,"usd":0.002,"plus_per_slice":{"credits":0.05,"usd":0.005}},"unit":"per company, plus each slice with data","note":"the employer in depth","call":{"mcp_tool":"get_company","arguments":{"id":3798387},"rest":"https://alion.io/mcp/rest/get_company?id=3798387"},"human":"https://alion.io/catalog?offer=company.slices&for=job%2Fzoop-one-senior-data-scientist"},{"id":"market.stats","title":"A market slice: pay, demand and time to fill","status":"live","price":{"credits":1,"usd":0.1},"unit":"per slice","note":"pay, demand and time to fill for this role and place","call":{"mcp_tool":"market_stats"},"human":"https://alion.io/catalog?offer=market.stats&for=job%2Fzoop-one-senior-data-scientist"},{"id":"job.search","title":"Open jobs by role, technology, place, pay and visa","status":"live","price":{"credits":0.02,"usd":0.002},"unit":"per posting in a list","note":"similar open postings","call":{"mcp_tool":"search_jobs"},"human":"https://alion.io/catalog?offer=job.search&for=job%2Fzoop-one-senior-data-scientist"},{"id":"company.verify","title":"Is this company real and active right now","status":"pilot","price":null,"unit":"per company","request":{"url":"https://alion.io/catalog/request","method":"POST","body":"{\"offer\": \"company.verify\", \"for\": \"job/zoop-one-senior-data-scientist\", \"note\": \"what you need it for\"}"},"human":"https://alion.io/catalog?offer=company.verify&for=job%2Fzoop-one-senior-data-scientist"}]}