{"id":1780007,"url":"https://alion.io/job/vserve-ebusiness-solutions-india-private-limited-python-data-extraction-engineer","title":"Python Data Extraction Engineer","company":{"id":3801527,"name":"Vserve Ebusiness Solutions","domain":"vservesolution.com","url":"https://alion.io/company/vserve-ebusiness-solutions-india-private-limited","size_band":"201-500","is_staffing_agency":false,"employer_type":"direct","is_intermediary":false,"listed_via":null,"ats_vendor":"Zoho Recruit","truth_index":null},"role":"Backend","role_family":"Backend","seniority":"middle","employment_type":"full_time","work_mode":"on_site","remote_scope":null,"remote_scope_basis":null,"remote_working_hours":null,"hiring_geo_confidence":"structured","locations":["Coimbatore, India"],"countries":["IN"],"hiring_countries":[],"hiring_countries_total":0,"salary":null,"salary_estimate":null,"experience_years_min":3,"visa_sponsorship":false,"relocation_package":false,"has_equity":false,"technologies":[{"name":"Beautiful Soup","optional":false},{"name":"ETL/ELT","optional":false},{"name":"Git","optional":false},{"name":"JavaScript","optional":false},{"name":"Pandas","optional":false},{"name":"Playwright","optional":false},{"name":"Python","optional":false},{"name":"Rest API","optional":false},{"name":"Selenium","optional":false},{"name":"SQL","optional":false}],"status":"live","first_seen_at":"2026-10-03T17:14:13Z","employer_posted_date":"2026-10-03","last_verified_at":"2026-10-05T22:44:37Z","board_verified":true,"closed_at":null,"days_open":2,"trust":{"level":"ok","repost_count":null,"flags":[],"days_open":2},"description":"Python Data Extraction Engineer - Web Scraping & Government Data\nLocation: Coimbatore\nExperience: 3-7 years\nRole Type: Full-time\nAbout the Role\nWe are building a data intelligence platform that looks to analyze fragmented public information into structured, actionable business data.\nWe are looking for a strong Python Data Extraction Engineer who can discover, extract, clean, normalize and integrate data from government portals, public websites, PDFs, APIs and other open data sources.\nThis is not a conventional application-development role. The ideal candidate enjoys solving difficult data-acquisition problems involving poorly structured websites, inconsistent government portals, changing schemas, PDFs, JavaScript-rendered pages and large volumes of semi-structured information.\nWhat You Will Own\nYou will build and maintain the data-acquisition layer of the platform.\nKey responsibilities include:\nIdentify and evaluate government and public data sources relevant to property, businesses and commercial activity.\nBuild Python-based crawlers and extraction pipelines for government portals and public websites.\nExtract structured information from HTML pages, tables, PDFs, downloadable files and publicly accessible APIs.\nWork with JavaScript-rendered websites and multi-step public search interfaces.\nAutomate recurring extraction from multiple sources while respecting applicable access rules, rate limits, and terms.\nClean, standardize and normalize inconsistent data from different government authorities.\nPerform entity matching and record linkage across datasets using fields such as owner name, company name, address, survey number, coordinates and project information.\nDevelop mechanisms to detect website/schema changes and extraction failures.\nBuild validation and QA processes to measure completeness and accuracy.\nStore extracted information in structured databases and expose clean datasets to downstream applications.\nWork closely with GIS, product and engineering teams to combine location-based signals with government/public records.\nResearch new public and open-data sources that can improve the accuracy and completeness of our intelligence.\nExamples of Data Sources\nThe work may involve sources such as:\nMunicipal corporation portals\nState planning and development authorities\nDTCP and similar planning authorities\nRERA databases\nBuilding and planning permission records\nLand and property records\nTender and procurement portals\nCompany/business registries\nGovernment open-data portals\nEnvironmental and regulatory approvals\nPublic notices and downloadable government documents\nMaps and geospatial datasets\nOther legally accessible public and open-source information\nRequired Technical Skills\nStrong hands-on experience with:\nPython\nWeb scraping and crawling\nRequests / HTTP clients\nBeautifulSoup / lxml\nSelenium and/or Playwright\nREST APIs and JSON\nHTML/XML parsing\nPandas\nSQL\nData cleaning and transformation\nRegex and text processing\nETL/data pipelines\nGit\nWhat We Are Looking For\nWe particularly want someone who is a problem solver rather than simply a Python programmer.\nThe candidate should be able to investigate the available sources, understand how the underlying website works, determine the best extraction approach, build the pipeline and validate the resulting data.\nIdeal Background\nCandidates may come from backgrounds such as:\nWeb scraping / data extraction companies\nAlternative-data companies\nPropTech / real-estate data companies\nMarket-intelligence companies\nOSINT/data intelligence companies\nGovernment-data projects\nData aggregation platforms\nLead/data enrichment companies\nGIS/location-intelligence companies\nSuccess in This Role\nWithin the first few months, the successful candidate should be able to:\nMap relevant government/public data sources.\nBuild reliable extraction pipelines across multiple portals.\nConvert fragmented information into standardized records.\nCross-reference records from multiple sources.\nEstablish automated QA and monitoring.\nContinuously discover additional datasets that improve our product's coverage and accuracy.\nThe objective is not simply to scrape websites. It is to build a scalable public-data acquisition and enrichment engine that becomes a core component of our intelligence platform.\nRequirements\nPython, pandas, webscraping, web crawling,selenium postman","description_format":"text","description_chars":4348,"description_truncated":false,"requirements":{"experience_years_min":3,"management_years_min":null,"team_size_min":null,"manages_managers":false,"education":null,"security_clearance":false,"languages":[]},"benefits":[],"hiring_locations":[],"hiring_excludes":[],"relocation_offered":false,"industries":["E-commerce Marketing","Data Entry & Processing"],"lifecycle":[{"event":"open","at":"2026-10-03T17:14:13Z"}],"visa":[],"liveness":{"score":99,"band":"hot","label":"Hiring now","p_open":1,"p_active":0.993,"p_room":1,"age_days":1,"expected_fill_days":37,"reasons":["conf:1","urgency","velocity","win:early"],"computed_at":"2026-10-05T05:45:15Z"},"pay":null,"html_url":"https://alion.io/job/vserve-ebusiness-solutions-india-private-limited-python-data-extraction-engineer","json_url":"https://alion.io/job/vserve-ebusiness-solutions-india-private-limited-python-data-extraction-engineer.json","meta":{"generated_at":"2026-10-06T00:47:45Z","cache_seconds":300,"methodology":"https://alion.io/methodology","terms":"https://alion.io/terms","contact":"https://alion.io/contact","api":"https://alion.io/developers","about":"Alion is a live layer of people, companies and AI agents: who they are, whether they are real and active right now, what they do and how to work with them, readable by people and by agents and paid per call.","catalog":"https://alion.io/catalog.json","usage":{"tier":"crawler","counted_by":"address","units_charged":1,"used_today":1119,"day_limit":5000,"remaining_today":3881,"minute_limit":60,"resets_at":"2026-10-07T00:00:00Z"}},"offers":[{"id":"company.slices","title":"One company in depth, by slice","status":"live","price":{"credits":0.02,"usd":0.002,"plus_per_slice":{"credits":0.05,"usd":0.005}},"unit":"per company, plus each slice with data","note":"the employer in depth","call":{"mcp_tool":"get_company","arguments":{"id":3801527},"rest":"https://alion.io/mcp/rest/get_company?id=3801527"},"human":"https://alion.io/catalog?offer=company.slices&for=job%2Fvserve-ebusiness-solutions-india-private-limited-python-data-extraction-engineer"},{"id":"market.stats","title":"A market slice: pay, demand and time to fill","status":"live","price":{"credits":1,"usd":0.1},"unit":"per slice","note":"pay, demand and time to fill for this role and place","call":{"mcp_tool":"market_stats"},"human":"https://alion.io/catalog?offer=market.stats&for=job%2Fvserve-ebusiness-solutions-india-private-limited-python-data-extraction-engineer"},{"id":"job.search","title":"Open jobs by role, technology, place, pay and visa","status":"live","price":{"credits":0.02,"usd":0.002},"unit":"per posting in a list","note":"similar open postings","call":{"mcp_tool":"search_jobs"},"human":"https://alion.io/catalog?offer=job.search&for=job%2Fvserve-ebusiness-solutions-india-private-limited-python-data-extraction-engineer"},{"id":"company.verify","title":"Is this company real and active right now","status":"pilot","price":null,"unit":"per company","request":{"url":"https://alion.io/catalog/request","method":"POST","body":"{\"offer\": \"company.verify\", \"for\": \"job/vserve-ebusiness-solutions-india-private-limited-python-data-extraction-engineer\", \"note\": \"what you need it for\"}"},"human":"https://alion.io/catalog?offer=company.verify&for=job%2Fvserve-ebusiness-solutions-india-private-limited-python-data-extraction-engineer"}]}