{"id":1374447,"url":"https://alion.io/job/resources-valley-lead-data-architect","title":"Lead Data Architect","company":{"id":3801855,"name":"Resources Valley","domain":"resourcesvalley.in","url":"https://alion.io/company/resources-valley","size_band":null,"is_staffing_agency":false,"employer_type":"direct","is_intermediary":false,"listed_via":null,"ats_vendor":null,"truth_index":null},"role":"Data Science","role_family":"Data Science","seniority":"lead","employment_type":null,"work_mode":"on_site","remote_scope":null,"remote_scope_basis":null,"remote_working_hours":null,"hiring_geo_confidence":"structured","locations":["Jaipur, India","Indore, India","India"],"countries":["IN"],"hiring_countries":[],"hiring_countries_total":0,"salary":null,"salary_estimate":{"min_usd":21000,"max_usd":38000,"period":"year","method":"role_seniority_country_remote_cell","sample_n":16},"experience_years_min":8,"visa_sponsorship":false,"relocation_package":false,"has_equity":false,"technologies":[{"name":"Amazon Redshift","optional":false},{"name":"AWS","optional":false},{"name":"Azure","optional":false},{"name":"BigQuery","optional":false},{"name":"Dagster","optional":false},{"name":"Dask","optional":false},{"name":"Dimensional Modeling","optional":false},{"name":"GCP","optional":false},{"name":"Google BigQuery","optional":false},{"name":"Google Cloud Spanner","optional":false},{"name":"LLM","optional":false},{"name":"PostgreSQL","optional":false},{"name":"Prefect","optional":false},{"name":"Python","optional":false},{"name":"Spark","optional":false}],"status":"live","first_seen_at":"2026-09-28T05:12:40Z","employer_posted_date":null,"last_verified_at":"2026-09-28T05:12:40Z","board_verified":false,"closed_at":null,"days_open":2,"trust":{"level":"not_scored","repost_count":null,"flags":[],"days_open":2},"description":"Role Overview :\n\nWe are looking for a Lead Data Architect to design, build, and scale our data pipelines and entity resolution systems. This role combines deep technical expertise in data engineering with hands-on experience in AI-assisted tooling, entity matching, and data integration from diverse sources. You will lead architectural decisions for our data platform, mentor engineers, and ensure our pipelines are reliable, scalable, and production-grade.\n\nKey Responsibilities :\n\n- Architect, build, and maintain robust, scalable data pipelines that ingest, transform, and serve data from multiple internal and external sources.\n\n- Own the end-to-end orchestration of data workflows using tools like Dagster, ensuring observability, reliability, and maintainability of pipelines.\n\n- Design and implement entity resolution workflows - including matching, merging, and survivorship logic - using tools such as Splink, to produce clean, deduplicated, golden records.\n\n- Build and maintain web scrapers to source data from external providers, ensuring resilience to source changes, rate limits, and data quality issues.\n\n- Integrate and reconcile data coming from multiple, often inconsistent, sources into unified, trustworthy datasets.\n\n- Design and maintain data models and schemas across transactional and analytical systems, ensuring consistency, scalability, and performance.\n\n- Leverage AI/LLM-based tools and techniques to enhance data pipeline capabilities - e.g., intelligent data extraction, automated data quality checks, or AI-assisted entity matching.\n\n- Define and enforce best practices around pipeline design, testing, monitoring, and documentation.\n\n- Collaborate closely with data engineers, product managers, and other stakeholders to translate business requirements into scalable data architecture.\n\n- Provide technical leadership and mentorship to the data engineering team.\n\nRequired Skills & Experience :\n\n- Strong hands-on experience building and maintaining production-grade data pipelines at scale.\n\n- Practical experience with Dagster (or similar orchestration tools like Airflow/Prefect) for pipeline orchestration.\n\n- Experience with Splink or similar probabilistic/deterministic record linkage tools for entity matching, merging, and survivorship.\n\n- Strong proficiency in Python, including experience writing and maintaining web scrapers.\n\n- Proven experience integrating and maintaining data pipelines that pull from multiple, heterogeneous data sources.\n\n- Experience applying AI/ML tools within data engineering workflows (e.g., LLM-assisted data cleaning, extraction, or matching).\n\n- Hands-on experience with relational and distributed databases such as PostgreSQL and Google Cloud Spanner.\n\n- Strong understanding of data modeling principles (normalization, dimensional modeling, schema design) across OLTP and OLAP systems.\n\n- Experience with cloud data warehousing platforms such as BigQuery, Redshift, and cloud platforms (GCP/AWS/Azure).\n\n- Strong communication skills and experience working cross-functionally with engineering and product teams.\n\n- Experience with distributed data processing frameworks (e.g., Spark, Dask).\n\n- Familiarity with data governance, lineage, and cataloging tools.\n\n- Prior experience in a lead or architect-level role guiding a data engineering team.\n\nWhat We're Looking For :\n\nA technically strong, hands-on leader who can balance architectural thinking with the practical grit of debugging a flaky scraper or tuning a matching algorithm - someone who's comfortable owning both the big picture and the messy details of real-world data.\n\nSkills\nData Architect, Python, Data Engineering, Data Modeling, Data Pipeline, Data Observability, OLTP, OLAP, PostgreSQL, Data Warehousing","description_format":"text","description_chars":3749,"description_truncated":false,"requirements":{"experience_years_min":8,"management_years_min":null,"team_size_min":null,"manages_managers":false,"education":null,"security_clearance":false,"languages":[]},"benefits":[],"hiring_locations":[],"hiring_excludes":[],"relocation_offered":false,"industries":[],"lifecycle":[{"event":"open","at":"2026-09-28T06:01:22Z"}],"liveness":{"score":90,"band":"hot","label":"Hiring now","p_open":1,"p_active":0.903,"p_room":1,"age_days":2,"expected_fill_days":23,"reasons":["seen:2","velocity","win:early"],"computed_at":"2026-09-30T05:45:00Z"},"pay":null,"html_url":"https://alion.io/job/resources-valley-lead-data-architect","json_url":"https://alion.io/job/resources-valley-lead-data-architect.json","meta":{"generated_at":"2026-10-01T04:34:41Z","cache_seconds":300,"methodology":"https://alion.io/methodology","terms":"https://alion.io/terms","contact":"https://alion.io/contact","api":"https://alion.io/developers","usage":{"tier":"crawler","counted_by":"address","units_charged":1,"used_today":3964,"day_limit":5000,"remaining_today":1036,"minute_limit":60,"resets_at":"2026-10-02T00:00:00Z"}}}