{"id":1456382,"url":"https://alion.io/job/java-r-d-lead-data-analytics-engineer","title":"Lead Data & Analytics Engineer","company":{"id":3839609,"name":"Java R & D","domain":"enploy.in","url":"https://alion.io/company/java-r-and-d","size_band":null,"is_staffing_agency":true,"employer_type":"agency","is_intermediary":false,"listed_via":null,"ats_vendor":null,"truth_index":null},"role":"Data Science","role_family":"Data Science","seniority":"lead","employment_type":null,"work_mode":"on_site","remote_scope":null,"remote_scope_basis":null,"remote_working_hours":null,"hiring_geo_confidence":"structured","locations":["Delhi, India"],"countries":["IN"],"hiring_countries":[],"hiring_countries_total":0,"salary":null,"salary_estimate":{"min_usd":28000,"max_usd":49000,"period":"year","method":"role_seniority_country_remote_cell","sample_n":16},"experience_years_min":10,"visa_sponsorship":false,"relocation_package":false,"has_equity":false,"technologies":[{"name":"Airflow","optional":false},{"name":"Apache Kafka","optional":false},{"name":"Azure","optional":false},{"name":"CI/CD","optional":false},{"name":"Copilot","optional":false},{"name":"Databricks","optional":false},{"name":"dbt","optional":false},{"name":"Delta Lake","optional":false},{"name":"Dimensional Modeling","optional":false},{"name":"Flink","optional":false},{"name":"GitHub Actions","optional":false},{"name":"GitLab CI","optional":false},{"name":"Great Expectations","optional":false},{"name":"Kubernetes","optional":false},{"name":"Microsoft Fabric","optional":false},{"name":"Power BI","optional":false},{"name":"Presto","optional":false},{"name":"pySpark","optional":false},{"name":"Snowflake","optional":false},{"name":"Spark","optional":false},{"name":"SQL","optional":false},{"name":"Trino","optional":false},{"name":"Apache Hudi","optional":true},{"name":"Apache Iceberg","optional":true},{"name":"Dagster","optional":true},{"name":"Docker","optional":true},{"name":"Python","optional":true}],"status":"live","first_seen_at":"2026-09-29T09:53:39Z","employer_posted_date":null,"last_verified_at":"2026-09-29T09:53:39Z","board_verified":false,"closed_at":null,"days_open":2,"trust":{"level":"not_scored","repost_count":null,"flags":[],"days_open":2},"description":"Designation: Lead Data and Analytics Engineer\n\nPosition Overview:\n\nWe are seeking a hands-on Senior Data and Analytics professional who will design, build, and operate the enterprise data and analytics ecosystem covering:\n\n- Data Engineering and Data Platforms - open-source lakehouse stack (primary), with Fabric / Snowflake / Databricks as additional platforms\n\n- Business Intelligence (BI) and Analytics delivery\n\n- AI-ready data preparation and modeling\n\n- Embedded and platform-native AI capabilities within analytics tools\n\nThis role is focused on doing rather than managing - the individual will personally build pipelines, write transformation code, model data, create semantic layers, and develop BI dashboards. The platform foundation is open-source first: the candidate must be comfortable building and maintaining open-source data infrastructure (Apache Spark, Delta Lake, dbt, Airflow, Trino/Presto, Great Expectations and similar). Experience with managed platforms (Microsoft Fabric, Snowflake, Databricks) is a strong add-on but secondary to open-source depth.\n\nRequired Skills and Competencies:\n\nTechnical: Must-Have:\n\n- Open-source data engineering stack (hands-on, production-grade):\n\n1. Apache Spark / PySpark - personally wrote and optimised Spark jobs in production\n\n2. Delta Lake / Iceberg / Hudi - built and operated medallion-architecture lakes\n\n3. dbt - authored transformation models, tests and documentation\n\n4. Apache Airflow (or Prefect / Dagster) - built and maintained DAGs in production\n\n5. Docker - containerised data platform components; basic Kubernetes familiarity\n\n6. CI/CD for data pipelines - GitHub Actions / GitLab CI\n\n- Strong SQL - complex transformations, window functions, query optimisation.\n\n- Data modelling - dimensional modelling (star/snowflake schemas), semantic layer design.\n\n- Power BI - hands-on dashboard and semantic model development.\n\n- End-to-end ownership - personally built and maintained pipelines and BI artefacts in production, not just designed or reviewed them.\n\nTechnical - Strong Add-on:\n\n- Microsoft Fabric - Lakehouse, Delta tables, Direct Lake, Fabric Copilot\n\n- Snowflake - data loading, transformation, performance tuning\n\n- Azure Databricks - Delta Live Tables, Unity Catalog, Workflows\n\n- Apache Kafka / Flink - streaming ingestion\n\n- Trino / Presto - federated query\n\n- Great Expectations / Soda - open-source data quality frameworks\n\n- OpenMetadata / DataHub / Apache Atlas - data cataloguing and lineage\nSkills\nApache Spark, PySpark, Delta Lake, Apache Airflow, SQL, Power BI, Databricks, Data Build Tool, Data Modeling, Data Analytics","description_format":"text","description_chars":2616,"description_truncated":false,"requirements":{"experience_years_min":10,"management_years_min":null,"team_size_min":null,"manages_managers":false,"education":null,"security_clearance":false,"languages":[]},"benefits":[],"hiring_locations":[],"hiring_excludes":[],"relocation_offered":false,"industries":[],"lifecycle":[{"event":"open","at":"2026-09-29T10:02:18Z"}],"liveness":{"score":52,"band":"ok","label":"Likely open","p_open":1,"p_active":0.516,"p_room":1,"age_days":1,"expected_fill_days":24,"reasons":["seen:1","agency","win:early"],"computed_at":"2026-10-01T05:45:00Z"},"pay":null,"html_url":"https://alion.io/job/java-r-d-lead-data-analytics-engineer","json_url":"https://alion.io/job/java-r-d-lead-data-analytics-engineer.json","meta":{"generated_at":"2026-10-01T11:20:43Z","cache_seconds":300,"methodology":"https://alion.io/methodology","terms":"https://alion.io/terms","contact":"https://alion.io/contact","api":"https://alion.io/developers","usage":{"tier":"crawler","counted_by":"address","units_charged":1,"used_today":2930,"day_limit":5000,"remaining_today":2070,"minute_limit":60,"resets_at":"2026-10-02T00:00:00Z"}}}