{"id":1515407,"url":"https://alion.io/job/hcltech-senior-apache-spark-technical-lead-scala-python","title":"Senior Apache Spark Technical Lead - Scala, Python","company":{"id":225,"name":"HCLTech","domain":"hcltech.com","url":"https://alion.io/company/hcltech","size_band":"5000+","is_staffing_agency":false,"employer_type":"direct","is_intermediary":false,"listed_via":null,"ats_vendor":"SuccessFactors","truth_index":null},"role":"Backend","role_family":"Backend","seniority":"lead","employment_type":null,"work_mode":"on_site","remote_scope":null,"remote_scope_basis":null,"remote_working_hours":null,"hiring_geo_confidence":"inferred","locations":[],"countries":[],"hiring_countries":[],"hiring_countries_total":0,"salary":null,"salary_estimate":null,"experience_years_min":12,"visa_sponsorship":false,"relocation_package":false,"has_equity":false,"technologies":[{"name":"Azure","optional":false},{"name":"GraphQL","optional":false},{"name":"Spark","optional":false},{"name":"SQL","optional":false},{"name":"Apache Kafka","optional":true},{"name":"AWS","optional":true},{"name":"Azure Cosmos DB","optional":true},{"name":"Bamboo","optional":true},{"name":"Bitbucket","optional":true},{"name":"CI/CD","optional":true},{"name":"Db2","optional":true},{"name":"Delta Lake","optional":true},{"name":"ETL/ELT","optional":true},{"name":"Flink","optional":true},{"name":"Git","optional":true},{"name":"Hadoop","optional":true},{"name":"Jenkins","optional":true},{"name":"Oracle","optional":true},{"name":"Python","optional":true},{"name":"Scala","optional":true}],"status":"live","first_seen_at":"2026-09-30T09:23:05Z","employer_posted_date":"2026-09-30","last_verified_at":"2026-10-01T06:19:46Z","board_verified":true,"closed_at":null,"days_open":1,"trust":{"level":"ok","repost_count":null,"flags":[],"days_open":1},"description":"Job Summary\nIn this role, you will support the Data Engineering team at Prudential Singapore in setting up the Data Lake on Cloud and the implementation of standardized Data Model, single view of customer. You will develop data pipelines for new sources, data transformations within the Data Lake , implementing graphql , work on no sql database, CI/CDand data delivery as per the business requirements.\nKey Responsibilities\n\nJob Description:\nBuild pipelines to bring in wide variety of data from multiple sources within the organization as well as from social media and public data sources.\nCollaborate with cross functional teams to source data and make it available for downstream consumption.\nWork with the team to provide an effective solution design to meet business needs.\nEnsure regular communication with key stakeholders, understand any key concerns in how the initiative is being delivered or any risks/issues that have either not yet been identified or are not being progressed.\nEnsure dependencies and challenges (risks) are escalated and managed. Escalate critical issues to the Sponsor and/or Head of Data Engineering.\nEnsure timelines (milestones, decisions and delivery) are managed and value of initiative is achieved, without compromising quality and within budget.\nEnsure an appropriate and coordinated communications plan is in place for initiative execution and delivery, both internal and external.\nEnsure final handover of initiative to business-as-usual processes, carry out a post implementation review (as necessary) to ensure initiative objectives have been delivered, and any lessons learned are fed into future initiative management processes.\n\nSkill Requirements\n\nWho we are looking for:\nCompetencies & Personal Traits\nWork as a team player\nExcellent problem analysis skills\nExperience with at least one Cloud Infra provider (Azure/AWS)\nExperience in building data pipelines using batch processing with Apache Spark (Spark SQL, Dataframe API) or Hive query language (HQL)\nExperience in building streaming data pipeline using Apache Spark Structured Streaming or Apache Flink on Kafka & Delta Lake\nKnowledge of NOSQL databases. Good to have experience in Cosmos DB, Restful API’s and GraphQL\nKnowledge of Big data ETL processing tools, Data modelling and Data mapping.\nExperience with Hive and Hadoop file formats (Avro / Parquet / ORC)\nBasic knowledge of scripting (shell / bash)\nExperience of working with multiple data sources including relational databases (SQL Server / Oracle / DB2 / Netezza), NoSQL / document databases, flat files\nBasic understanding of CI CD tools such as Jenkins, JIRA, Bitbucket, Artifactory, Bamboo and Azure Dev-ops.\nBasic understanding of DevOps practices using Git version control\nAbility to debug, fine tune and optimize large scale data processing jobs\nWorking Experience\n12-15 years of broad experience of working with Enterprise IT applications in cloud platform and big data environments.\nProfessional Qualifications\nCertifications related to Data and Analytics would be an added advantage\n\nEducation\nMaster/bachelor’s degree in STEM (Science, Technology, Engineering, Mathematics)\n\nLanguage\nFluency in written and spoken English\n\nOther Requirements\n1.Relevant certifications in apache spark, scala, or python are a plus","description_format":"text","description_chars":3287,"description_truncated":false,"requirements":{"experience_years_min":12,"management_years_min":null,"team_size_min":null,"manages_managers":false,"education":{"level":"bachelor","optional":false},"security_clearance":false,"languages":[{"language":"English","level":"All levels","optional":false}]},"benefits":[],"hiring_locations":[],"hiring_excludes":[],"relocation_offered":false,"industries":["IT Consulting & Digital Transformation","IT Outsourcing & Dedicated Teams","Systems Integrators"],"lifecycle":[{"event":"open","at":"2026-09-30T09:23:05Z"}],"liveness":{"score":90,"band":"hot","label":"Hiring now","p_open":1,"p_active":0.903,"p_room":1,"age_days":0,"expected_fill_days":20,"reasons":["conf:10","velocity","win:early","comp:brand"],"computed_at":"2026-10-01T05:45:00Z"},"pay":null,"html_url":"https://alion.io/job/hcltech-senior-apache-spark-technical-lead-scala-python","json_url":"https://alion.io/job/hcltech-senior-apache-spark-technical-lead-scala-python.json","meta":{"generated_at":"2026-10-01T10:44:10Z","cache_seconds":300,"methodology":"https://alion.io/methodology","terms":"https://alion.io/terms","contact":"https://alion.io/contact","api":"https://alion.io/developers","usage":{"tier":"crawler","counted_by":"address","units_charged":1,"used_today":1915,"day_limit":5000,"remaining_today":3085,"minute_limit":60,"resets_at":"2026-10-02T00:00:00Z"}}}