{"id":1955671,"url":"https://alion.io/job/capgemini-data-engineer-27","title":"Data Engineer","company":{"id":145,"name":"Capgemini","domain":"capgemini.com","url":"https://alion.io/company/capgemini","size_band":"5000+","is_staffing_agency":false,"employer_type":"services","is_intermediary":false,"listed_via":null,"ats_vendor":"Career site","truth_index":{"grade":"B","score":81,"open_postings":1665,"ghost_share":0,"stale_share":0.74,"repost_share":0.001,"time_to_fill_p50_days":27,"computed_at":"2026-10-09T06:01:00Z"}},"role":"Data Science","role_family":"Data Science","seniority":"middle","employment_type":"full_time","work_mode":"on_site","remote_scope":null,"remote_scope_basis":null,"remote_working_hours":null,"hiring_geo_confidence":"structured","locations":["Cairo, Egypt"],"countries":["EG"],"hiring_countries":[],"hiring_countries_total":0,"salary":null,"salary_estimate":{"min_usd":42000,"max_usd":113000,"period":"year","method":"global_role_cell_scaled_by_country","sample_n":662},"experience_years_min":4,"visa_sponsorship":false,"relocation_package":false,"has_equity":false,"technologies":[{"name":"pySpark","optional":false},{"name":"Spark","optional":false},{"name":"Airflow","optional":true},{"name":"Apache Iceberg","optional":true},{"name":"Apache Kafka","optional":true},{"name":"Apache NiFi","optional":true},{"name":"Azure","optional":true},{"name":"CI/CD","optional":true},{"name":"ETL/ELT","optional":true},{"name":"Git","optional":true},{"name":"Informatica","optional":true},{"name":"Linux","optional":true},{"name":"Platform Engineering","optional":true},{"name":"Python","optional":true},{"name":"SQL","optional":true},{"name":"Teradata","optional":true}],"status":"live","first_seen_at":"2026-10-06T12:34:29Z","employer_posted_date":"2026-10-06","last_verified_at":"2026-10-09T21:50:49Z","board_verified":true,"closed_at":null,"days_open":3,"trust":{"level":"ok","repost_count":null,"flags":[],"days_open":3},"description":"Job Description\nWe are seeking a highly skilled and motivated Senior Data Engineer to join our Data & Analytics practice. The successful candidate will build and operate enterprise-scale data solutions using PySpark, Apache Spark and modern data engineering technologies, with a strong grounding in data warehousing, lakehouse architecture and large-scale data integration.\nThe target environment uses Cloudera Data Platform (CDP). Prior Cloudera experience is preferred but is not a mandatory requirement. Candidates with strong PySpark/Spark expertise and relevant experience on comparable enterprise data platforms will be considered, and Cloudera CDP upskilling will be provided to the selected candidate.\nThis role requires hands-on engineering capability, technical ownership and effective collaboration with architects, analysts, source-system teams and client stakeholders. Experience in banking and financial services is highly desirable.\nCritical hiring priority: Strong, production-grade PySpark/Spark engineering skills and the ability to rapidly learn and work effectively within the Cloudera CDP environment.\nKey Responsibilities\nData Engineering & Development\nDesign, develop, test and maintain scalable batch and streaming data pipelines using PySpark.\nBuild and optimize ETL/ELT processes for large-volume enterprise data workloads.\nDevelop reusable ingestion, transformation and validation frameworks.\nImplement data quality controls, reconciliation checks, monitoring and operational logging.\nSupport dimensional, warehouse, data lake and lakehouse data modeling activities.\nApply Spark performance optimization techniques including partitioning, join optimization, caching and handling data skew.\nCloudera Platform Engineering\nDevelop and run Spark workloads within the Cloudera Data Platform (CDP) environment.\nWork with components including Cloudera Data Engineering (CDE), Cloudera Data Warehouse (CDW), Hive, Impala, Ozone, Ranger, Atlas and NiFi.\nConfigure, execute, monitor and troubleshoot Spark workloads.\nPerformance-tune PySpark jobs, Spark configurations, SQL queries and Hive/Impala workloads.\nDiagnose platform, ingestion, transformation, storage and workload execution issues.\nApply security, access-control, metadata and lineage practices using Ranger and Atlas.\nCandidates without prior Cloudera experience will be expected to complete the required Cloudera enablement/upskilling and apply their existing Spark and data engineering knowledge to the platform.\nArchitecture, Governance & Delivery\nContribute to lakehouse implementations using modern table formats such as Apache Iceberg.\nApply engineering standards covering modular design, code quality, testing, version control and CI/CD.\nParticipate in requirements analysis, solution design, technical estimation and design reviews.\nSupport system integration testing, UAT, production deployment and operational readiness.\nProduce clear technical documentation and deliver structured knowledge-transfer sessions.\nCollaborate effectively with distributed, multicultural teams and technical and business stakeholders.\nCandidate Profile\nEducation & Experience\nBachelor's degree in Computer Science, Information Systems, Engineering or a related discipline.\nAt least 4 years of professional experience in data engineering.\nAt least 3 years of hands-on PySpark / Apache Spark development experience.\nExperience delivering large-scale enterprise data pipelines and ETL/ELT solutions.\nCloudera CDP experience is an advantage, but not mandatory. Strong candidates from comparable Spark-based data platforms are encouraged to apply.\nMandatory Technical Skills\nPySpark\nApache Spark\nPython\nAdvanced SQL and query optimization\nETL/ELT development\nData warehousing concepts\nLinux\nGit\nData modeling fundamentals\nExperience working with large-scale enterprise data platforms\nPreferred / Upskillable Skills\nCloudera CDP\nCloudera CDE and CDW\nHive and Impala\nRanger and Atlas\nApache Iceberg\nApache NiFi and Kafka\nApache Airflow\nDenodo and Informatica\nOracle and Teradata\nAzure data services\nDevOps and CI/CD pipelines\nCloudera CDP knowledge is preferred, not mandatory. The selected candidate will be provided with Cloudera enablement/upskilling where required.\nPreferred Domain Background\nBanking and financial services\nDigital banking platforms\nCustomer analytics\nRegulatory reporting\nEnterprise data warehousing\nData lakehouse implementation and migration programs\nBehavioural Competencies\nStrong analytical thinking and structured problem solving.\nOwnership mindset with the ability to work independently and deliver reliably.\nClear written and verbal communication with technical and non-technical stakeholders.\nAbility to mentor junior engineers and contribute to team capability building.\nAbility and willingness to quickly learn new data platforms and technologies.\nComfort working in distributed and multicultural delivery teams.","description_format":"text","description_chars":4908,"description_truncated":false,"requirements":{"experience_years_min":4,"management_years_min":null,"team_size_min":null,"manages_managers":false,"education":{"level":"bachelor","optional":false},"security_clearance":false,"languages":[]},"benefits":[],"hiring_locations":[],"hiring_excludes":[],"relocation_offered":false,"industries":["Cybersecurity","Government","Science & Engineering","Research Institutes"],"lifecycle":[{"event":"open","at":"2026-10-06T12:34:35Z"}],"visa":[],"liveness":{"score":40,"band":"fade","label":"Fading","p_open":1,"p_active":0.417,"p_room":1,"age_days":2,"expected_fill_days":27,"reasons":["conf:0","stale_co","evergreen","wave","velocity","win:early","comp:brand"],"computed_at":"2026-10-09T06:01:00Z"},"pay":null,"html_url":"https://alion.io/job/capgemini-data-engineer-27","json_url":"https://alion.io/job/capgemini-data-engineer-27.json","meta":{"generated_at":"2026-10-10T00:32:08Z","cache_seconds":300,"methodology":"https://alion.io/methodology","terms":"https://alion.io/terms","contact":"https://alion.io/contact","api":"https://alion.io/developers","about":"Alion is a live layer of people, companies and AI agents: who they are, whether they are real and active right now, what they do and how to work with them, readable by people and by agents and paid per call.","catalog":"https://alion.io/catalog.json","usage":{"tier":"crawler","counted_by":"address","units_charged":1,"used_today":818,"day_limit":5000,"remaining_today":4182,"minute_limit":60,"resets_at":"2026-10-11T00:00:00Z"}},"offers":[{"id":"company.slices","title":"One company in depth, by slice","status":"live","price":{"credits":0.02,"usd":0.002,"plus_per_slice":{"credits":0.05,"usd":0.005}},"unit":"per company, plus each slice with data","note":"the employer in depth","call":{"mcp_tool":"get_company","arguments":{"id":145},"rest":"https://alion.io/mcp/rest/get_company?id=145"},"human":"https://alion.io/catalog?offer=company.slices&for=job%2Fcapgemini-data-engineer-27"},{"id":"market.stats","title":"A market slice: pay, demand and time to fill","status":"live","price":{"credits":1,"usd":0.1},"unit":"per slice","note":"pay, demand and time to fill for this role and place","call":{"mcp_tool":"market_stats"},"human":"https://alion.io/catalog?offer=market.stats&for=job%2Fcapgemini-data-engineer-27"},{"id":"job.search","title":"Open jobs by role, technology, place, pay and visa","status":"live","price":{"credits":0.02,"usd":0.002},"unit":"per posting in a list","note":"similar open postings","call":{"mcp_tool":"search_jobs"},"human":"https://alion.io/catalog?offer=job.search&for=job%2Fcapgemini-data-engineer-27"},{"id":"company.verify","title":"Is this company real and active right now","status":"pilot","price":null,"unit":"per company","request":{"url":"https://alion.io/catalog/request","method":"POST","body":"{\"offer\": \"company.verify\", \"for\": \"job/capgemini-data-engineer-27\", \"note\": \"what you need it for\"}"},"human":"https://alion.io/catalog?offer=company.verify&for=job%2Fcapgemini-data-engineer-27"}]}