{"id":1115281,"url":"https://alion.io/job/jpmorganchase-data-engineer-iii-databricks-pyspark-python-aws","title":"Data Engineer III - Databricks, Pyspark, Python, AWS","company":{"id":257,"name":"JPMorganChase","domain":"jpmorganchase.com","url":"https://alion.io/company/jpmorganchase","size_band":"5000+","is_staffing_agency":false,"employer_type":"direct","is_intermediary":false,"listed_via":null,"ats_vendor":"Oracle","truth_index":{"grade":"A","score":100,"open_postings":337,"ghost_share":0.009,"stale_share":0.009,"repost_share":0.024,"time_to_fill_p50_days":4,"computed_at":"2026-09-25T05:45:01Z"}},"role":"Data Science","role_family":"Data Science","seniority":"middle","employment_type":null,"work_mode":"on_site","remote_scope":null,"remote_scope_basis":null,"remote_working_hours":null,"hiring_geo_confidence":"structured","locations":["Bengaluru, India"],"countries":["IN"],"hiring_countries":[],"hiring_countries_total":0,"salary":null,"salary_estimate":{"min_usd":27000,"max_usd":69000,"period":"year","method":"global_role_cell_scaled_by_country","sample_n":431},"experience_years_min":3,"visa_sponsorship":false,"relocation_package":false,"has_equity":false,"technologies":[{"name":"Agile","optional":false},{"name":"Amazon S3","optional":false},{"name":"AWS","optional":false},{"name":"Bitbucket","optional":false},{"name":"Databricks","optional":false},{"name":"Dimensional Modeling","optional":false},{"name":"GitHub","optional":false},{"name":"pySpark","optional":false},{"name":"Python","optional":false},{"name":"Spark","optional":false},{"name":"SQL","optional":false},{"name":"CloudFormation","optional":true},{"name":"Terraform","optional":true}],"status":"live","first_seen_at":"2026-09-22T13:30:39Z","employer_posted_date":"2026-09-23","last_verified_at":"2026-09-25T22:51:45Z","board_verified":true,"closed_at":null,"days_open":3,"trust":{"level":"ok","repost_count":null,"flags":[],"days_open":3},"description":"Be part of a dynamic team where your distinctive skills will contribute to a winning culture and team.\nAs a Data Engineer III - Databricks, Pyspark, Python, AWS at JPMorgan Chase within the Commercial & Investment Bank, you'll serve as a seasoned member of an agile team to design and deliver trusted data collection, storage, access, and analytics solutions in a secure, stable, and scalable way. You are responsible for developing, testing, and maintaining critical data pipelines and architectures across multiple technical areas within various business functions in support of the firm’s business objectives.\nJob responsibilities\nDesign, develop, and maintain big data pipelines (batch and streaming) using PySpark/Spark and Databricks.\nLead/own data modelling and solution design for data products, including defining target-state architecture, data flows, and transformation patterns.\nBuild scalable ingestion and transformation workflows for high-volume datasets, ensuring reliability, quality, and performance.\nDevelop and optimize complex SQL transformations, reconciliation queries, and analytical datasets; perform query tuning for large-scale workloads.\nApply strong data warehousing concepts (dimensional modeling, SCDs, partitioning strategies, etc.) to build well-structured, analytics-ready data layers and Leverage common AWS services, with strong emphasis on S3 and AWS data processing capabilities, to support scalable storage and processing.\nPerform advanced debugging and troubleshooting across distributed Spark workloads (data skew, shuffle tuning, memory/compute optimization).\nImplement engineering best practices: modular design, efficient coding, code reviews, and CI-friendly development approaches.\nUse GitHub/Bitbucket and standard version control workflows to manage codebase, peer reviews, and releases.\nPartner with cross-functional stakeholders to convert requirements into robust big data solutions.\nUses enterprise-authorized AI capabilities within the work environment to accelerate data pipeline/design analysis and documentation, validating outputs and handling data according to sensitivity and security requirements.\nApplies reuse-first, AI-assisted practices to strengthen SDLC-quality routines for data pipelines (e.g., test generation and control validation), ensuring traceability/auditability and alignment to resiliency and security expectations.\n\nRequired qualifications, capabilities, and skills\nFormal training or certification on data engineering concepts and 3+ years applied experience\n Experience in data engineering / big data engineering, with strong hands-on delivery and Expert-level SQL skills (must be extremely strong): complex joins, window functions, CTEs, optimization, and analytical problem solving at scale.\nStrong hands-on coding experience with Python in production environments; demonstrated ability to write efficient, maintainable code.\nDeep expertise in Apache Spark (in depth) and PySpark, including performance tuning and distributed processing fundamentals.\nStrong experience with Databricks for large-scale data processing and pipeline development.\nStrong understanding of data warehousing concepts and best practices; proven capability in data modelling and solution design for scalable, maintainable data platforms/products.\nExperience implementing both batch and streaming data processing solutions.\nFamiliarity with AWS S3 and common AWS services used in data platforms and processing\nExcellent debugging, troubleshooting, problem-solving skills and experience with GitHub, Bitbucket, and version control best practices.\nDemonstrated experience using enterprise-authorized AI capabilities within the work environment to support data engineering workflows with strong validation habits and awareness of data sensitivity.\nAbility to review and validate AI-assisted outputs (e.g., query suggestions, test ideas, or model change summaries) before use, escalating when uncertain and following data handling requirements.\n\nPreferred qualifications, capabilities, and skills\nGood to have: infrastructure provisioning in AWS using Infrastructure as Code (IaC) (e.g., Terraform, AWS CloudFormation).","description_format":"text","description_chars":4171,"description_truncated":false,"requirements":{"experience_years_min":3,"management_years_min":null,"team_size_min":null,"manages_managers":false,"education":null,"security_clearance":false,"languages":[]},"benefits":[],"hiring_locations":[],"hiring_excludes":[],"relocation_offered":false,"industries":["Commercial & Retail Banks","Wealth Management & Financial Advisors","Asset Management & Funds","Investment Banking & M&A Advisory"],"lifecycle":[{"event":"open","at":"2026-09-22T15:06:28Z"},{"event":"close","at":"2026-09-23T04:41:23Z"},{"event":"reopen","at":"2026-09-23T14:28:04Z"}],"liveness":{"score":69,"band":"ok","label":"Likely open","p_open":1,"p_active":0.772,"p_room":0.9,"age_days":2,"expected_fill_days":4,"reasons":["conf:1","velocity","win:mid","comp:brand"],"computed_at":"2026-09-25T05:45:01Z"},"pay":null,"html_url":"https://alion.io/job/jpmorganchase-data-engineer-iii-databricks-pyspark-python-aws","json_url":"https://alion.io/job/jpmorganchase-data-engineer-iii-databricks-pyspark-python-aws.json","meta":{"generated_at":"2026-09-26T04:06:10Z","cache_seconds":300,"methodology":"https://alion.io/methodology","terms":"https://alion.io/terms","contact":"https://alion.io/contact","api":"https://alion.io/developers","usage":{"tier":"crawler","counted_by":"address","units_charged":1,"used_today":4453,"day_limit":5000,"remaining_today":547,"minute_limit":60,"resets_at":"2026-09-27T00:00:00Z"}}}