368,634open jobs
9,437companies
50,578added this week
Browse all
Salary
$20k – $38k per year (Estimated)
Location
In office (Bengaluru)
Seniority
Senior · 7+ years exp
Employment
Full-Time
Overview
Company
Impact
Profile match
Synechron is a global digital transformation, technology, and management consulting firm headquartered in New York City. Founded in 2001 by Faisal Husain, Tanveer Saulat, and Zia Bhutta, the company specializes exclusively in end-to-end IT solutions, systems integration, and business consulting for the Financial Services, Banking, Asset Management, and Insurance (BFSI) sectors.

Job Summary

Synechron is seeking a PySpark Data Engineer with 7+ years of overall experience and at least 5+ years of commercial experience in data-driven roles. The role will design, develop, test, deploy, and support scalable data pipelines, data marts, and data warehousing solutions using Python, PySpark, SQL, and related data technologies.The position will contribute to business objectives by delivering reliable data solutions, improving data quality and accessibility, supporting analytics and reporting, and ensuring effective data processing across the full software development lifecycle.

Software Requirements

Required

  • 7+ years of overall professional experience in data engineering, software development, or related technology roles.

  • 5+ years of commercial experience in a data-driven role.

  • Hands-on experience building data marts and ETL pipelines.

  • Strong expertise in Python and PySpark for ETL scripting.

  • Experience writing clean, maintainable, robust, and testable Python code.

  • Hands-on experience with Spark, PySpark, Hadoop, MapReduce, Hive, and Pandas.

  • Strong knowledge of SQL and Oracle query development.

  • Experience working with SQL and NoSQL database management systems.

  • Experience across the end-to-end software development lifecycle, including:

    • Build and development.

    • User acceptance testing.

    • UAT defect resolution.

    • Production deployment.

    • Post-production support.

  • Experience debugging PySpark code and investigating data processing issues.

  • Strong understanding of data warehousing and data pipeline production practices.

  • Ability to process structured, semi-structured, and unstructured data.

  • Familiarity with Git, CI/CD processes, data testing, and validation.

  • Experience with data analysis, data cleansing, data linking, imputation, and feature engineering.

  • Familiarity with workflow orchestration and scheduling tools.

  • Experience collaborating with multiple technical and business teams.

Preferred

  • Experience with Apache Airflow, Oozie, and Jenkins pipelines.

  • Experience using Jupyter for data exploration, prototyping, and analysis.

  • Knowledge of cloud-based data engineering platforms and services.

  • Experience with data lake, lakehouse, distributed processing, and streaming concepts.

  • Familiarity with automated data quality monitoring and pipeline observability.

  • Experience in banking, financial services, or other regulated, data-intensive industries.

  • Knowledge of data governance, metadata management, lineage, security, and privacy practices.

  • Experience leading technical workstreams or coordinating delivery across multiple teams.

Overall Responsibilities

  • Design, develop, test, deploy, and support scalable ETL pipelines and data marts using Python and PySpark.

  • Build data processing solutions for structured, semi-structured, and unstructured data.

  • Develop clean, maintainable, robust, and reusable Python and PySpark code.

  • Analyze business and technical requirements and translate them into data engineering solutions.

  • Develop and optimize SQL and Oracle queries for data extraction, transformation, validation, and analysis.

  • Integrate data from multiple sources, databases, files, and systems.

  • Apply data cleansing, data linking, imputation, transformation, validation, and feature engineering techniques.

  • Support data warehouse development, data modeling, data integration, and reporting requirements.

  • Participate in build, UAT, UAT defect resolution, production deployment, and post-production support activities.

  • Debug PySpark code, investigate pipeline failures, and resolve data quality and processing issues.

  • Validate data outputs, reconcile results, and ensure that pipelines meet defined quality and business requirements.

  • Collaborate with technical and non-technical stakeholders to clarify requirements, resolve dependencies, and deliver agreed outcomes.

  • Participate in code reviews, technical discussions, testing, deployment planning, and production support activities.

  • Identify opportunities to improve pipeline performance, automation, reliability, maintainability, and resource efficiency.

  • Maintain technical documentation covering data flows, pipeline logic, data models, dependencies, test evidence, and operational procedures.

  • Consider security, data privacy, cost management, and sustainability when designing and operating data solutions.

Technical Skills (By Category)

Programming Languages

Essential

  • Python using a current and supported version.

  • PySpark for distributed data processing and ETL development.

  • Strong understanding of Python functions, modules, object-oriented programming, exception handling, testing, and package management.

  • Ability to write clean, maintainable, robust, reusable, and testable code.

  • SQL for data extraction, transformation, validation, analysis, and query optimization.

  • Understanding of data structures, algorithms, and software engineering principles.

Preferred

  • Shell scripting for automation and operational support.

  • Experience developing reusable Python packages and data-processing utilities.

  • Knowledge of programming practices for distributed and production-scale data applications.

Databases/Data Management

Essential

  • Strong knowledge of relational databases and Oracle query development.

  • Experience with SQL and NoSQL database management systems.

  • Understanding of data warehousing, data marts, data modeling, and data integration.

  • Knowledge of structured, semi-structured, and unstructured data processing.

  • Experience with data cleansing, data linking, imputation, reconciliation, transformation, and validation.

  • Understanding of data quality, data integrity, data lifecycle, and metadata requirements.

  • Ability to analyze large datasets and identify data inconsistencies or processing issues.

Preferred

  • Experience with dimensional modeling, fact and dimension tables, and analytical data warehouse design.

  • Knowledge of data lake and lakehouse architectures.

  • Familiarity with data lineage, metadata management, and data governance.

  • Experience with feature engineering and preparing data for analytics or machine learning use cases.

  • Knowledge of database performance tuning and query optimization.

Cloud Technologies

Essential

  • Understanding of cloud-based data engineering concepts and distributed data processing.

  • Awareness of cloud storage, compute, networking, access management, monitoring, and deployment considerations.

  • Ability to support data pipelines across development, test, UAT, and production environments.

Preferred

  • Experience developing and deploying PySpark data pipelines on cloud platforms.

  • Familiarity with cloud-based data lakes, data warehouses, managed databases, and workflow services.

  • Knowledge of cloud monitoring, infrastructure automation, identity management, and security controls.

  • Understanding of cost-efficient and sustainable use of cloud data-processing resources.

Frameworks and Libraries

Essential

  • Apache Spark and PySpark.

  • Hadoop, MapReduce, and Hive.

  • Pandas for data analysis and transformation.

  • Python libraries for database connectivity, file handling, data validation, and automation.

  • Experience developing ETL and data-processing frameworks.

  • Understanding of distributed processing, partitioning, transformations, actions, and performance considerations.

Preferred

  • Apache Airflow or Oozie for workflow orchestration.

  • Jupyter for data analysis, exploration, and prototyping.

  • Libraries supporting data quality, testing, feature engineering, and statistical analysis.

  • Familiarity with streaming or near-real-time data-processing frameworks.

Development Tools and Methodologies

Essential

  • Experience across the end-to-end SDLC, including build, UAT, defect fixing, deployment, and post-production support.

  • Git for source code versioning, branching, merging, and code review.

  • Familiarity with CI/CD processes and automated build or deployment workflows.

  • Experience with data testing, validation, reconciliation, and defect management.

  • Knowledge of Agile or iterative software delivery practices.

  • Ability to document data flows, transformation logic, data dependencies, test results, and operational procedures.

  • Experience coordinating with multiple teams to resolve dependencies and deliver project outcomes.

Preferred

  • Jenkins pipeline experience.

  • Experience with automated data quality checks and test execution.

  • Familiarity with pipeline monitoring, logging, alerting, and incident management.

  • Knowledge of infrastructure as code and automated environment deployment.

  • Experience with performance monitoring and optimization of production data pipelines.

Security Protocols

Essential

  • Understanding of secure data handling and data protection principles.

  • Awareness of authentication, authorization, identity and access management, encryption, secrets management, and secure connectivity.

  • Ability to apply appropriate access controls to data pipelines, databases, files, and processing environments.

  • Understanding of data privacy, data integrity, auditability, and secure transfer practices.

Preferred

  • Experience implementing security controls across cloud and on-premises data environments.

  • Knowledge of data masking, tokenization, role-based access control, and audit logging.

  • Familiarity with vulnerability management, security testing, and compliance-related data controls.

  • Understanding of secure configuration and monitoring practices for data platforms.

Experience Requirements

  • 7+ years of overall experience in data engineering, software development, or related technology roles.

  • 5+ years of commercial experience in a data-driven role.

  • Experience building data marts and ETL pipelines.

  • Strong hands-on experience with Python and PySpark for ETL scripting.

  • Experience with Spark, Hadoop, MapReduce, Hive, Pandas, SQL, and Oracle queries.

  • Experience working with SQL and NoSQL database technologies.

  • Experience across build, UAT, UAT defect resolution, production deployment, and post-production support.

  • Experience debugging PySpark code and resolving data pipeline, data quality, and production issues.

  • Strong understanding of data warehousing and production data pipeline practices.

  • Experience handling structured, semi-structured, and unstructured data.

  • Experience with CI/CD, Git, data testing, validation, workflow scheduling, and pipeline support.

  • Experience in banking, financial services, or other regulated data-intensive industries is preferred.

  • Candidates may also qualify through equivalent practical experience, relevant certifications, professional training, or demonstrated delivery of complex data engineering solutions.

Day-to-Day Activities

  • Design, develop, test, and maintain Python and PySpark ETL pipelines, data marts, data transformations, and data warehouse components.

  • Collaborate with technical and non-technical stakeholders, participate in Agile meetings, clarify requirements, and resolve cross-team dependencies.

  • Perform data analysis, Oracle query development, PySpark debugging, data validation, UAT defect fixing, and production support.

  • Review pipeline results, monitor delivery progress, document technical outcomes, recommend improvements, and make implementation decisions within approved standards.

Qualifications

  • Degree in Computer Science, Information Technology Engineering, or an equivalent discipline; equivalent professional experience may be considered.

  • Minimum of 7+ years of overall experience, including at least 5+ years of commercial experience in data-driven roles.

  • Certifications in data engineering, cloud technologies, Python, Spark, Agile, or database technologies are preferred.

  • Complete Synechron-required training related to information security, data protection, data governance, workplace conduct, and responsible technology use.

  • Maintain continuous professional development in Python, PySpark, Spark, data warehousing, cloud data platforms, SQL, automation, security, and data engineering practices.

Professional Competencies

  • Critical thinking, data analysis, technical investigation, and structured problem-solving.

  • Technical ownership, teamwork, dependency coordination, and delivery accountability.

  • Clear communication with technical and non-technical stakeholders.

  • Adaptability, continuous learning, and effective response to changing data and delivery requirements.

  • Innovation focused on reliable, maintainable, automated, scalable, and sustainable data solutions.

  • Effective prioritization, organization, time management, and delivery under multiple deadlines.

S YNECHRON’S DIVERSITY & INCLUSION STATEMENT

Diversity & Inclusion are fundamental to our culture, and Synechron is proud to be an equal opportunity workplace and is an affirmative action employer. Our Diversity, Equity, and Inclusion (DEI) initiative ‘Same Difference’ is committed to fostering an inclusive culture - promoting equality, diversity and an environment that is respectful to all. We strongly believe that a diverse workforce helps build stronger, successful businesses as a global company. We encourage applicants from across diverse backgrounds, race, ethnicities, religion, age, marital status, gender, sexual orientations, or disabilities to apply. We empower our global workforce by offering flexible workplace arrangements, mentoring, internal mobility, learning and development programs, and more.

All employment decisions at Synechron are based on business needs, job requirements and individual qualifications, without regard to the applicant’s gender, gender identity, sexual orientation, race, ethnicity, disabled or veteran status, or any other characteristic protected by law.

Candidate Application Notice

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
368,634 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
Bengaluru
$21k – $57k per year (Estimated) • Remote/Hybrid • Full-Time • 6+ years exp • Bachelor's Degree • Noida
C#
SQL
TypeScript
JavaScript
C#
.NET
Frontend
Angular
DevOps
Azure
Azure DevOps
CI/CD
Git
TeamCity
Apply
$16k – $42k per year (Estimated) • Remote/Hybrid • Full-Time • 3+ years exp • Bachelor's Degree • Noida
C#
SQL
TypeScript
JavaScript
C#
.NET
Frontend
Angular
DevOps
Azure
Azure DevOps
CI/CD
Git
TeamCity
Apply
$21k – $57k per year (Estimated) • Remote/Hybrid • Full-Time • 8+ years exp • Bachelor's Degree • Noida
C#
SQL
TypeScript
JavaScript
C#
.NET
Frontend
Angular
DevOps
Azure
Azure DevOps
CI/CD
Git
TeamCity
Apply
$16k – $42k per year (Estimated) • Remote/Hybrid • Full-Time • 3+ years exp • Bachelor's Degree • Noida
C#
SQL
TypeScript
JavaScript
C#
.NET
Frontend
Angular
DevOps
Azure
Azure DevOps
CI/CD
Git
TeamCity
Apply
$18k – $48k per year (Estimated) • Remote/Hybrid • Full-Time • 3+ years exp • Bachelor's Degree • Noida
C#
SQL
TypeScript
JavaScript
C#
.NET
Frontend
Angular
DevOps
Azure
Azure DevOps
CI/CD
Git
TeamCity
Apply
UI Engineer 4 days ago
In office • Full-Time • Sydney
TypeScript
JavaScript
Frontend
Material UI
Next.js
React Hook Form
React Query
React.js
Redux
Redux Toolkit
shadcn/ui
Tailwind CSS
Zod
Radix UI
Apply
Platform Engineer 4 days ago
$44k – $111k per year (Estimated) • In office • Full-Time • 5+ years exp • Melbourne
PowerShell
Python
DevOps
Amazon EC2
AWS
AWS Fargate
Azure
Azure DevOps
Canary Release
CI/CD
CloudFormation
Datadog
Docker
FinOps
GitHub Actions
GitLab CI
Grafana
Jenkins
Kubernetes
Prometheus
Splunk
Terraform
Amazon CloudWatch
Amazon ECS
GitHub
GitLab
IAM
Apply
$83k – $90k per year • In office • Full-Time • 7+ years exp • Mississauga
Go
Java
Node JS
Python
JavaScript
Databases
Apache Kafka
DevOps
GCP
Apply
Business Analyst 4 days ago
$100k – $110k per year • In office • Full-Time • Montreal
Management
Confluence
Jira
Apply
Solution Architect 4 days ago
$58k – $139k per year (Estimated) • In office • Full-Time • 10+ years exp • Sydney
C#
C#
.NET
DevOps
AWS
Azure
GCP
Apply
$16k – $34k per year (Estimated) • Remote/Hybrid • Full-Time • 2+ years exp • Bachelor's Degree • Mumbai • Bengaluru
JavaScript
PowerShell
SQL
C#
C#
.NET
Databases
Azure SQL Database
MS SQL
DevOps
Azure
Rest API
Cybersecurity
Microsoft Entra ID
QA
Postman
Swagger
Apply
$41k – $89k per year (Estimated) • Remote/Hybrid • Full-Time • 8+ years exp • Bengaluru
C#
TypeScript
JavaScript
C#
.NET
Databases
Apache Kafka
AI/ML
Copilot
LLM
OpenAI
Frontend
Angular
GraphQL
DevOps
Azure
Azure AKS
Azure DevOps
CI/CD
Docker
GitHub
GitHub Actions
Grafana
Kubernetes
Prometheus
Rest API
Apply
$38k – $83k per year (Estimated) • In office • Full-Time • 12+ years exp • Bachelor's Degree • Bengaluru
Databases
Oracle
DevOps
AWS
Platform Engineering
Apply
Data Architect 3 hours ago
$38k – $91k per year (Estimated) • In office • Full-Time • 3+ years exp • Bengaluru • Pune
Node JS
Python
SQL
JavaScript
Databases
Databricks
MongoDB
Redis
Apply
$28k – $71k per year (Estimated) • In office • Full-Time • 5+ years exp • Bengaluru
DevOps
CI/CD
Platform Engineering
Apply
See all jobs
This is one of many
368,634 more open roles from verified company boards, updated every day.