We are seeking a highly experienced Senior Data Engineer with strong expertise in Google Cloud Platform (GCP), Apache Spark, and Scala to design, develop, and optimize large-scale data processing solutions. The ideal candidate should have hands-on experience building enterprise-grade data pipelines, real-time and batch processing systems, cloud-native data platforms, and modern lakehouse architectures. The candidate will work closely with data architects, data scientists, product owners, and business stakeholders to deliver scalable and reliable data solutions.
The core responsibilities for the job include the following:
Data Engineering and Development:
- Design, build, and maintain scalable ETL/ELT pipelines.
- Develop large-scale data processing applications using Apache Spark and Scala.
- Build batch and real-time data ingestion frameworks.
- Implement data transformation, cleansing, enrichment, and aggregation pipelines.
- Ensure data quality, consistency, and governance across platforms.
GCP Cloud Engineering:
- Build and support cloud-native data solutions on Google Cloud Platform.
- Work with: BigQuery, Cloud Storage, Dataflow, Dataproc, Pub/Sub, Composer (Airflow), Cloud Functions, and Cloud Run.
- Optimize cloud resources for performance and cost efficiency.
Architecture and Design:
- Participate in data lake and lakehouse architecture design.
- Design scalable distributed systems for high-volume datasets.
- Implement data models supporting analytics and reporting req.
Performance Optimization:
- Optimize Spark jobs for large datasets.
- Tune SQL queries and Spark transformations.
- Improve cluster utilization and execution efficiency.
- Implement partitioning and data distribution strategies.
Collaboration:
- Work closely with: Data Architects, Data Scientists, DevOps Teams, Product Owners, and Business Stakeholders.
- Provide technical leadership and mentoring to junior engineers.
Requirements:
- Experience: 9-14 Years.

