Echos. AI is looking for an experienced Lead Data Engineer to lead the design, development, and optimization of enterprise-scale cloud data platforms on AWS. The ideal candidate will possess deep expertise in building scalable data pipelines, architecting modern data lakes, implementing distributed data processing frameworks, and leading engineering teams in delivering high-quality data solutions. You will collaborate closely with global stakeholders, solution architects, data scientists, and business teams to develop highly scalable, secure, and reliable data platforms capable of supporting analytics, reporting, AI, and machine learning initiatives. This role demands excellent technical leadership, strong problem-solving skills, and hands-on expertise in AWS cloud services and modern data engineering technologies.
Responsibilities:
- Lead the design, development, and deployment of enterprise-scale data engineering solutions on AWS.
- Architect and implement robust ETL/ELT pipelines using PySpark and AWS Glue.
- Develop highly scalable data ingestion, transformation, and processing frameworks.
- Design and maintain cloud-native data lakes and data warehouse architectures.
- Build and manage workflow orchestration using Apache Airflow.
- Collaborate with data scientists, analysts, product managers, and business stakeholders to understand business requirements.
- Optimize data processing performance, scalability, and cost efficiency.
- Design batch and near real-time data pipelines.
- Ensure high data quality, governance, lineage, security, and compliance across enterprise data platforms.
- Implement monitoring, logging, alerting, and troubleshooting mechanisms for production data pipelines.
- Lead code reviews and establish engineering best practices.
- Mentor junior and mid-level data engineers while promoting technical excellence.
- Work closely with DevOps teams to automate deployments using CI/CD pipelines.
- Participate in architectural discussions and technology selection.
- Collaborate with cross-functional teams across India and the United States.
- Drive continuous improvement initiatives and modernize legacy data platforms.
Requirements:
- Mandatory Technical Skills: Cloud Platform, Amazon Web Services (AWS) (Mandatory).
- Big Data Technologies: PySpark, Apache Spark, AWS Glue, Apache Airflow.
- AWS Services: Amazon S3 AWS Glue, EMR, Redshift, Athena, Lambda, IAM, CloudWatch.
- Programming: Python, SQL.
- Data Engineering: ETL/ELT, Data Lakes, Data Warehousing, Data Modeling, Workflow Automation.
- DevOps & Infrastructure: Docker, Kubernetes, Terraform, Git, Jenkins, CI/CD Pipelines.
- Nice to Have: Kafka, Databricks, Snowflake, Delta Lake, Iceberg, AWS Step Functions, REST APIs, MLOps, Machine Learning pipeline integration.
Preferred Candidate Profile:
- The ideal candidate should possess:
- Strong leadership and mentoring skills.
- Excellent communication and stakeholder management.
- Ability to solve complex technical challenges.
- Strong analytical and problem-solving capabilities.
- Customer-centric mindset.
- Ownership and accountability.
- Passion for continuous learning and innovation.
- Ability to thrive in a fast-paced, collaborative environment.
- Mandatory Skills: AWS, PySpark, Apache Airflow, AWS Glue, Python, SQL. Good to have: EMR, Redshift, Athena, Lambda, S3 Kafka, Docker, and Kubernetes.

