Overview
Technical skills
Timeline
Roles

Overview

Data Engineer with 5+ years building cloud data pipelines, ETL/ELT workflows, and data warehousing solutions across Azure, AWS, and GCP. Experienced in SQL, Python, PySpark/Spark, Databricks, Airflow, and analytics enablement with governed, validated datasets and monitoring practices. Recent work includes ML infrastructure modernization using TensorFlow and PyTorch and integration of data quality into data workflows.

Technical skills

Bash
PowerShell
SQL• Senior • 6y+
Python• Middle
Python
pySpark
Databases
Azure SQL Database
MySQL
PostgreSQL
Snowflake
MS SQL
Google BigQuery• 6y+
Amazon Redshift
Apache Kafka
Databricks
Delta Lake
AI/ML
Airflow• 6y+
Scikit-learn
PyTorch
Spark
TensorFlow
DevOps
Amazon EC2
Amazon EKS
Ansible
Azure
Bicep
CloudFormation
Docker
GCP
GitHub Actions
Google GKE
Jenkins
Kubernetes
New Relic
SLI/SLO/SLA
Git• 6y+
AWS
AWS Lambda
CI/CD
Terraform
Azure DevOps
Cybersecurity
CIS Benchmarks
Microsoft Entra ID
NIST 800-53
PCI DSS
Analytics
Power BI
Tableau
Cryptography
Vault

Timeline

Senior Data Engineer Senior
Morgan Stanley Full-Time
Apr 2025 to Present 1 Year 4 Months New York In office
Built and modernized ML and data infrastructure using TensorFlow and PyTorch to improve model training efficiency and reduce production resource usage. Integrated data quality metrics into workflows to improve prediction reliability and system dependability. Designed data warehousing and Azure pipeline solutions using Azure Data Factory with Databricks and SQL, and implemented PySpark transformations with Delta Lake for curated datasets. Coordinated Azure Synapse and Power BI models with validation and access controls, and applied Git-based CI/CD practices to standardize releases.
TensorFlow
PyTorch
Databricks
Spark
Delta Lake
pySpark
SQL
Azure DevOps
Git
CI/CD
Data Engineer Middle
McKesson Full-Time
Jan 2024 to Mar 2025 1 Year 2 Months Irving In office
Developed a feature engineering framework and improved data pipeline performance with better testing and validation, reducing processing time across distributed systems. Standardized SQL transformation documentation and validation checks to increase data processing speed and availability for analytics teams. Implemented ML workflows with scikit-learn to reduce inference and training time, and supported AI/ML use cases with faster deployment cycles. Built AWS Glue ETL jobs with S3 and Redshift, integrated Kafka-based streaming ingestion patterns, and automated Terraform plus CI/CD deployments for AWS data services.
Scikit-learn
Python
SQL
AWS
Amazon Redshift
AWS Lambda
Terraform
CI/CDsince 2024
Apache Kafka
Associate Data Engineer Middle
Accenture Full-Time
Jun 2020 to Jul 2023 3 Years 1 Month Hyderabad In office
Resolved operational issues by applying collaborative development practices and improving reliability and incident response for critical services. Optimized database performance by consolidating schemas and views, reducing query execution time. Built and maintained GCP analytics assets including BigQuery data marts and warehouse layers, and validated lineage, governance, and data quality rules to reduce reporting discrepancies. Orchestrated batch and real-time workflows using Airflow with Git and Agile delivery practices, improving scheduled workflow reliability and documentation quality.
Google BigQuery
Airflow
SQLsince 2020
Gitsince 2020