Data Scientist
6+ years exp
5+ years ML exp
SQL
Python
Java
Scala
Invite to interview
Download CVCV
Overview
Technical skills
Timeline
Roles
Overview
Data Engineer with 6+ years building cloud-based data platforms, ETL/ELT pipelines, and enterprise analytics across Azure, AWS, and GCP. Experience designing lakehouse architectures with Delta Lake and enabling streaming/near real-time ingestion with Kafka and Spark. Implemented data governance for healthcare compliance and delivered analytics dashboards using Power BI while optimizing performance and automation in production pipelines.
Technical skills
Languages
4
SQL
Python
Java
Scala
Python
3
FastAPI
Flask
pySpark
Databases
11
Snowflake
PostgreSQL
Apache Kafka
Delta Lake
MySQL
Azure SQL Database
Google BigQuery
Teradata
Amazon Redshift
DynamoDB
ElasticSearch
DevOps
10
GCP
Kubernetes
Git
Docker
Azure DevOps
AWS Lambda
Kibana
Bitbucket
AWS
Azure
AI/ML
6
Pandas
Spark
Airflow
NLTK
Hadoop
Great Expectations
Cybersecurity
3
Microsoft Entra ID
GDPR
HIPAA
Other
14
Power BI
Pytest
Matplotlib
Jackson
Azure Cosmos DB
MS SQL
Databricks
Azure AKS
GitHub Actions
Rest API
CI/CD
Jenkins
NLP
SLI/SLO/SLA
Timeline
Senior Data Engineer
•
Senior
Jackson Health System
•
Full-Time
Designed and implemented enterprise healthcare data pipelines integrating clinical, claims, provider, and member data into centralized analytics on Azure. Built a governed lakehouse with Delta Lake and Databricks, supporting Bronze/Silver/Gold layers and self-service analytics. Developed PySpark and Spark SQL datasets for healthcare reporting and enabled near real-time ingestion using Kafka-based streaming. Implemented HIPAA-oriented governance using Entra ID and access controls, and delivered KPI dashboards in Power BI.
pySpark
Spark
Apache Kafka
Delta Lake
SQL
Power BI
Microsoft Entra ID
HIPAA
GDPR
Data Engineer
•
Middle
IBM
•
Full-Time
Built scalable, fault-tolerant AWS data pipelines using DynamoDB, Lambda, S3-based ingestion, and workflow orchestration. Optimized Spark SQL processing in AWS Glue/EMR to improve query execution efficiency and supported downstream real-time reporting use cases. Led migrations from on-prem databases to AWS services including Redshift while ensuring secure data transfer. Implemented CI/CD and deployment automation using Azure DevOps for data pipeline delivery.
DynamoDB
AWS Lambda
Amazon Redshift
Spark
SQL
Azure DevOps
Data Engineer
•
Middle
Blue Cross and Blue Shield
•
Full-Time
Delivered healthcare analytics workflows by translating business and regulatory requirements into scalable data pipelines. Designed and optimized data models using Azure SQL Database, PostgreSQL, and MySQL, and built dashboards in Power BI, Kibana, and Matplotlib. Architected cloud warehouse solutions with Azure Synapse Analytics and Snowflake, and automated ETL using Azure Data Factory and Apache Airflow. Implemented streaming and governance patterns with Kafka, Delta Lake, Entra ID RBAC, and compliance controls (HIPAA/GDPR), and created NLP extraction pipelines using NLTK with testing and validation via Great Expectations and pytest.
Azure SQL Database
PostgreSQL
MySQL
SQL
Spark
pySpark
Apache Kafka
Airflow
Snowflake
Delta Lake
Kubernetes
Docker
Git
Bitbucket
Azure DevOps
Jira
Power BI
Kibana
Matplotlib
Great Expectations
Pytest
Pandas
Microsoft Entra ID
HIPAA
GDPR
NLTK
Flask
FastAPI
Hadoop
GCP Data Engineer / Big Data Engineer
•
Middle
Target Corporation
•
Full-Time
Supported production deployments by coordinating with DevOps teams for release success, monitoring, and security controls. Created Oozie workflows for Hadoop-based processing and developed Hive/HQL table definitions for loading and querying datasets. Integrated and governed data across Snowflake and BigQuery, and performed data validation and troubleshooting using SQL on Teradata. Used Git and Jira for version control and task tracking, and supported reporting with Power BI during data warehouse migrations.
Google BigQuery
Snowflake
Teradata
SQL
Git
Jira
Power BI
