Data Scientist
6+ years exp
5+ years ML exp
Python
Scala
Java
SQL
Active 14 days ago
Invite to interview
Message
Download CVCV
Overview
Technical skills
Timeline
Roles
Overview
Data Engineer with 6+ years building cloud-based data platforms, ETL/ELT pipelines, and enterprise analytics across Azure, AWS, and GCP. Experience designing lakehouse architectures with Delta Lake and enabling streaming/near real-time ingestion with Kafka and Spark. Implemented data governance for healthcare compliance and delivered analytics dashboards using Power BI while optimizing performance and automation in production pipelines.
Technical skills
Python
Scala
Java
SQL• Senior • 6y+
Python
FastAPI• 5y+
Flask• 5y+
pySpark• 5y+
Java
Jackson
Databases
Azure Cosmos DB
MS SQL
ElasticSearch
Google BigQuery• 6y+
Snowflake• 6y+
Teradata• 6y+
Apache Kafka• 5y+
Azure SQL Database• 5y+
Delta Lake• 5y+
MySQL• 5y+
PostgreSQL• 5y+
Amazon Redshift• 3y+
DynamoDB• 3y+
Databricks
AI/ML
NLP
Airflow• 5y+
Great Expectations• 5y+
Hadoop• 5y+
NLTK• 5y+
Pandas• 5y+
Spark• 5y+
DevOps
AWS
Azure AKS
CI/CD
GitHub Actions
Jenkins
SLI/SLO/SLA
Rest API
Git• 6y+
Azure DevOps• 5y+
Bitbucket• 5y+
Docker• 5y+
GCP• 5y+
Kibana• 5y+
Kubernetes• 5y+
AWS Lambda• 3y+
Azure
Cybersecurity
GDPR• 5y+
HIPAA• 5y+
Microsoft Entra ID• 5y+
Analytics
Power BI• 6y+
Matplotlib• 5y+
QA
Pytest• 5y+
Timeline
Senior Data Engineer
•
Senior
Jackson Health System
•
Full-Time
Designed and implemented enterprise healthcare data pipelines integrating clinical, claims, provider, and member data into centralized analytics on Azure. Built a governed lakehouse with Delta Lake and Databricks, supporting Bronze/Silver/Gold layers and self-service analytics. Developed PySpark and Spark SQL datasets for healthcare reporting and enabled near real-time ingestion using Kafka-based streaming. Implemented HIPAA-oriented governance using Entra ID and access controls, and delivered KPI dashboards in Power BI.
pySpark
Spark
Apache Kafka
Delta Lake
SQL
Power BI
Microsoft Entra ID
HIPAA
GDPR
Data Engineer
•
Middle
IBM
•
Full-Time
Built scalable, fault-tolerant AWS data pipelines using DynamoDB, Lambda, S3-based ingestion, and workflow orchestration. Optimized Spark SQL processing in AWS Glue/EMR to improve query execution efficiency and supported downstream real-time reporting use cases. Led migrations from on-prem databases to AWS services including Redshift while ensuring secure data transfer. Implemented CI/CD and deployment automation using Azure DevOps for data pipeline delivery.
DynamoDB
AWS Lambda
Amazon Redshift
Spark
SQL
Azure DevOps
Data Engineer
•
Middle
Blue Cross and Blue Shield
•
Full-Time
Delivered healthcare analytics workflows by translating business and regulatory requirements into scalable data pipelines. Designed and optimized data models using Azure SQL Database, PostgreSQL, and MySQL, and built dashboards in Power BI, Kibana, and Matplotlib. Architected cloud warehouse solutions with Azure Synapse Analytics and Snowflake, and automated ETL using Azure Data Factory and Apache Airflow. Implemented streaming and governance patterns with Kafka, Delta Lake, Entra ID RBAC, and compliance controls (HIPAA/GDPR), and created NLP extraction pipelines using NLTK with testing and validation via Great Expectations and pytest.
Azure SQL Database
PostgreSQL
MySQL
SQL
Sparksince 2021
pySparksince 2021
Apache Kafkasince 2021
Airflow
Snowflake
Delta Lakesince 2021
Kubernetes
Docker
Git
Bitbucket
Azure DevOpssince 2021
Jira
Power BI
Kibana
Matplotlib
Great Expectations
Pytest
Pandas
Microsoft Entra IDsince 2021
HIPAAsince 2021
GDPRsince 2021
NLTK
Flask
FastAPI
Hadoop
GCP Data Engineer / Big Data Engineer
•
Middle
Target Corporation
•
Full-Time
Supported production deployments by coordinating with DevOps teams for release success, monitoring, and security controls. Created Oozie workflows for Hadoop-based processing and developed Hive/HQL table definitions for loading and querying datasets. Integrated and governed data across Snowflake and BigQuery, and performed data validation and troubleshooting using SQL on Teradata. Used Git and Jira for version control and task tracking, and supported reporting with Power BI during data warehouse migrations.
Google BigQuery
Snowflakesince 2020
Teradata
SQLsince 2020
Gitsince 2020
Jirasince 2020
Power BIsince 2020
