Overview
Technical skills
Timeline
Roles

Overview

Lead Data Engineer specializing in cloud-native healthcare data platforms and analytics. Experienced in architecting lakehouse ETL/ELT pipelines, implementing data governance and HIPAA-aligned security controls, and improving reliability with automated data quality and monitoring. Works across Azure and other cloud data systems, with additional experience in streaming integrations and AI-ready data preparation.

Technical skills

Scala
Java
Python• Senior • 10y+
SQL• Senior • 10y+
Python
pySpark• 10y+
Databases
Amazon Redshift
Apache Iceberg
Cassandra
DynamoDB
Google BigQuery
Milvus
MySQL
Oracle
PostgreSQL
Weaviate
Databricks• 10y+
Delta Lake• 7y+
Apache Kafka• 6y+
Snowflake• 6y+
Pinecone• 5y+
AI/ML
Dagster
Flink
Hadoop
LlamaIndex
NumPy
Prefect
PyTorch
RAG
Scikit-learn
TensorFlow
Spark• 10y+
Airflow• 6y+
dbt• 6y+
Great Expectations• 6y+
MLFlow• 6y+
Copilot• 5y+
ChatGPT• 4y+
LangChain• 4y+
Claude• 3y+
DevOps
AWS
AWS Lambda
Azure DevOps
Datadog
GCP
GitOps
Grafana
Vector
Azure• 10y+
CI/CD• 10y+
Docker• 10y+
Git• 10y+
Jenkins• 10y+
Terraform• 10y+
GitHub Actions• 6y+
Kubernetes• 6y+
Cybersecurity
GDPR
HIPAA
Analytics
Tableau
Power BI• 10y+
Cryptography
Vault

Timeline

Lead Data Engineer Lead
Evidation Health Full-Time
Apr 2023 to Present 3 Years 5 Months Columbus In office
Architected and delivered cloud-native healthcare data platforms on Azure Databricks and Azure Data Factory, integrating FHIR R4, HL7, claims, and wearable data into a unified lakehouse. Built scalable ETL/ELT pipelines with Python, PySpark, Spark, and SQL, including Bronze/Silver/Gold models for analytics and ML workloads. Implemented HIPAA-focused security and governance controls (RBAC, encryption, audit logging, PHI de-identification) and added automated data quality, validation, lineage, and monitoring with Great Expectations.
Azure
Databricks
Delta Lake
Python
pySpark
SQL
Spark
Apache Kafka
Great Expectations
Senior Data Engineer Senior
Leaflink Full-Time
Feb 2020 to Mar 2023 3 Years 1 Month In office
Designed and implemented cloud-native lakehouse solutions using Snowflake, Databricks, Delta Lake, dbt, and Apache Airflow to support analytics, AI/ML, and self-service reporting. Built end-to-end ETL/ELT with CDC and streaming integrations using Kafka, Spark Structured Streaming, and connector tools, while managing schema evolution and lineage. Created AI-ready feature pipelines and reusable data products for predictive analytics and RAG using MLflow, LangChain, and vector databases, and supported DevOps/DataOps with GitHub Actions, Terraform, Docker, and Kubernetes.
Snowflake
Databricks
Delta Lake
dbt
Airflow
Apache Kafkasince 2020
Spark
MLFlow
LangChain
Pinecone
Azure
GitHub Actions
Terraform
Docker
Kubernetes
CI/CD
Git
Great Expectationssince 2020
Copilot
Claude
ChatGPT
Data Engineer Middle
Benchling Full-Time
Mar 2016 to Jan 2020 3 Years 10 Months In office
Built cloud-native data pipelines on Azure using Azure Data Factory and Azure Databricks with Delta Lake, ingesting lab, clinical, EHR/FHIR, and research datasets into a centralized lakehouse. Developed automated ETL/ELT workflows with validation, lineage, and quality monitoring, and applied Bronze/Silver/Gold modeling for trusted analytics outputs. Improved distributed Spark performance with Delta Lake optimizations and delivered semantic layers and dashboards in Power BI, while enforcing HIPAA/GDPR security controls and automating CI/CD with Git, Jenkins, Docker, and Terraform.
Azuresince 2016
Databrickssince 2016
Delta Lakesince 2016
Pythonsince 2016
pySparksince 2016
SQLsince 2016
Sparksince 2016
Power BI
Gitsince 2016
Jenkins
Dockersince 2016
Terraformsince 2016
CI/CDsince 2016