Data Scientist
6+ years exp
6+ years ML exp
Python
Scala
SQL
TypeScript
Java
JavaScript
Active 14 days ago
+1 (989) 3307909 Invite to interview
Message
Download CVCV
Overview
Technical skills
Timeline
Roles
Overview
Data Engineer with 5+ years of experience building ETL pipelines and data infrastructure for AI/ML-ready, user-facing data products. Experienced in AWS-based ingestion and orchestration using Spark, PySpark, Airflow, and dbt with Redshift modeling, validation, and monitoring. Delivered improvements including a reported 40% cost reduction and 40%+ performance gains, with experience spanning financial-services and healthcare domains.
Technical skills
Python
Scala
SQL
TypeScript
Java
JavaScript
Python
pySpark
Databases
Delta Lake
PostgreSQL
Apache Kafka• 6y+
Google BigQuery• 6y+
Amazon Redshift• 3y+
AI/ML
Dagster
Prefect
Spark• 6y+
Airflow• 3y+
dbt• 3y+
DevOps
Azure
GCP
GitHub Actions
Terraform
CI/CD• 6y+
Git• 6y+
Jenkins• 6y+
Amazon EC2• 3y+
AWS• 3y+
Grafana• 3y+
AWS Lambda
Timeline
Data Engineer
•
Middle
UnitedHealth Group
•
Full-Time
Designed scalable AWS ETL pipelines using Spark and PySpark to process large datasets in S3 for healthcare analytics. Built metadata-driven orchestration using AWS Step Functions with Lambda and Glue jobs, including logging and error handling. Implemented data validation and schema enforcement in PySpark with monitoring and alerts for ingestion anomalies. Modeled star/snowflake schemas in Amazon Redshift for analytics and improved batch performance with Spark optimization techniques.
Spark
pySpark
AWS
AWS Lambda
Amazon Redshift
Data Engineer
•
Middle
Capital One
•
Full-Time
Built real-time ingestion pipelines on AWS for fraud-signal detection using streaming with Spark Streaming and AWS Kinesis. Developed and managed dbt data models with Redshift and automated CI/CD testing for transformation and auditing needs. Implemented lakehouse-style S3 architecture concepts and ran orchestration workflows in Airflow on EC2 to meet compliance SLAs. Reduced processing costs through Spark and Redshift tuning and improved operational visibility with Grafana dashboards and custom metrics.
Spark
dbt
Amazon Redshiftsince 2023
Airflow
Amazon EC2
AWSsince 2023
Grafana
CI/CD
Data Engineer
•
Middle
Target
•
Full-Time
Developed ingestion workflows from relational sources and Kafka using Spark SQL and DataFrame transformations with cleansing, validation, and deduplication. Maintained Git-based version control and built Jenkins CI pipelines to support testing and safe deployments across environments. Created BigQuery views for executive reporting and designed Hive external tables with partitioning, compression, and bucketing for efficient historical analytics. Implemented data quality checks and followed Agile delivery practices to meet backlog and roadmap needs.
Sparksince 2020
Apache Kafka
Google BigQuery
Git
Jenkins
CI/CDsince 2020
