Data Scientist
6+ years exp
4+ years ML exp
Apex
Java
Python
SQL
Active 14 days ago
Invite to interview
Message
Download CVCV
Overview
Technical skills
Timeline
Roles
Overview
Data engineer with 5+ years of experience building governed data platforms and analytics systems, including Snowflake-based warehousing and dbt semantic modeling. Worked on cloud-native ETL and orchestration across GCP, Databricks and production reliability improvements using CI/CD and infrastructure-as-code. Delivered AI-enabled analytics using dbt mesh, and built LLM-powered agents with LangChain alongside data science and reporting workflows.
Technical skills
Apex
Java• 6y+
Python• Middle • 5y+
SQL• Junior • 5y+
Apex
MuleSoft• 6y+
Python
pySpark• 5y+
Databases
Amazon Redshift
Snowflake
PostgreSQL
Apache Kafka• 4y+
Google BigQuery
SAP HANA
Databricks
Delta Lake
Neo4j
AI/ML
AI Agents
Hadoop
Spark
dbt• 4y+
LangChain
LLM
XGBoost
Mobile
Braze
DevOps
CI/CD
Cortex
SLI/SLO/SLA
Prometheus
Git• 6y+
Rest API• 6y+
AWS• 5y+
Docker• 5y+
GCP• 4y+
Azure
GitHub Actions
Kubernetes
Terraform
Cybersecurity
HIPAA
Analytics
Tableau• 5y+
Power BI
Timeline
Senior Data Engineer
•
Senior
ID.me
•
Full-Time
Built cloud-native ETL pipelines to synchronize ERP, API and graph data into a governed analytics stack on GCP and Databricks. Orchestrated multiple production pipelines with Airflow on Kubernetes and implemented reliability features like retries, alerting, and backfills. Improved deployment lead time using reusable CI/CD and infrastructure-as-code with automated validation and testing. Delivered standardized KPI semantics via dbt mesh to support Snowflake Cortex, Looker and Sigma dashboards.
Google BigQuery
GCP
Databricks
Delta Lake
dbt
Kubernetes
GitHub Actions
Terraform
SQL
Braze
Salesforce
Neo4j
Senior Data Engineer
•
Senior
Stride Inc
•
Full-Time
Architected a Snowflake data warehouse using dbt and Python, establishing reusable transformations and modeling patterns across many analytics tables. Migrated large datasets from SAP HANA, BigQuery, and Azure storage into Snowflake while tuning processing for full, incremental, and CDC flows. Built LLM-powered agents with LangChain to automate student risk identification and reporting. Enhanced analytics reporting with Power BI by defining KPIs aligned to ROMI, LTV, and CAC and improving monitoring.
dbt
Python
SAP HANA
Google BigQuerysince 2024
Azure
LLM
LangChain
Power BI
Data Engineer Intern
•
Junior
Skyworks Solutions Inc.
•
Internship
Built XGBoost forecasting models using AWS Redshift datasets to support inventory risk analysis with strong predictive accuracy. Developed production-grade Snowflake ETL pipelines ingesting data from Oracle, S3, REST APIs and mixed file formats with data quality tests and audit logging. Ran A/B testing on warehouse region strategies to optimize inventory thresholds and improve turnover. Produced Power BI dashboards and DAX KPI models for large user groups by reducing reliance on manual reporting.
XGBoost
Python
AWS
SQL
Rest API
Senior Data Engineer
•
Senior
Snowflake
•
Full-Time
Delivered PySpark and Python-based ETL migrations for a fintech SaaS platform processing large transaction volumes across multiple business domains. Implemented cloud data pipelines on GCP with batch and near-real-time analytics using event streaming. Maintained and monitored orchestration workflows including SLA and dbt failure alerting, and supported data movement into Snowflake. Collaborated with client engineering teams to standardize Git-based branching strategies for parallel QA testing.
pySpark
Python
GCPsince 2022
Apache Kafka
dbtsince 2022
Data Engineer
•
Middle
Snowflake
•
Full-Time
Optimized GCP batch and backfill processing using PySpark and auditing patterns to improve ETL runtime. Built a Snowflake financial data warehouse model with SQL and supported data validation workflows to reduce downstream reporting errors. Implemented infrastructure-as-code access control policies for AWS Glue pipelines to meet data governance and HIPAA compliance. Integrated Python with Tableau for reporting and applied cost controls using Docker autoscaling and AWS Redshift lifecycle management.
pySparksince 2021
Pythonsince 2021
SQLsince 2021
Tableau
AWSsince 2021
Docker
Software Developer
•
Middle
NTT Data Services
•
Full-Time
Automated CRM data integration and reporting using MuleSoft alongside Java to improve data cycle time. Improved development workflows by integrating REST API calls with Git version control platforms and releasing new features through repository automation. Supported faster, more reliable data pipeline deployments by combining Git-based workflows with Linux automation scripts.
MuleSoft
Java
Git
Rest APIsince 2020
