Overview
Technical skills
Timeline
Roles

Overview

A data engineer building end-to-end real-time ETL pipelines focused on streaming e-commerce data with practical skills in Spark, Kafka and Postgres. The strongest proven skill is streaming ETL implementation, evidenced by a Spark Structured Streaming job that parses nested order JSON, flattens arrays, and uses a foreachBatch upsert function into Postgres. There is little evidence of testing, robust error handling, environment pinning or production-grade observability in the public code.
Phone

Technical skills

Languages
2
Python
SQL
AI/ML
3
Airflow
LLM
Spark
Other
21
AWS
Rest API
PostgreSQL
Apache Kafka
NumPy
Pandas
Snowflake
Dagster
Streamlit
DuckDB
Metabase
Plotly.js
GitHub
GitHub Actions
Docker
Tableau
Groq
Model Context Protocol
CI/CD
dbt
ETL/ELT

Timeline

Data Engineer (Contract/Freelance) • Middle
Multiple Clients • Freelance
Jan 2026 to Present 9 Months In office
Delivered data engineering, architecture, and AI integration work for multiple concurrent client engagements under NDA. Built a unified analytics platform consolidating CRM, issue tracking, and financial systems into a DuckDB warehouse with a dbt semantic layer. Implemented an MCP layer for AI-driven natural-language querying, and maintained production batch pipelines for crypto market data orchestration. Developed an AI validation MVP using Streamlit with a Groq API and added data quality remediation and Python/SQL automation workflows.
DuckDB
dbt
Model Context Protocol
GitHub Actions
Dagster
Python
SQL
Streamlit
Groq
Customer Success Lead (Data Analyst) • Lead
NorthGravity • Full-Time
Dec 2022 to Nov 2025 2 Years 11 Months In office
Managed enterprise customer accounts, covering onboarding, adoption, renewals, pre-sales, and expansions, with regular KPI-focused business reviews. Built Python/SQL automation and data ingestion pipelines using APIs, scraping sources, and legacy data into an AWS-based internal platform. Created analytics dashboards for churn, forecasting, and adoption metrics, and supported enterprise release and onboarding UAT cycles. Coordinated technical delivery with developers and created customer automations using a Zapier/Make-style low-code approach.
Python
SQL
AWS
Zapier
Anti-Fraud Analyst • Middle
ZEN.COM • Full-Time
Aug 2022 to Dec 2022 4 Months In office
Analyzed transaction data to detect and prevent fraudulent activity and supported improvements to AML detection logic. Built fraud detection rules to identify relevant patterns and collaborated with compliance teams to refine processes. Focused on improving detection accuracy through KPI-driven iteration of the logic.
Trader / Quantitative Analyst (Proprietary Trading) • Middle
STARBETA & FTMO • Full-Time
Jan 2021 to Aug 2022 1 Year 7 Months In office
Executed trading activities for proprietary trading firms across equities, futures, and options. Developed systematic trading strategies using technical and fundamental analysis. Worked with detailed knowledge of derivatives contracts and associated pricing and risk considerations for decision-making.
Financial Reporting Intern • Junior
Brown Brothers Harriman • Internship
Jan 2020 to Sep 2020 8 Months In office
Supported fund accounting processes including NAV reconciliation and related financial reporting workflows. Worked with financial datasets and compliance-oriented reporting materials. Contributed to routine reconciliation tasks and assisted with reporting preparation under internship scope.
Middle Data Scientist Confidence: Medium Data Engineer
A data engineer building end-to-end real-time ETL pipelines focused on streaming e-commerce data with practical skills in Spark, Kafka and Postgres. The strongest proven skill is streaming ETL implementation, evidenced by a Spark Structured Streaming job that parses nested order JSON, flattens arrays, and uses a foreachBatch upsert function into Postgres. There is little evidence of testing, robust error handling, environment pinning or production-grade observability in the public code.
Statistical Rigor
Correct use of statistics
Not evidenced in public code
Data Wrangling & Cleaning
5/10
Preparing and cleaning data
Clear schema-driven ingestion and flattening of nested JSON plus explicit handling of duplicate records before upsert shows competent data wrangling for streaming ETL workloads.
Evidence
realtime-ecommerce-pipeline/spark/stream_orders.py: order_schema and customer_schema StructType definitions
realtime-ecommerce-pipeline/spark/stream_orders.py: flattening logic using explode and withColumn to extract items and customer fields
Exploratory Analysis & Visualization
Exploring and visualizing data
Not evidenced in public code
Predictive Modeling
Building models that predict
Not evidenced in public code
Business Insight & Impact
Turning analysis into business value
Not evidenced in public code
Reproducibility & Notebook Hygiene
3/10
Clean, repeatable analysis
Basic reproducibility and deployment signals exist such as explicit SparkSession configuration and a runnable Kafka producer, but there is limited error handling, no environment pinning, and no CI or tests.
Evidence
realtime-ecommerce-pipeline/spark/stream_orders.py: SparkSession.builder including spark.jars.packages and forceDeleteTempCheckpointLocation
realtime-ecommerce-pipeline/generator/generate_orders.py: KafkaProducer configuration and standalone main loop for synthetic order generation
Expertise
Big Data• Middle
Streaming• Middle
Industries
Commerce• Middle
Technologies
Python• since 2022 • Middle
PostgreSQL
Spark
Apache Kafka
Metabase
Recommendations
  • Hardening streaming upserts by batching, using server-side COPY or COPY-like approaches, adding retries and connection pooling, and avoiding large collect() calls in executors.
  • Add structured logging, monitoring and alerting for streaming jobs, plus idempotency checks and backpressure handling to make the pipeline production-ready.
  • Introduce automated tests and CI, as well as pinned environment specifications (requirements or container images) and a reproducible deployment manifest.
  • Add error handling and schema evolution handling for incoming messages, and consider using a feature or metadata table to prevent schema drift and data loss.
Repositories
The developer's experience in this domain has been verified based on AI analysis of the following repositories:
Junior AI/ML Engineer Confidence: Medium Data-centric
Data-centric engineer (mid-level) specializing in real-time data engineering pipelines for e-commerce. The strongest proven skill is building streaming ingestion and ETL with Spark Structured Streaming and Kafka as shown in spark/stream_orders.py and the Kafka producer in generator/generate_orders.py. There is no evidence of ML model training, structured evaluation pipelines, unit tests, or production-grade error handling and scalability hardening in the human-authored code.
Model Architecture & Training
How well models are designed and trained
Not evidenced in public code
Data Pipeline & Feature Engineering
5/10
How data is prepared for models
Clear, practical end-to-end real-time data pipeline work: schema design, Kafka ingestion, Spark Structured Streaming transformations and flattening, and an upsert path into Postgres are implemented, but advanced streaming patterns (watermarking, late data handling), scale optimizations and robust partitioning are missing.
Evidence
realtime-ecommerce-pipeline/spark/stream_orders.py: schema definitions and Spark readStream -> from_json -> flatten -> writeStream.foreachBatch(save_to_postgres)
realtime-ecommerce-pipeline/generator/generate_orders.py: KafkaProducer usage and structured synthetic order generation
Experimentation & Evaluation
1/10
How results are measured and tested
Minimal experimentation or evaluation artifacts; only basic logging/print statements are present and there is no structured validation, metrics collection, or experiment tracking.
Evidence
realtime-ecommerce-pipeline/generator/generate_orders.py: print(order) debug output in main loop
realtime-ecommerce-pipeline/spark/stream_orders.py: batch-level print statements (f"==== BATCH {batch_id} START ====") but no metrics or evaluation pipelines
MLOps & Deployment
4/10
How models are shipped to production
Solid practical MLOps-adjacent engineering for data pipelines: Spark is configured with Kafka and Postgres connectors and the streaming job is wired to a foreachBatch upsert routine; however, deployment hardening, retries, credential management, and observability are not implemented.
Evidence
realtime-ecommerce-pipeline/spark/stream_orders.py: SparkSession.builder with spark-sql-kafka and org.postgresql jars and writeStream.foreachBatch(save_to_postgres)
realtime-ecommerce-pipeline/generator/generate_orders.py: KafkaProducer bootstrap and main loop for continuous ingestion
Computational Efficiency
1/10
How efficiently computing resources are used
Minimal computational efficiency work; code uses driver-side collect() and per-row DB inserts which will not scale and there are no batching or parallelized write optimizations.
Evidence
realtime-ecommerce-pipeline/spark/stream_orders.py: use of .collect() on DataFrame partitions and iterative psycopg2 inserts inside save_to_postgres
Research Depth & Innovation
1/10
Depth of research and new ideas
No research depth or novel algorithmic work is present; artifacts are engineering-focused and there are no paper implementations, custom layers, or experimental ML contributions.
Evidence
realtime-ecommerce-pipeline/spark/stream_orders.py and realtime-ecommerce-pipeline/generator/generate_orders.py: pipeline and generator code with no custom ML layers, training loops, or research experiments
Expertise
MLOps & Model Lifecycle• Junior
Industries
Commerce• Middle
Telecommunications• Junior
Technologies
SQL• since 2022 • Middle
Snowflake
DuckDB• since 2026
Groq• since 2026
Model Context Protocol• since 2026
Dagster• since 2026
GitHub Actions• since 2026
dbt• since 2026
CI/CD
Pandas
NumPy
AWS• since 2022
Docker
LLM
Streamlit• since 2026
GitHub
Recommendations
  • Develop production-grade streaming sinks by replacing collect() and per-row inserts with batched upserts or COPY-based writes and applying idempotent/transactional patterns.
  • Add robust error handling and retry/backoff for Kafka and Postgres interactions and move credentials to a secrets manager rather than hardcoded values.
  • Introduce observability: emit metrics, instrument latency/error logs, and integrate a metrics backend and alerting for streaming job health.
  • Create automated tests and CI for the pipeline logic, and add a small-scale performance test to measure throughput and identify bottlenecks.
Repositories
The developer's experience in this domain has been verified based on AI analysis of the following repositories:
Intern Frontend Developer Confidence: Low Generalist
Early-career developer at an intern level focused on data processing and visualization tasks. Proven experience with interactive visualizations using Plotly as described in the project README. No frontend component code, UI architecture, accessibility work, tests, or production frontend state management are present for assessment.
UI Component Architecture
How interface parts are built
Not evidenced in public code
Responsive & Cross-browser
Works on all screens and browsers
Not evidenced in public code
Performance Optimization
Speed of the interface
Not evidenced in public code
Accessibility & Semantics
Usable for everyone
Not evidenced in public code
State Management & Data Flow
Managing data in the app
Not evidenced in public code
UX & Visual Polish
Look and feel quality
Not evidenced in public code
Technologies
Plotly.js
Recommendations
  • Implement data cleaning and ETL scripts or small Python-based data pipelines using the existing data processing patterns.
  • Develop interactive analytics dashboards and visualizations leveraging Plotly for exploratory analysis.
  • Work on CSV export, PII detection scripts, and data transformation tasks where the developer can apply data manipulation skills.
Repositories
The developer's experience in this domain has been verified based on AI analysis of the following repositories: