Overview
Technical skills
Projects
Timeline
Roles

Overview

A practical data engineer working at a mid-senior level who focuses on ETL pipelines and enterprise data dimensions and who excels at building repeatable enrichment pipelines and master dimensions. The strongest proven skill is enterprise ETL and data modeling demonstrated by the Enterprise Country Dimension and the Country Weather ETL (01E and 01F notebooks) which show ISO standardization, surrogate keys, validation gates, geocoding fallbacks and robust retry logic. Public code does not evidence production infrastructure hardening such as environment pinning, automated tests, CI/CD, or any predictive-modeling experimentation and validation frameworks.
Phone

Technical skills

Languages
4
Python
SQL
C++
C
AI/ML
3
NumPy
Pandas
Scikit-learn
Analytics
3
Matplotlib
ETL/ELT
Seaborn
Frontend
3
React.js
Material UI
React Router
Other
14
PostgreSQL
Flask
Jupyter Notebook
MySQL
SQLite
styled-components
Power BI
Beautiful Soup
Plotly
GitHub
Tableau
Git
Rest API
Frontend

Projects

Aug 2026 to Present 2 Months

SupplySight AI is an enterprise-scale procurement analytics platform designed to transform traditional procurement reporting into an intelligent decision-support system.

Conventional procurement systems primarily rely on historical Enterprise Resource Planning (ERP) transactions, limiting their ability to explain changing market conditions or anticipate future procurement risks. Critical external factors such as commodity price volatility, fuel costs, currency exchange rates, weather conditions, and public holidays significantly influence procurement performance but are typically excluded from operational reporting.

To address this gap, this project integrates internal ERP procurement transactions with multiple external economic datasets through a structured enterprise ETL pipeline. The resulting analytical data warehouse provides a unified foundation for descriptive, diagnostic, predictive, and prescriptive analytics.

The project follows enterprise data engineering principles including standardized data modeling, dimensional integration, data quality validation, reproducible ETL workflows, and analytical dataset generation. The integrated dataset supports advanced feature engineering, machine learning models, executive dashboards, and procurement decision intelligence.

Jan 2026 to Mar 2026 2 Months

•Analyzed 10K+ retail records by integrating 6 datasets to build a centralized vendor analytics model

•Performed EDA and engineered 15+ KPIs including profit margin, stock turnover, and sales-to-purchase ratio

•Identified $2.7M in unsold inventory, 72% bulk-purchase savings opportunity, and 65% vendor dependency risk

•Developed an interactive Business Intelligence dashboard with 10+ visuals, improving decision-making and query performance by 30%

Nov 2025 to Dec 2025 1 Month

•Engineered a Python-based PDF extraction pipeline processing unstructured reports into structured datasets

•Built automated data cleaning and validation workflows, improving data consistency and usability for downstream reporting

•Designed scalable APIs and batch pipelines for automated ingestion and storage in PostgreSQL

•Enabled near real-time BI reporting by integrating processed data with Power BI dashboards

Sep 2025 to Oct 2025 1 Month

•Analyzed 15K+ sales transactions to uncover revenue trends and regional performance insights

•Engineered 10+ KPIs including total sales, average price, and transaction metrics using DAX

•Built an interactive dashboard with filters and slicers for dynamic exploration across cities and brands

•Delivered interactive data storytelling dashboards supporting inventory planning and marketing decisions

Timeline

Guru Gobind Singh Indraprastha University (GGSIPU)
Bachelor's Degree • Bachelor of Technology in Information Technology | CGPA: 9.21 / 10
2021–2025 New Delhi, India
Data & Analytics Intern • Junior
Learning Folks • Internship
Jun 2024 to Dec 2024 6 Months Delhi In office

Built Python-based ETL pipelines to integrate, transform, and validate datasets, improving analytics processing speed. Automated data cleaning and transformation workflows to enhance data consistency and reduce manual effort. Performed SQL-based data analysis and statistical exploration to identify business trends and insights. Prepared and delivered data-driven reports to cross-functional stakeholders to support decision-making and process improvement.

Python
SQL
ETL/ELT
Senior Data Scientist Confidence: High Data Engineer
A practical data engineer working at a mid-senior level who focuses on ETL pipelines and enterprise data dimensions and who excels at building repeatable enrichment pipelines and master dimensions. The strongest proven skill is enterprise ETL and data modeling demonstrated by the Enterprise Country Dimension and the Country Weather ETL (01E and 01F notebooks) which show ISO standardization, surrogate keys, validation gates, geocoding fallbacks and robust retry logic. Public code does not evidence production infrastructure hardening such as environment pinning, automated tests, CI/CD, or any predictive-modeling experimentation and validation frameworks.
Statistical Rigor
2/10
Correct use of statistics
Limited formal statistical rigor applied to the delivered pipelines; statistical theory content exists in learning notebooks but the production ETL notebooks contain minimal hypothesis testing, uncertainty quantification or causal analysis.
Evidence
data-analysis-with-python/Module 8 - Statistics For Data Analysis.ipynb: cells covering hypothesis testing and inferential statistics
chavijain332/SupplySight-AI/Notebooks/01F_Country_Weather_ETL.ipynb: observational text discussing data quality but no formal statistical tests
Exploratory Analysis & Visualization
4/10
Exploring and visualizing data
Exploratory analysis and written interpretation appear, but are mostly instructional or descriptive; the notebooks include clear narrative observations and some diagnostic checks rather than deep, analytic visual storytelling tied to downstream decisions.
Evidence
data-analysis-with-python/Module 5 - EDA.ipynb: multiple distribution and correlation visualizations with interpretation
chavijain332/SupplySight-AI/Notebooks/01F_Country_Weather_ETL.ipynb: notebook narrative and validation sections describing dataset completeness and business tradeoffs
Predictive Modeling
1/10
Building models that predict
Practically no predictive modeling workflow is present; only foundational ML/feature-engineering notes in tutorial notebooks but no experiments, CV design, error analysis or model calibration for production.
Evidence
data-analysis-with-python/Module 5 - EDA.ipynb: feature engineering basics and ML primer (educational content)
Business Insight & Impact
7/10
Turning analysis into business value
Clear business framing linked to procurement analytics and explicit treatment of business impact (why weather/holidays matter for procurement); artifacts show thought about integration keys and downstream analytics but no cost-of-error modelling or KPIs quantified.
Reproducibility & Notebook Hygiene
3/10
Clean, repeatable analysis
Notebooks include reproducible exports and project-path configuration, but rely on hard-coded local paths, in-notebook pip installs and lack pinned environments, tests, CI or data/versioning best practices.
Expertise
Big Data• Middle
Industries
Commerce• Middle
Data & Analytics• Middle
Technologies
Python• since 2024 • Senior
PostgreSQL
Jupyter Notebook
Seaborn
Matplotlib
Pandas
NumPy
ETL/ELT• since 2024
SQL• mentioned only
Recommendations
  • Develop and own enterprise ETL pipelines and canonical dimensions (country, holiday, weather) for analytics and reporting
  • Implement data-quality and monitoring tooling around API-driven enrichment (rate-limit handling, backoff, alerting, retries, idempotency)
  • Build reproducible deployment artifacts: containerized extraction jobs, pinned environments, unit/integration tests and a simple CI pipeline
  • Implement scalable storage/partitioning and orchestration (e.g., job scheduler / Airflow) for larger-volume enrichments and productionization
Repositories
The developer's experience in this domain has been verified based on AI analysis of the following repositories:
Middle AI/ML Engineer Confidence: High Data-centric
Data-centric engineer at a Middle level specializing in enterprise data engineering and ETL for procurement analytics. The strongest proven skill is building reliable, validated ETL and dimensional modeling pipelines as shown by the country dimension and weather ETL notebooks (SupplySight-AI/Notebooks/01E_Build_Enterprise_Country_Dimension.ipynb and 01F_Country_Weather_ETL.ipynb). There is no evidence of custom model training, distributed training, production model serving, or end-to-end ML lifecycle tooling in the available artifacts.
Model Architecture & Training
1/10
How well models are designed and trained
Minimal model-building; mostly instructional references and basic ML mentions rather than custom architectures or training loops.
Evidence
data-analysis-with-python/Module 5 - EDA.ipynb: import from sklearn.preprocessing (LabelEncoder) and ML concept references
Data Pipeline & Feature Engineering
7/10
How data is prepared for models
Strong, production-oriented data engineering and ETL work: standardized dimensions, cross-join request matrices, geocoding, API extraction, validation and programmatic correction layers.
Evidence
SupplySight-AI/Notebooks/01E_Build_Enterprise_Country_Dimension.ipynb: translator + pycountry mapping and manual country_corrections to achieve 100% ISO coverage
SupplySight-AI/Notebooks/01F_Country_Weather_ETL.ipynb: weather_matrix cross join (country × calendar) and get_coordinates(city,country) geocoding function
SupplySight-AI/Notebooks/01F_Country_Weather_ETL.ipynb: retry_weather loop with backoff and retry logic to handle 429 rate limits
Experimentation & Evaluation
3/10
How results are measured and tested
Good validation and data-quality checks (counts, uniqueness, missing-value audits) but no formal experiment tracking, A/B, or reproducible ML experiment infrastructure.
Evidence
SupplySight-AI/Notebooks/01F_Country_Weather_ETL.ipynb: enterprise validation blocks (rows, countries, duplicate keys, missing iso2 checks)
SupplySight-AI/Notebooks/01E_Build_Enterprise_Country_Dimension.ipynb: ISO validation and final validation sections
MLOps & Deployment
2/10
How models are shipped to production
Basic operationalization: CSV export, clear project paths and reproducible notebooks, but no serving, model versioning, monitoring or CI/CD for pipelines.
Evidence
SupplySight-AI/Notebooks/01F_Country_Weather_ETL.ipynb: dim_weather.to_csv(...) export to PROCESSED path
SupplySight-AI/Notebooks/01E_Build_Enterprise_Country_Dimension.ipynb: dim_country.to_csv(...) and use of PROJECT_PATH organization
Computational Efficiency
3/10
How efficiently computing resources are used
Some efficiency awareness and pragmatic optimizations (requesting full historical ranges per city, reusing geocoding, retry/backoff) but no low-level optimization or profiling.
Evidence
SupplySight-AI/Notebooks/01F_Country_Weather_ETL.ipynb: design choice to request full historical period per country to reduce API calls
SupplySight-AI/Notebooks/01F_Country_Weather_ETL.ipynb: time.sleep usage and staged retry logic to manage rate limits
Research Depth & Innovation
1/10
Depth of research and new ideas
Little to no research depth or novel algorithmic work; implementation follows established, practical ETL and data engineering patterns.
Evidence
SupplySight-AI/Notebooks/01E_Build_Enterprise_Country_Dimension.ipynb: standard use of pycountry and manual correction maps
SupplySight-AI/Notebooks/01F_Country_Weather_ETL.ipynb: straightforward API consumption and data normalization steps
Industries
Commerce• Middle
Technologies
SQL• since 2024 • Junior
C++
MySQL
Scikit-learn
Beautiful Soup
Git
SQLite
GitHub
SQL• mentioned only
Recommendations
  • Develop robust CI/CD and reproducible pipeline automation (Airflow/Prefect/Argo) to run and schedule ETL notebooks and retries.
  • Add formal experiment and artifact tracking (W&B, MLflow) if moving toward ML features, and include unit/integration tests for core transformation functions.
  • Modularize notebook logic into reusable Python modules or packages and parameterize file paths to remove hardcoded local paths and improve portability.
  • Introduce monitoring and observability for external API-dependent jobs (request success rates, rate-limit metrics, error budgets) and credential management for external services.
Repositories
The developer's experience in this domain has been verified based on AI analysis of the following repositories:
Middle Frontend Developer Confidence: Medium UI Engineer
UI-focused frontend developer at a Middle level specializing in building component-driven, responsive React interfaces. The strongest proven skill is implementing styled-components based UI components and responsive layouts as shown in src/Components/Navbar/index.js and src/Components/HeroSection/HeroStyle.js. There is little evidence of complex state management, testing, measured performance work, or formal accessibility auditing in public code.
UI Component Architecture
4/10
How interface parts are built
Component-based UI built with styled-components and small focused React components, but no evidence of a custom design system, cross-component contracts, or advanced composition patterns.
Evidence
Portfolio/src/Components/Navbar/index.js: self-contained styled-components and a single-file Navbar component with local state
Portfolio/src/Components/HeroBgAnimation/index.js: standalone HeroBgAnimation React component encapsulating an SVG animation
Portfolio/src/Components/Certificates/index.js: Cards implemented as composed styled components with hover interactions
Responsive & Cross-browser
5/10
Works on all screens and browsers
Responsive layouts and breakpoints are applied consistently across components using media queries and component-level styling, demonstrating practical cross-device considerations.
Evidence
Portfolio/src/Components/HeroSection/HeroStyle.js: multiple media query breakpoints for Hero layout and image sizes
Portfolio/src/Components/Navbar/index.js: mobile menu visibility driven by media query and local isOpen state
Netflix-Clone/style.css: media queries adjusting layout and image/video sizing for different viewports
Performance Optimization
1/10
Speed of the interface
Little to no explicit performance engineering, no measured optimizations, code-splitting, or lazy-loading; uses heavy inline SVG animation and videos without optimization evidence.
Evidence
Portfolio/src/Components/HeroBgAnimation/index.js: inline animated SVG with animateMotion but no performance instrumentation or fallbacks
Netflix-Clone/index.html and Netflix-Clone/style.css: autoplaying video elements included without obvious lazy-loading or size-optimizations
Accessibility & Semantics
2/10
Usable for everyone
Basic semantic markup and form attributes are present, but there is no evidence of ARIA on custom widgets, focus management, keyboard handling, or automated a11y tooling in CI.
Evidence
Portfolio/src/Components/Contact/index.js: semantic form inputs with required attributes but no ARIA or focus-management logic
Portfolio/src/Components/Navbar/index.js: uses nav element and anchor links but no keyboard/menu accessibility enhancements
State Management & Data Flow
2/10
Managing data in the app
Uses simple local React state and refs for UI interactions and a direct third-party email send; no evidence of advanced server-state discipline, caching, request cancellation, optimistic updates, or state machines.
Evidence
Portfolio/src/Components/Navbar/index.js: React.useState for mobile menu open/close state
Portfolio/src/Components/Contact/index.js: useRef form handling and emailjs.sendForm for direct network call without error-retry, cancellation, or optimistic rollback
UX & Visual Polish
4/10
Look and feel quality
Visual polish and micro-interactions are present with hover effects, gradients and animated SVGs; however UX edge states like loading skeletons, empty states, undo patterns or progressive enhancement are largely absent.
Evidence
Portfolio/src/Components/HeroSection/HeroStyle.js: ResumeButton gradient, hover transform and responsive typography
Portfolio/src/Components/Certificates/index.js: card hover lift, shadow and transition animations
Expertise
React• Middle
Industries
Media & Entertainment• Junior
Technologies
Frontend
React.js
Material UI
styled-components
React Router
Recommendations
  • Build marketing microsites, landing pages and portfolio sites where responsive UI and visual polish are primary requirements.
  • Implement component libraries or design-system work (design-to-code conversions) using styled-components and documented tokens.
  • Convert static UIs into production-ready apps by adding measured performance work, lazy-loading, code-splitting and basic accessibility improvements.
  • Develop small-to-medium React features that need clear UI/UX implementation rather than complex backend integration or distributed systems.
Repositories
The developer's experience in this domain has been verified based on AI analysis of the following repositories: