Overview
Technical skills
Timeline
Roles

Overview

AI Engineer | LLM Specialist | Building Autonomous Agents & Speech Analytics | Python • FastAPI • n8n Applied AI Engineer focused on building production-grade LLM solutions and autonomous agents. I transform business workflows through advanced automation (n8n), speech recognition (STT/Diarization), and computer vision. From deploying self-hosted models on GPU servers to designing RAG-based assistants, I build the full cycle of intelligent software.

Technical skills

C++
JavaScript
C
Python• Middle
SQL• Middle
C
FFmpeg
Python
Asyncio
Pydantic
aiogram
FastAPI
Databases
FAISS
PostgreSQL
Redis
SQLite
AI/ML
Computer Vision
NLP
ChatGPT
PaddlePaddle
Jupyter Notebook
Statsmodels
Time Series Forecasting
Classic ML
LightGBM
CatBoost
Keras
NumPy
OpenCV
Pandas
PyTorch
Scikit-learn
TensorFlow
XGBoost
Claude
CLIP
CUDA Toolkit
Gemini
Llama
llama.cpp
LLM
NER
ONNX
OpenRouter
PaddleOCR
Prompt Engineering
Qwen
RAG
DevOps
Git
Docker
Grafana
Nginx
Prometheus
Rest API
WebSockets
Analytics
Matplotlib
Plotly
Seaborn
Frontend
Bootstrap

Timeline

Data Quality Assessment Specialist Middle
Yandex Croud Part-Time
May 2023 to Present 3 Years 3 Months Moscow Remote only

Collected, preprocessed, and analyzed data for research and exploratory purposes. Conducted analysis and labeling of data in a search quality assessment domain to support ML and neural network training and testing. Formulated and validated hypotheses based on analysis results.

Analytics Engineer Middle
Etm Full-Time
May 2026 to Present 3 Months Saint Petersburg Remote only

Developed business process analytics and workflow descriptions using BPMN. Built and maintained automation workflows in n8n enhanced with LLM-based nodes, connected n8n to internal systems, and handled testing and deployment. Debugged workflow logic and rapidly iterated process changes based on operational needs.

AI Specialist (Data Scientist) Middle
Fianit-Lombard Full-Time
May 2025 to May 2026 1 Year Chelyabinsk Remote only

Designed and built end-to-end AI services from business requirements to production deployment with Docker and GPU infrastructure. Developed and integrated LLM solutions with prompt engineering, multi-model pipelines, and fallback/streaming strategies. Built NLP/NLU pipelines, document and fraud-related computer vision services (OCR and visual search with CLIP+FAISS), speech pipelines, and microservices using FastAPI with WebSocket and Redis queues; also implemented monitoring and self-hosted LLM operations and customer-facing Telegram bots and web interfaces.

Python
FastAPI
Docker
PostgreSQL
SQLite
Redis
WebSockets
LLM
Claude
Gemini
Llama
Qwen
OpenRouter
llama.cpp
RAG
Prompt Engineering
FAISS
CLIP
PyTorch
ONNX
PaddleOCR
CUDA Toolkit
NER
FFmpeg
aiogram
Nginx
Prometheus
Grafana
Git
Rest API
Data Analyst / Data Scientist Middle
Decartus Full-Time
Oct 2024 to Apr 2025 6 Months Sevastopol In office
Integrated and preprocessed datasets from multiple sources, including cleaning missing values and anomalies and performing transformations for modeling readiness. Conducted exploratory analysis and time-series analysis to identify seasonal and cyclic patterns and built forecasting models such as ARIMA/SARIMA and LSTM. Developed prediction use cases (including yield and frost risk) and delivered AI assistant solutions with an OpenAI-based approach using RAG plus a web interface with voice input and spoken streaming responses.
Pythonsince 2024
SQL
Pandas
Scikit-learn
NumPy
Matplotlib
TensorFlow
PyTorchsince 2024
Keras
Seaborn
Plotly
XGBoost
CatBoost
OpenCV
Gitsince 2024
Middle AI/ML Engineer Confidence: Medium Data-centric
A data-focused ML practitioner at a solid middle level who builds end-to-end classical ML solutions in notebooks. The strongest proven skill is applied tabular and time-series modeling with concrete artifacts such as the churn classification notebook (upsample/downsample, RF hyperparameter search and evaluation) and the taxi forecasting notebook (resampling, lag features, TimeSeriesSplit and GridSearchCV). There is little evidence of production MLOps, model serving, experiment tracking, or custom deep learning work in public code.
Model Architecture & Training
4/10
How well models are designed and trained
Competent end-to-end application of classical ML models with systematic hyperparameter search and imbalance handling; no custom model architectures, deep learning training loops or novel loss/optimizer work.
Evidence
churn_bank_customers/churn_bank_customers.ipynb: grid search loops for RandomForestClassifier (n_estimators, max_depth) and DecisionTreeClassifier depth sweep
churn_bank_customers/churn_bank_customers.ipynb: final RandomForestClassifier training and metric calculation (F1, AUC-ROC)
predictions_orders_taxi/prediction_orders_taxi.ipynb: GridSearchCV for RandomForestRegressor and model selection with TimeSeriesSplit
Data Pipeline & Feature Engineering
5/10
How data is prepared for models
Clear, practical data work - cleaning, encoding, scaling, resampling and time-series feature engineering (lags, rolling mean) were implemented and evaluated.
Evidence
churn_bank_customers/churn_bank_customers.ipynb: data cleaning, pd.get_dummies, StandardScaler usage and removal of anomalous Tenure rows
churn_bank_customers/churn_bank_customers.ipynb: upsample and downsample functions to address class imbalance
predictions_orders_taxi/prediction_orders_taxi.ipynb: make_features function creating lag_*, rolling_mean, day/dayofweek/hour and resample('1H')
Experimentation & Evaluation
4/10
How results are measured and tested
Reasonable experimentation and evaluation practices - cross-validation, time-series CV, baseline checks and multiple metrics are present, but there is no formal experiment tracking or reproducibility tooling.
Evidence
predictions_orders_taxi/prediction_orders_taxi.ipynb: use of TimeSeriesSplit and GridSearchCV with printed best scores and CPU timing
churn_bank_customers/churn_bank_customers.ipynb: confusion_matrix, precision/recall/F1 and ROC-AUC plotted and compared across methods
MLOps & Deployment
1/10
How models are shipped to production
Almost no MLOps or deployment artifacts - models are trained and evaluated interactively but there is no saving/serialization, serving code, CI/CD, or monitoring.
Evidence
churn_bank_customers/churn_bank_customers.ipynb: model_final and model_test are trained and evaluated but no model.save/pickle/export steps are present
Computational Efficiency
1/10
How efficiently computing resources are used
Minimal computational-efficiency work - basic GridSearchCV reporting and occasional timing outputs appear, but no GPU/parallel tuning, batching, quantization, or memory profiling.
Evidence
predictions_orders_taxi/prediction_orders_taxi.ipynb: GridSearchCV cell includes CPU times in outputs for the search
Research Depth & Innovation
1/10
Depth of research and new ideas
No research-depth artifacts - no custom layers, paper re-implementations, or novel algorithms; work applies standard, well-known methods.
Evidence
Both notebooks: use of standard scikit-learn estimators (DecisionTree, RandomForest, LogisticRegression, Ridge) and common regressors (CatBoostRegressor, LGBMRegressor) without custom model code
Industries
Financial Services• Middle
Transportation & Logistics• Middle
Technologies
SQL• Middle
C++
PostgreSQL
Redis
Rest API
llama.cpp
Claude
Qwen
ChatGPT
OpenCV
CUDA Toolkit
Jupyter Notebook
XGBoost
FastAPI
Prometheus
WebSockets
Prompt Engineering
Computer Vision
NLP
NER
ONNX
Llama
TensorFlow
NumPy
Keras
Git
SQLite
PaddlePaddle
PyTorch
Docker
Gemini
Nginx
Grafana
LLM
RAG
CLIP
PaddleOCR
OpenRouter
Recommendations
  • Lead development of production-ready ML pipelines for tabular and time-series models including model serialization, CI/CD for models, and reproducible experiment tracking (MLflow or W&B).
  • Build and harden time-series forecasting services for transportation or demand prediction with scheduled retraining and drift monitoring.
  • Implement MLOps practices: add model versioning, unit tests for data pipelines, and automated evaluation and deployment workflows.
  • Extend skills toward deep-learning productionization if needed by adding reproducible training scripts (notebooks -> python modules), GPU-aware training, and experiment logging.
Repositories
The developer's experience in this domain has been verified based on AI analysis of the following repositories:
Middle Data Scientist Confidence: Medium ML Practitioner
A practical ML practitioner at an early middle level who independently builds end-to-end modeling experiments and evaluations. The strongest proven skill is supervised-model development and evaluation, evidenced by the end-to-end churn classification notebook (data cleaning, imbalance handling, ROC/confusion analysis, up/downsampling) and the time-series forecasting notebook (resampling, feature lags, GridSearchCV with TimeSeriesSplit). Limitations include limited production/hardening evidence - no dependency pinning, pipeline or deployment artifacts, and several hard-coded local paths that reduce reproducibility.
Statistical Rigor
5/10
Correct use of statistics
Reasonable metric-driven evaluation: baseline checks, ROC/confusion-matrix analysis and a stationarity test are present, but there is little uncertainty quantification, hypothesis testing beyond ADF, or causal/experimental design.
Evidence
predictions_orders_taxi/prediction_orders_taxi.ipynb: adfuller stationarity test (statsmodels.tsa.stattools.adfuller)
churn_bank_customers/churn_bank_customers.ipynb: confusion_matrix/roc_auc_score and baseline constant-model accuracy checks
Data Wrangling & Cleaning
6/10
Preparing and cleaning data
Solid practical data wrangling: NA handling, anomaly filtering, resampling, scaling and explicit upsampling/downsampling functions and feature engineering; some hard-coded local paths reduce portability.
Evidence
churn_bank_customers/churn_bank_customers.ipynb: Tenure missing/zero handling and df.dropna + df.query('Tenure != 0') and StandardScaler fit/transform
predictions_orders_taxi/prediction_orders_taxi.ipynb: resample('1H') usage and make_features(...) function creating lags and rolling means
Exploratory Analysis & Visualization
6/10
Exploring and visualizing data
Exploratory plots are numerous and accompanied by written interpretation (trend/seasonality comments, class-balance discussion); analyses are descriptive and targeted at model selection rather than deep causal storytelling.
Evidence
predictions_orders_taxi/prediction_orders_taxi.ipynb: seasonal_decompose plots, daily/hourly plots and rolling mean/std visualizations with comments
churn_bank_customers/churn_bank_customers.ipynb: class frequency bar plots, ROC curves and textual interpretation after metrics
Predictive Modeling
6/10
Building models that predict
Good predictive workflow: baseline-first checks, hyperparameter search, TimeSeriesSplit for time series, explicit imbalance strategies (class_weight, up/downsampling) and final test evaluation; lacks advanced model diagnostics (calibration, feature-importance stability or systematic error analysis).
Evidence
predictions_orders_taxi/prediction_orders_taxi.ipynb: GridSearchCV with TimeSeriesSplit for RandomForestRegressor and other models
churn_bank_customers/churn_bank_customers.ipynb: upsample/downsample functions and grid loops for RandomForestClassifier hyperparameter search; use of class_weight='balanced'
Business Insight & Impact
4/10
Turning analysis into business value
Some connection to business objectives is stated (marketing retention for churn, RMSE threshold and driver staffing for taxi), but there is no quantitative cost-benefit, FP/FN cost modeling, or SLAs defined.
Evidence
churn_bank_customers/churn_bank_customers.ipynb: conclusion relating model output to marketing actions and noting precision/recall tradeoffs
predictions_orders_taxi/prediction_orders_taxi.ipynb: explicit RMSE target (<=48) and discussion about driver allocation
Reproducibility & Notebook Hygiene
3/10
Clean, repeatable analysis
Some reproducibility practices present (explicit random_state values, seeded shuffles), but notebooks use hard-coded local paths, have no dependency/environment pinning, no pipeline scripts or data/versioning tooling.
Evidence
predictions_orders_taxi/prediction_orders_taxi.ipynb: RANDOM_STATE = 12345 and GridSearchCV but uses '/datasets/taxi.csv' path
churn_bank_customers/churn_bank_customers.ipynb: StandardScaler with fit/transform and many uses of random_state but also uses an absolute local path '/Users/...' in data load
Industries
Financial Services• Middle
Transportation & Logistics• Middle
Technologies
Classic ML
CatBoost
Scikit-learn
Seaborn
Matplotlib
LightGBM
Plotly
Pandas
Statsmodels
Time Series Forecasting
Recommendations
  • Package the notebooks into repeatable scripts or a small pipeline (data ingestion - preprocessing - training - evaluation) and add a requirements.txt or environment.yml.
  • Replace absolute/local file paths with parametrized dataset locations and add lightweight data-versioning (DVC or explicit checksums) to improve reproducibility.
  • Add model diagnostics beyond point metrics - calibration plots, permutation feature importance stability, and a simple FP/FN cost analysis for business decisioning.
  • For time-series work, add rolling-backtest evaluation and persistence (model save/load) steps to demonstrate production readiness.
Repositories
The developer's experience in this domain has been verified based on AI analysis of the following repositories:
Middle Backend Developer Confidence: Medium API Engineer
A practical backend developer at a middle level who builds conversational AI integrations and simple production-ready bots. The strongest proven skill is integrating LLMs and vector search for contextual responses, as shown by VineyardAssistant.initialize_vector_store and process_query wiring to OpenAI and FAISS. The code lacks production-grade operational concerns such as retries with backoff, metrics/tracing, migration/versioning for data, and automated tests.
API Design
3/10
How well APIs are designed
Basic, pragmatic API/interaction design for a chat assistant and Telegram handlers is present (command registration, consistent handler patterns and a message contract to the LLM) but lacks formal versioning, idempotency keys, pagination strategies, or standardized error contracts; many choices are single-process and application-specific rather than platform-grade.
Data Layer & Database
3/10
Working with databases
Shows a simple data layer with local FAISS vector store, text-splitting and document ingestion; there is no migration history, transactional semantics, or explicit index tuning beyond saving/loading a local FAISS index.
Scalability & Performance
3/10
Handling load and speed
Some performance-minded choices are present - local FAISS index, lru_cache, and a simple in-memory cache - but there is no distributed caching/invalidation strategy, no queue-based decoupling, nor documented load testing or connection pooling for production workloads.
System Architecture
3/10
Overall system structure
Code is split into logical modules (bot, assistant core, console, config) and uses a SessionManager abstraction; however the architecture remains single-process, tightly-coupled and lacks service contracts, config-driven service decomposition, or explicit secret/config management beyond dotenv.
Security & Auth
3/10
Protecting data and access
Basic security hygiene is present - use of environment variables and dotenv, minimal input validation for queries and text cleaning - but there is no secret rotation, dependency vulnerability auditing, explicit input sanitization against prompt injection, nor token lifecycle handling for API keys.
Reliability & Observability
4/10
Stability and monitoring
Reasonable reliability and observability practices for a small app are included: structured logging, graceful shutdown handlers, session expiration, and exception logging; however there are no metrics, traces, retry/backoff strategies for external calls, or alerting/SLIs defined.
Expertise
Backend AI & LLM• Middle
Databases & Vector Storage• Middle
Messaging & Real-time• Middle
Python• Middle
Industries
Farming & Agriculture• Middle
Technologies
Python• Middle
FAISS
Asyncio
Pydantic
aiogram
Recommendations
  • Develop AI-backed conversational services and Telegram/chatbot integrations that use FAISS or vector DBs for semantic retrieval and an LLM gateway.
  • Harden LLM production pipelines by adding retry/backoff with jitter for external API calls, circuit breakers, metrics and tracing, and a distributed cache or vector DB with invalidation strategy.
  • Implement secure secret handling and key lifecycle (avoid dotenv in production), add input sanitization and prompt-safety layers to reduce prompt-injection risks.
  • Extend the SessionManager and storage to use a durable store (Redis or a database) and add tests and CI to validate behavior and regression safety.
Repositories
The developer's experience in this domain has been verified based on AI analysis of the following repositories: