MLOps Engineer
6+ years exp
JavaScript
SQL
Bash
TypeScript
Python
Data Pipeline & Feature Engineering: 5/10
Experimentation & Evaluation: 5/10
MLOps & Deployment: 5/10
Active 19 days ago
+7 (988) 5224597 Invite to interview
Message
Download CVCV
Overview
Technical skills
Timeline
Roles
Overview
A capable LLM-focused engineer at a lower-senior level who builds agentic pipelines, RAG components and production MLOps glue. The strongest proven skill is building LLM-based code-review and RAG pipelines - evidenced by the LangGraph pipeline, dedup/enrichment logic (src/agent/graph.py) and a schema-enforced security scanner (src/tools/security_scanner.py). The public code shows limited evidence of custom neural-model engineering, GPU/quantization work, or advanced serving/performance benchmarking.
Technical skills
JavaScript
SQL
Bash
TypeScript• 6y+
Python• Senior
Python
FastAPI
Flask
Pydantic
Uvicorn
pySpark• 3y+
Databases
Azure SQL Database
MySQL
MS SQL
Chroma
FAISS
MongoDB
Databricks• 3y+
Delta Lake• 3y+
Pinecone
AI/ML
Claude
LangChain
LLM
MLFlow
NumPy
Pandas
Prompt Engineering
PyTorch
Scikit-learn
Spark
Streamlit
Anthropic SDK
Embeddings
LangGraph
NLP
RAG
Sentence-Transformers
Tokenization
Transformers
DevOps
AWS Fargate
Git
Rest API
AWS• 7y+
CI/CD• 6y+
Docker• 6y+
GitHub Actions• 6y+
Azure• 3y+
Azure DevOps• 3y+
Analytics
Matplotlib
Power BI
Frontend
Angular• 6y+
Timeline
Data Engineer
•
Middle
Globestar Software Limited
•
Full-Time
Designed and built production data pipelines using Azure Databricks and Azure Data Factory to support downstream analytics and ML workloads. Optimized Spark and Azure Synapse processing through partitioning and indexing, improving runtime and streamlining Dev/Prod CI/CD with Azure DevOps. Created reusable, metadata-driven ADF workflows to onboard datasets securely within network boundaries.
Databricks
Azure
pySpark
Delta Lake
Azure DevOps
Docker
GitHub Actions
CI/CD
Moscow Institute of Physics and Technology
Master's Degree •
Artificial Intelligence
Python
SQL
Azure DevOps
Azure
Databricks
pySpark
Frontend Software Engineer
•
Middle
NR Technologies Pvt Ltd
•
Full-Time
Built and maintained enterprise Angular applications integrated with backend REST APIs, focusing on performance improvements and faster page load. Streamlined build and deployment workflows by adopting Docker and GitHub Actions for consistent release lifecycles. Worked across frontend delivery and operational automation to improve reliability of deployments.
Angular
TypeScript
Dockersince 2020
GitHub Actionssince 2020
CI/CDsince 2020
SRM Institute of Science and Technology (ex SRM University)
Bachelor's Degree •
Computer Science
Senior AI/ML Engineer
Confidence: Medium LLM Engineer
A capable LLM-focused engineer at a lower-senior level who builds agentic pipelines, RAG components and production MLOps glue. The strongest proven skill is building LLM-based code-review and RAG pipelines - evidenced by the LangGraph pipeline, dedup/enrichment logic (src/agent/graph.py) and a schema-enforced security scanner (src/tools/security_scanner.py). The public code shows limited evidence of custom neural-model engineering, GPU/quantization work, or advanced serving/performance benchmarking.
Model Architecture & Training
3/10
How well models are designed and trained
Some training and model-selection work for classical models (scikit-learn) is present, but no custom neural architectures, custom training loops, or advanced optimization. Evidence shows model selection, hyperparameter grids and MLflow logging rather than research-grade model engineering.
Evidence
phishing-url-detector-mlops/networksecurity/components/model_trainer.py: ModelTrainer.train_model / track_mlflow (sklearn model selection and MLflow logging)
phishing-url-detector-mlops/networksecurity/components/model_trainer.py: evaluate_models usage and hyperparameter grids
Data Pipeline & Feature Engineering
5/10
How data is prepared for models
Clear, modular data pipeline and preprocessing code with ingestion, feature store export, train/test split and a transformation pipeline using scikit-learn Pipeline/KNNImputer.
Evidence
phishing-url-detector-mlops/networksecurity/components/data_ingestion.py: export_collection_as_dataframe / split_data_as_train_test
phishing-url-detector-mlops/networksecurity/components/data_transformation.py: get_data_transformer_object / initiate_data_transformation (preprocessor, save transformed arrays)
Experimentation & Evaluation
5/10
How results are measured and tested
Concrete evaluation and experiment tracking: MLflow integration for classic models plus an LLM evaluation harness with precision/recall/F1 computation and MLflow logging of eval runs.
Evidence
llm-code-review-agent/src/eval/mlflow_logger.py: log_eval_run uses MLflow to record eval metrics and artifacts
llm-code-review-agent/src/eval/run_eval.py and src/eval/metrics.py: structured eval harness computing precision/recall/F1
MLOps & Deployment
5/10
How models are shipped to production
Practical MLOps and deployment plumbing is present - FastAPI/Streamlit frontends, S3 sync in training pipeline, CI/CD references and MLflow tracking; but no evidence of advanced serving optimizations or production-grade orchestration beyond standard patterns.
Evidence
llm-code-review-agent/streamlit_app/app.py: Streamlit front-end calling the review API
phishing-url-detector-mlops/networksecurity/pipeline/training_pipeline.py: sync_artifact_dir_to_s3 / sync_saved_model_dir_to_s3 and pipeline orchestration
Computational Efficiency
2/10
How efficiently computing resources are used
Small signals of efficiency-awareness (embedding function selection, limited concurrency for API calls) but no measured GPU/quantization/batching work or profiling evidence.
Evidence
llm-code-review-agent/src/agent/graph.py: use of SentenceTransformerEmbeddingFunction and ThreadPoolExecutor to cap concurrent enrichment workers
llm-code-review-agent/src/rag/build_index.py: uses chromadb / embedding setup (implicit efficiency choices)
Research Depth & Innovation
2/10
Depth of research and new ideas
Good engineering patterns and practical heuristics (deduplication, UI-leak scrubbing, OWASP-grounding) but no novel algorithms or paper-level reproductions.
Evidence
llm-code-review-agent/src/agent/graph.py: deduplicate_findings with composite embedding + category short-circuit and _scrub_ui_leak
llm-code-review-agent/src/tools/security_scanner.py: combination of deterministic regex checks and LLM schema-enforced tool calls grounded in OWASP
Verified artifacts
Expertise
AI Agents & Agentic Workflows• Senior
MLOps & Model Lifecycle• Senior
RAG• Senior
LLM• All Tiers
Industries
Artificial Intelligence• Senior
Cybersecurity• Senior
Technologies
SQL
MySQL
LangGraph
Rest API
LangChain
Claude
Spark
Databricks• 3y+
Delta Lake• 3y+
Flask
Azure DevOps• 3y+
GitHub Actions• 6y+
Embeddings
Prompt Engineering
NLP
Azure• 3y+
MS SQL
CI/CD• 6y+
Transformers
Git
PyTorch
AWS• 7y+
Docker• 6y+
LLM
RAG
pySpark• 3y+
Azure SQL Database
AWS Fargate
Tokenization
AWS• mentioned only
CI/CD• mentioned only
CI/CD• mentioned only
Claude• mentioned only
Docker• mentioned only
LangGraph• mentioned only
LLM• mentioned only
MLOps• mentioned only
Recommendations
- Lead development of LLM-powered code analysis or RAG systems (agent orchestration, retrieval, prompt templates, tool integration).
- Implement production MLOps for classical ML models and experiments (MLflow tracking, training pipelines, CI/CD and S3 artifact sync).
- Own front-end + API integration for ML services (Streamlit or FastAPI dashboards that call model inference).
- Work on security-focused automation that combines deterministic checks with LLM grounding (OWASP-grounded scanners and remediation suggestions).
Repositories
The developer's experience in this domain has been verified based on AI analysis of the following repositories:
Senior Data Scientist
Confidence: Medium Data Engineer
A pragmatic data-engineering-focused developer at a solid mid/senior level who builds production MLOps pipelines and LLM-driven tooling. The strongest proven skill is production MLOps and pipeline design - evidenced by the end-to-end phishing URL detection pipeline (data ingestion, transformation, model training, MLflow logging and S3 syncing) and the LLM code-review agent architecture (tree-sitter parsing, embedding-driven deduplication, RAG retrieval). What is not evidenced is rigorous statistical experimentation, advanced causal analysis, or formal error-cost / business-metric treatment in the codebase.
Statistical Rigor
2/10
Correct use of statistics
Basic evaluation metrics and simple evaluation helpers are present, but there is little evidence of formal statistical tests, uncertainty quantification, or causal/experimental design.
Evidence
ml_project_student_math_score_predict/notebook/model_training.ipynb: evaluate_model function and model comparison
phishing-url-detector-mlops/networksecurity/utils/ml_utils/metric/classification_metric.py: get_classification_score
Data Wrangling & Cleaning
6/10
Preparing and cleaning data
Concrete ETL and ingestion components, train/test splitting, imputation and feature pipeline code show solid data engineering and cleaning practice suitable for production MLOps.
Evidence
phishing-url-detector-mlops/networksecurity/components/data_ingestion.py: export_collection_as_dataframe, split_data_as_train_test
phishing-url-detector-mlops/networksecurity/components/data_transformation.py: get_data_transformer_object and initiate_data_transformation
phishing-url-detector-mlops/networksecurity/components/data_validation.py: detect_dataset_drift usage (ks_2samp)
Exploratory Analysis & Visualization
3/10
Exploring and visualizing data
Notebooks and some inline commentary produce visualizations and observations, but most EDA is conventional; deeper hypothesis-driven analysis is limited in human-authored artifacts used for scoring.
Evidence
ml_project_student_math_score_predict/notebook/model_training.ipynb: plotting and dataframe exploration
Predictive Modeling
5/10
Building models that predict
Multiple modeling pipelines, hyperparameter grids and MLflow experiment logging are present, showing an applied modeling workflow; however, cross-validation design, leakage checks and advanced error analysis are lightweight.
Evidence
phishing-url-detector-mlops/networksecurity/components/model_trainer.py: evaluate_models integration and MLflow tracking
ml_project_student_math_score_predict/src/components/model_trainer.py: evaluate_models usage and model selection
Business Insight & Impact
2/10
Turning analysis into business value
There is pipeline and deployment awareness (MLflow, S3 syncing, FastAPI, Streamlit) but little explicit framing of business metrics, error-cost analysis, or decision-oriented recommendations in the scored code.
Evidence
phishing-url-detector-mlops/networksecurity/pipeline/training_pipeline.py: sync_saved_model_dir_to_s3 and run_pipeline orchestration
phishing-url-detector-mlops/networksecurity/components/model_trainer.py: track_mlflow method
Reproducibility & Notebook Hygiene
5/10
Clean, repeatable analysis
Reproducibility and MLOps practices are evident - requirements/setup, MLflow tracking, saved preprocessors and model artifacts - although notebooks still contain iterative cells and some hardcoded values.
Evidence
phishing-url-detector-mlops/setup.py and requirements.txt: pinned deps and packaging
phishing-url-detector-mlops/networksecurity/components/model_trainer.py: MLflow logging and model artifact saving
Expertise
Analytics• Middle
Big Data• All Tiers
Data Quality• All Tiers
Data Warehouse• All Tiers
Pipelines• All Tiers
Streaming• All Tiers
Cloud Data• All Tiers
Industries
Artificial Intelligence• Middle
Cybersecurity• Middle
Technologies
Sentence-Transformers
MLFlow
Power BI
Scikit-learn
Matplotlib
Anthropic SDK
Pandas
NumPy
Streamlit
Claude• mentioned only
LangGraph• mentioned only
LLM• mentioned only
MLOps• mentioned only
MongoDB• mentioned only
RAG• mentioned only
Recommendations
- Lead development of production MLOps pipelines and model deployment work - build retrain, CI/CD and monitoring for models (use the existing MLflow and S3 patterns).
- Implement and harden LLM-based microservices - extend the LLM code-review agent’s RAG, rate-limit handling, and observability for safe production operation.
- Build robust data validation and drift-detection tooling (automated tests, signed schemas, dataset versioning) to improve model reliability and reduce training-serving skew.
Repositories
The developer's experience in this domain has been verified based on AI analysis of the following repositories:
Middle Backend Developer
Confidence: Medium API Engineer
A practical backend engineer at a middle level who builds API-driven ML/LLM tools and MLOps components. The strongest proven skill is building LLM-based pipelines and detectors, evidenced by the StateGraph pipeline, deduplication/enrichment logic in src/agent/graph.py and the security scanner integration in src/tools/security_scanner.py. The code shows limited evidence of large-scale distributed systems design, production-grade secrets management, or systematic resilience patterns like retries/backoff and circuit breakers.
API Design
3/10
How well APIs are designed
Basic, functional API surfaces exist (FastAPI endpoints and a Streamlit client) but there is no clear versioning, idempotency, pagination strategy or consistent error contract documented in code; most API behavior is straightforward glue.
Evidence
llm-code-review-agent/src/api/main.py: FastAPI review endpoint signatures (api entrypoints)
llm-code-review-agent/streamlit_app/app.py: client calling POST /review (httpx.post usage)
little-automation-buddy/web_runner.py: FastAPI endpoints /add_tool and /do_it
Data Layer & Database
3/10
Working with databases
Data access and persistence are implemented (pymongo usage, csv/np storage, sqlite examples), but no migration history, transaction handling or explicit isolation/optimistic concurrency strategies are present.
Evidence
phishing-url-detector-mlops/networksecurity/components/data_ingestion.py: pymongo client usage and CSV feature-store export
llm-code-review-agent/src/parsing/code_parser.py: tree-sitter based parsing and structured ParsedFile dataclass
llm-code-review-agent/src/eval/holdout/snippets/ho_clean_correct_01.py: parameterized sqlite3 query example (secure pattern)
Scalability & Performance
3/10
Handling load and speed
Some scalability and performance work is present - embedding precomputation, concurrent enrichment with ThreadPoolExecutor, and use of vector stores - but there is little evidence of measured load-testing, cache invalidation strategies, or sophisticated rate-limiting/connection pooling choices.
Evidence
llm-code-review-agent/src/agent/graph.py: ThreadPoolExecutor used to parallelize enrichment and cap simultaneous API calls
llm-code-review-agent/src/agent/graph.py: use of SentenceTransformerEmbeddingFunction + deduplication via embeddings
little-automation-buddy/smart_picker.py: FAISS index construction and embedding retrieval
System Architecture
4/10
Overall system structure
Deliberate modular architecture is evident - clear separation of parsing, detection, RAG, LLM client, and enrichment via a StateGraph pipeline - demonstrating intentional decomposition though deployed operational concerns are modest.
Evidence
llm-code-review-agent/src/agent/graph.py: StateGraph pipeline (analyze -> detect -> enrich) and build_graph()
llm-code-review-agent/src/parsing/code_parser.py: language-agnostic parser abstraction using tree-sitter
phishing-url-detector-mlops/networksecurity/pipeline/training_pipeline.py: modular ML pipeline orchestration (ingest -> validate -> transform -> train)
Security & Auth
3/10
Protecting data and access
Security awareness is built into the product - deterministic regex checks and schema-enforced LLM tool calls - but there are also risky patterns and some secret/credential hygiene issues in code that weaken overall posture.
Evidence
llm-code-review-agent/src/tools/security_scanner.py: deterministic _SECRET_PATTERN and LLM scan via chat_tool
llm-code-review-agent/src/agent/prompts.py: SECURITY_* prompts and explicit OWASP-guided detection vocabulary
phishing-url-detector-mlops/networksecurity/components/model_trainer.py: environment variables and hardcoded MLflow credentials in code (credentials handling)
Reliability & Observability
3/10
Stability and monitoring
Basic observability and reliability practices are present - tests, logging, MLflow experiment tracking, and timeouts - but robust operational patterns (retries with backoff+jitter, circuit breakers, graceful shutdown, structured correlation ids) are not systematically implemented.
Evidence
llm-code-review-agent/tests/test_*: unit tests for detector, parser, metrics and end-to-end graph test
phishing-url-detector-mlops/networksecurity/components/model_trainer.py: MLflow tracking and logging usage
llm-code-review-agent/streamlit_app/app.py: httpx.post with explicit timeout and error handling around API calls
Expertise
Databases & Vector Storage• Middle
Microservices & API Architecture• Middle
Python• Middle
Industries
Artificial Intelligence• Middle
Cybersecurity• Middle
Technologies
Python• Senior
MongoDB
Chroma
FAISS
FastAPI
Pydantic
Uvicorn
AWS• mentioned only
CI/CD• mentioned only
CI/CD• mentioned only
Docker• mentioned only
Recommendations
- Lead API and LLM integration features - implement new LLM-powered review endpoints, RAG index improvements and embedding-based deduplication logic.
- Own mid-size MLOps work - expand the training pipeline, MLflow experimentization, model packaging and S3 artifact sync for production deployments.
- Hardening security review tooling - improve secrets handling, remove exec-style patterns, and add automated tests for adversarial inputs and prompt-injection defenses.
- Implement operational resilience - add retry/backoff, timeouts at all external calls, structured logging with correlation ids, and basic rate-limiting for LLM calls.
Repositories
The developer's experience in this domain has been verified based on AI analysis of the following repositories:
