Overview
Technical skills
Timeline
Roles

Overview

A senior-level ML practitioner focused on evidence-grounded NLP tooling and biomedical ML pipelines, with a strength in building robust, testable retrieval-and-verification architectures. The strongest proven skill is engineering an evidence-grounded RAG/agent stack with retrieval, structured prompts, grounding verification, and a Streamlit UI as implemented in the PSC Evidence Assistant (src/agent.py, src/rag.py, src/llm.py, src/prompts.py). Limited or no public evidence of production deployment (CI/CD, containerization, cloud infra), systematic experimental design (formal multiple-comparison corrections), or extensive hyperparameter tuning and model-selection infrastructure is present.
Phone

Technical skills

Languages
7
Python
C++
JavaScript
SQL
C
Bash
MATLAB
AI/ML
17
XGBoost
LLM
NLP
RAG
NumPy
Pandas
Accelerate
Datasets
TRL
Scikit-learn
TF-Keras
Streamlit
Keras
PEFT
Hugging Face
Knowledge Graph
Reinforcement Learning
Other
23
MPI
Docker
Pydantic
PyTorch
TensorFlow
PyTorch C++
TensorFlow C++
GitHub
Matplotlib
Bootstrap
CUDA
CUDA Toolkit
SciPy
Deep Learning
AI/ML
Plotly
Seaborn
Git
Linux
Machine Learning
CI/CD
HPC
SLURM

Timeline

Biomedical Data Scientist • Middle
Biomedical Informatics and Genetic Redefinition (IKMB) • Full-Time
Aug 2023 to Present 3 Years 1 Month Kiel In office
Built and evaluated a deep reinforcement learning model for sequential decision-making, optimizing policies from multivariate state inputs while accounting for delayed effects and action interactions. Developed an LLM-powered assistant that combines RAG, structured extraction, knowledge graphs, and verification to produce traceable answers grounded in source evidence. Created reproducible data pipelines for large datasets and packaged workflows with Docker for consistent execution across environments.
Reinforcement Learning
LLM
RAG
Knowledge Graph
Docker
University of Potsdam
Doctoral Degree (PhD) • Bioinformatics
2019–2023 Potsdam, Brandenburg
Optimization Scientist • Middle
Mathematical Modeling & Systems Biology, Max Planck Institute (MPI Potsdam) • Full-Time
Jun 2019 to Jun 2023 4 Years Potsdam In office
Designed optimization and constraint-based modeling approaches, including mixed-integer linear programming and nonlinear optimization, to analyze large-scale network trade-offs and system constraints. Built an interpretable machine-learning framework using XGBoost to predict missing gene–protein–reaction associations in genome-scale metabolic networks from network-derived features. Delivered teaching on optimization and graph-based methods, focusing on translating complex real-world problems into computational and data-driven solutions.
XGBoost
Optimization Scientist • Middle
MPI Potsdam • Full-Time
Jun 2019 to Jun 2023 4 Years Potsdam In office
Designed and applied mathematical optimization methods, including MILP and nonlinear optimization, to analyze and improve trade-offs in large-scale network systems under competing objectives and constraints. Built an interpretable machine-learning approach using XGBoost to infer missing gene–protein–reaction associations from network-derived biological features. Conducted teaching focused on optimization and graph-based methods, emphasizing translation of real-world problems into computational models.
XGBoost
Machine Learning Engineer • Middle
Werner Siemens Imaging Center • Full-Time
Jun 2018 to Jun 2019 1 Year In office
Designed deep-learning pipelines for large-scale image analysis, using CNN-based and autoencoder-based architectures to support automated feature extraction and prediction. Improved performance and scalability through more efficient data handling and systematic benchmarking to maintain reproducible results across large datasets. Implemented computational geometry and image-processing methods to enable automated detection and segmentation for downstream analytical workflows.
Computational Biology Researcher • Middle
Computer Science Department, University of Tehran • Full-Time
May 2014 to May 2018 4 Years Tehran In office
Developed and evaluated classification models such as random forest and logistic regression for binary prediction tasks, emphasizing feature selection, validation, and model comparison. Worked with large high-dimensional datasets by integrating multiple data sources and transforming raw features into structured representations to identify meaningful patterns. Built NLP pipelines for extracting, normalizing, and structuring information from unstructured and semi-structured data sources.
NLP
Computational Biology Researcher • Middle
University of Tehran • Full-Time
May 2014 to May 2018 4 Years Tehran In office
Developed and evaluated machine-learning models for binary prediction tasks, focusing on feature selection, validation, and model comparison. Analyzed a large high-dimensional dataset by integrating multiple sources and transforming raw variables into structured representations to uncover meaningful patterns. Built NLP pipelines for automated extraction, normalization, and structuring of information from unstructured and semi-structured data.
NLP
University of Tehran
Master's Degree • Computer Science
2013–2016 Tehran, Iran
Software Engineer & Applied Scientist • Middle
Konur • Full-Time
Mar 2007 to Sep 2013 6 Years 6 Months Tehran In office
Developed reusable backend libraries in C++ that allowed engineers to define custom computational formulations and automate previously manual analytical workflows. Applied signal-processing and algorithmic methods to diagnose and resolve latency issues in live video-broadcasting systems. Designed optimization approaches for scheduling in video-conferencing across geographically distributed sites to reduce conflicts and improve resource utilization.
C++
Senior AI/ML Engineer Confidence: Medium LLM Engineer
LLM-focused engineer at a senior level with strong practical experience building retrieval-augmented agents and applied biomedical ML pipelines. The most proven strength is evidence-grounded RAG and agent engineering as shown by the PSC Evidence Assistant agent, prompt schemas and verification logic. What is not evidenced is production-grade distributed training, formal experiment tracking or rigorous GPU/quantization optimization work.
Model Architecture & Training
4/10
How well models are designed and trained
Model architecture and training evidence is solid for applied convolutional autoencoders and custom training loops but mostly at the experimental/script level rather than production-grade research frameworks.
Data Pipeline & Feature Engineering
5/10
How data is prepared for models
Extensive domain-specific data pipelines and preprocessing for whole-slide images are implemented, including thresholding, convex hull masking and patch extraction with attention to large-file handling.
Experimentation & Evaluation
4/10
How results are measured and tested
Reasonable experiment and evaluation code exists with PSNR/MSE metrics and saved training artifacts, plus unit tests for the PSC assistant, but no centralized experiment tracking or extensive reproducible experiment metadata.
MLOps & Deployment
4/10
How models are shipped to production
Clear MLOps/serving orientation for the RAG/agent project with a Streamlit app, a pluggable LLM client, config handling and CLI runners for DEviRank; suitable for demos and light deployment.
Computational Efficiency
3/10
How efficiently computing resources are used
Some attention to memory and batch handling for large WSIs and manual batching in training loops exists, but there is little evidence of measured GPU profiling, advanced batching strategies, quantization or multi-node/distributed optimization.
Research Depth & Innovation
5/10
Depth of research and new ideas
Strong research-implementation evidence in a paper-style bioinformatics method and a carefully designed evidence-grounded RAG agent with grounding verification and structured prompt schemas.
Verified artifacts
Expertise
AI Agents & Agentic Workflows• Senior
RAG• Senior
Industries
Biotechnology• Senior
Artificial Intelligence• Senior
Health Care• Senior
Technologies
SQL• since 2026 • Junior
C++• since 2007 • Junior
MATLAB• since 2026 • Junior
CUDA Toolkit
XGBoost• since 2019
Reinforcement Learning• since 2023
SciPy
NLP• since 2014
SLURM
CI/CD
TensorFlow• since 2026
Git
PyTorch• since 2026
Docker• since 2023
LLM• since 2023
RAG• since 2023
TF-Keras
TensorFlow C++
PyTorch C++
CUDA
GitHub
HPC
Knowledge Graph• since 2023
Linux
Machine Learning• since 2026
Recommendations
  • Build evidence-grounded RAG agents and tool-using LLM assistants that require query planning, grounding verification and knowledge-graph extraction.
  • Develop end-to-end biomedical image preprocessing and patch extraction pipelines for whole-slide imaging, including efficient IO and memory-aware chunking.
  • Implement reproducible computational-biology analysis pipelines and statistical proximity scoring for drug prioritization with robust CSV/io and CLI runners.
  • Prototype Streamlit-based demo apps and integrate them with pluggable LLM backends for user-facing workflows and internal research demos.
Repositories
The developer's experience in this domain has been verified based on AI analysis of the following repositories:
Senior Data Scientist Confidence: High ML Practitioner
A senior-level ML practitioner focused on evidence-grounded NLP tooling and biomedical ML pipelines, with a strength in building robust, testable retrieval-and-verification architectures. The strongest proven skill is engineering an evidence-grounded RAG/agent stack with retrieval, structured prompts, grounding verification, and a Streamlit UI as implemented in the PSC Evidence Assistant (src/agent.py, src/rag.py, src/llm.py, src/prompts.py). Limited or no public evidence of production deployment (CI/CD, containerization, cloud infra), systematic experimental design (formal multiple-comparison corrections), or extensive hyperparameter tuning and model-selection infrastructure is present.
Statistical Rigor
5/10
Correct use of statistics
Implements sampling-based significance and z/p-value calculations for proximity scoring but lacks formal uncertainty propagation, multiple-comparison corrections, and explicit statistical test assumptions.
Data Wrangling & Cleaning
6/10
Preparing and cleaning data
Robust read/write helpers and careful data-path handling plus numerous image-preprocessing scripts implementing patch extraction, thresholding and saving; some scripts use hard-coded absolute paths and ad-hoc loops.
Exploratory Analysis & Visualization
4/10
Exploring and visualizing data
Contains plotting and saved diagnostic images (training/validation plots, PSNR) and many image-output steps, but little written interpretation or structured EDA narrative accompanying results.
Predictive Modeling
5/10
Building models that predict
Multiple implemented predictive models (convolutional autoencoders, multi-decoder architectures) and explicit training loops; evaluation is basic (train/test split, PSNR) with limited cross-validation, systematic hyperparameter search, or calibration analysis.
Business Insight & Impact
2/10
Turning analysis into business value
Strong domain relevance (drug ranking, histopathology) but no explicit business-metric framing, cost-of-errors analysis, or operational impact quantification.
Reproducibility & Notebook Hygiene
7/10
Clean, repeatable analysis
Good reproducibility practices in places: CLI entrypoints, path resolution, atomic CSV writes, Streamlit app with caching, pydantic models and unit tests; environment requirements are declared for some projects.
Expertise
Data Science• Senior
Industries
Health Care• Senior
Technologies
AI/ML
Deep Learning
Scikit-learn• since 2026
PEFT
TRL
Accelerate
Datasets
Pandas
NumPy
Keras
Streamlit
Pydantic
Hugging Face
Recommendations
  • Lead development of RAG-based, evidence-grounded assistants and verification tooling for scientific corpora, leveraging the PSC Evidence Assistant architecture.
  • Implement and productionize reproducible pipelines and model serving (Docker, CI pipelines, model registry/MLFlow or similar) for the bioimaging autoencoder work.
  • Extend DEviRank into a reproducible analysis pipeline with formal statistical controls (multiple-testing correction, confidence intervals, and clear experiment logging) and scalable sampling (parallelization, checkpointing).
Repositories
The developer's experience in this domain has been verified based on AI analysis of the following repositories:
Senior Backend Developer Confidence: Medium Generalist
ML and bioinformatics-focused backend developer at a middle-to-senior research level with a strength for building reproducible experiment pipelines and algorithmic tooling. The strongest proven skill is implementing reproducible, domain-aware computational pipelines and analytics as shown by the DEviRank pipeline and its robust CSV/io and graph-processing utilities (seirana/DEviRank/scr/DEviRank.py). There is little evidence of production web APIs, deployment manifests, structured observability or database migration histories in public code.
API Design
2/10
How well APIs are designed
Minimal API design; mostly CLI entrypoints and script-style interfaces with no HTTP/REST contracts, versioning, or idempotency patterns.
Data Layer & Database
4/10
Working with databases
Reasonable data handling using pandas and robust CSV helpers with validation and atomic writes, but no database migrations or explicit transactional design.
Scalability & Performance
4/10
Handling load and speed
Some scalability-minded choices exist - chunking, degree-aware random sampling and caching - but no production-grade distributed or async scaling, rate limiting or measured optimization artifacts.
System Architecture
5/10
Overall system structure
Clear module boundaries and separation between scripts and library code, dataclass usage and small reusable components show deliberate, research-oriented architecture appropriate for reproducible experiments.
Security & Auth
2/10
Protecting data and access
Basic defensive checks and input validation are present but there is little evidence of authentication, secrets management, or comprehensive boundary security hardening.
Reliability & Observability
4/10
Stability and monitoring
Reliability for experiments is considered: atomic writes, deterministic seeds, model checkpointing and metrics dumps exist, but there is limited evidence of structured logging, retries with backoff, timeouts or alerting for production services.
Expertise
Backend AI & LLM• Middle
Python• Middle
Industries
Artificial Intelligence• Middle
Biotechnology• Middle
Health Care• Middle
Technologies
Python• since 2019 • Senior
Recommendations
  • Build research-to-production bridges: add deployment, structured logging, and simple REST/CLI-to-service wrappers around core pipelines to enable repeatable runs in CI/CD.
  • Harden data layer operations: add migration history, schema validation tests and idempotent ingest paths for CSV/NPZ artifacts.
  • Productize ML artifacts: add model serving endpoints or batch-serving runners, plus observability (metrics and alerts) around training pipelines.
  • Expose experiment configs and reproducibility: publish parameterized experiment runners and small integration tests that exercise end-to-end flows.
Repositories
The developer's experience in this domain has been verified based on AI analysis of the following repositories: