Overview
Technical skills
Timeline
Roles

Overview

LLM engineering practitioner (senior-level) focused on building evaluation and agent benchmarking adapters and production document-extraction pipelines. The strongest proven skill is engineering evaluation and verification pipelines, demonstrated by the CRMArena verifier (harbor/adapters/crmarena/template/tests/verify.py) and multiple adapter evaluation scripts that generate Docker-based test harnesses. There is little to no evidence of custom model training, novel architectures, or experiment tracking for training runs in public code.

Technical skills

Java
Scala
C
JavaScript• Senior • 10y+ • 20+ projects
C++• Senior • 10y+ • 20+ projects
PHP• Junior • 8y+
SQL• Senior • 8y+ • 5+ projects
C#• Senior • 7y+ • 5+ projects
Node JS• Junior • 7y+
TypeScript• Senior • 6y+
Python• Middle • 6y+
Go• Junior
PHP
Laravel• 8y+
Databases
Cassandra
MySQL
PostgreSQL
Redis
AI/ML
Spark
OpenAI SDK
Datasets
huggingface_hub
Pandas
Google GenAI SDK
Anomaly Detection• 8y+
Hadoop• 8y+
MLlib• 8y+
NLP• 8y+
PyTorch• 8y+
TensorFlow• 8y+
LLM
DevOps
Azure
Azure DevOps
GCP
gRPC
Containers
AWS• 8y+
Rest API• 8y+
CI/CD• 3y+
Docker• 3y+
CloudFormation
GitHub Actions
Kubernetes
Frontend
Angular
GraphQL
Material UI
Next.js
Vue.js
React.js• 3y+

Timeline

Senior Software Engineer & AI Engineer Senior
Hire FullStacks Full-Time
Apr 2024 to Present 2 Years 4 Months Partially remote
Designed and delivered custom software solutions for client digital transformation across multiple industries. Built scalable web applications and backend services using Golang, Python, React, and PostgreSQL, including secure REST APIs for third-party integrations. Implemented AI features with LLM workflows and automation, and architected containerized deployments on AWS using Docker, Kubernetes, and AWS ECS/ECR. Created DevOps automation pipelines with GitHub Actions and CloudFormation while improving reliability, observability, and operational performance.
Go
Python
React.js
PostgreSQL
Docker
Kubernetes
AWS
CloudFormation
GitHub Actions
Redis
Rest API
LLM
CI/CD
Full Stack Developer Middle
WorkFlex Full-Time
Jan 2023 to Mar 2024 1 Year 2 Months Partially remote
Developed fully remote web applications for distributed engineering teams and remote-first operations. Delivered frontend and backend features using React, TypeScript, and Node.js, including scalable APIs and microservices with an emphasis on availability. Integrated AI-driven functionality to improve user productivity and workflow automation. Implemented monitoring and observability and supported cloud-hosted deployments on AWS with containerization and automated pipelines.
React.jssince 2023
TypeScript
Node JS
AWS
Dockersince 2023
CI/CDsince 2023
Rest API
Full Stack Developer Middle
Emerline Full-Time
Jul 2018 to Dec 2022 4 Years 5 Months Partially remote
Built custom software solutions for enterprise and startup clients in healthcare, fintech, e-commerce, and manufacturing. Developed SaaS platforms and backend services using PHP and Laravel with REST-based architecture, plus mobile-integrated solutions. Designed and implemented AI/ML features such as recommendation engines, predictive analytics, NLP, and anomaly detection, using TensorFlow, PyTorch, Spark MLlib, and distributed processing. Supported cloud deployments on AWS and engineered large-scale analytics using Hadoop and Apache Spark.
PHP
Laravel
TensorFlow
PyTorch
MLlib
Hadoop
NLP
Anomaly Detection
AWSsince 2018
Rest APIsince 2018
Georgian Technical University (GTU, ГТУ, ГПИ)
Bachelor's Degree Computer Science
2014–2018 Tbilisi, Georgia
Middle AI/ML Engineer Confidence: Medium LLM Engineer
LLM engineering practitioner (senior-level) focused on building evaluation and agent benchmarking adapters and production document-extraction pipelines. The strongest proven skill is engineering evaluation and verification pipelines, demonstrated by the CRMArena verifier (harbor/adapters/crmarena/template/tests/verify.py) and multiple adapter evaluation scripts that generate Docker-based test harnesses. There is little to no evidence of custom model training, novel architectures, or experiment tracking for training runs in public code.
Model Architecture & Training
1/10
How well models are designed and trained
Minimal custom model architecture or training; code is focused on LLM/SDK usage and inference prompts rather than building or training models.
Evidence
PDFRead/backend/app/services/lab_openai_extract.py: extract_with_openai_two_pass uses OpenAI Responses/chat API for JSON extraction rather than custom training
Multi-Agent-Swarm-Benchmark/harbor/adapters/algotune/src/algotune/utils.py: calibrate_problem_size and _create_calibration_script focus on runtime calibration rather than model training
Data Pipeline & Feature Engineering
4/10
How data is prepared for models
Practical data pipeline and parsing work; heuristics and pipeline code to extract and normalize text from pages and to prepare benchmark datasets are present.
Evidence
PDFRead/backend/app/services/page_pipeline.py: extract_pages_best_effort implements page-level extraction and filtering logic
PDFRead/backend/app/services/lab_schema.py: extract_lab_schema_heuristic contains line-level regex heuristics to parse biomarker rows
Experimentation & Evaluation
6/10
How results are measured and tested
Strong evidence of engineering evaluation pipelines and grading harnesses with reproducible verification scripts and multi-pass extraction/evaluation logic.
Evidence
Multi-Agent-Swarm-Benchmark/harbor/adapters/crmarena/template/tests/verify.py: full verifier with fuzzy/exact/privacy evaluators and LLM parse fallbacks
Multi-Agent-Swarm-Benchmark/harbor/adapters/bfcl/src/bfcl_adapter/adapter.py: dynamic evaluate script generation and run() logic for benchmark grading
MLOps & Deployment
5/10
How models are shipped to production
Good operational and deployment engineering: Docker-based calibration, task Dockerfiles, and FastAPI endpoints for extraction and job management are implemented.
Evidence
Multi-Agent-Swarm-Benchmark/harbor/adapters/algotune/src/algotune/utils.py: _build_docker_calibration_image and _run_docker_calibration_subprocess for Docker-based timing calibration
PDFRead/backend/app/routers/extraction.py: FastAPI extraction endpoint that integrates page pipeline, LLM extraction and result caching
Computational Efficiency
5/10
How efficiently computing resources are used
Practical efficiency work around runtime measurement and container resource controls; attention to CPU/threading and OOM behaviors is evident.
Evidence
Multi-Agent-Swarm-Benchmark/harbor/adapters/algotune/src/algotune/utils.py: creates calibration script, sets OMP_NUM_THREADS and parses RUNTIME_MS to steer problem sizing
Multi-Agent-Swarm-Benchmark/harbor/adapters/algotune/src/algotune/utils.py: calibrate_problem_size binary-search logic to hit target runtime
Research Depth & Innovation
3/10
Depth of research and new ideas
Some research-aligned implementations and careful replication of benchmark evaluation semantics, but no novel algorithmic contributions or custom model layers.
Evidence
Multi-Agent-Swarm-Benchmark/harbor/adapters/ace-bench/adapter.py: adapter implements upstream paper semantics and structured instruction generation
Multi-Agent-Swarm-Benchmark/harbor/adapters/algotune/src/algotune/utils.py: attempts to mirror test_outputs evaluation behavior for calibration
Expertise
AI / LLM Engineering (Agents)• Middle
Document Intelligence & OCR• Middle
Industries
Artificial Intelligence• Middle
Health Care• Middle
Technologies
Containers
Scala
MySQL
Redis
Rest API• 8y+
gRPC
Hadoop• 8y+
Spark
GCP
Cassandra
Azure DevOps
GitHub Actions
CloudFormation
NLP• 8y+
MLlib• 8y+
Azure
huggingface_hub
Datasets
CI/CD• 3y+
TensorFlow• 8y+
Pandas
PyTorch• 8y+
AWS• 8y+
Docker• 3y+
Kubernetes
LLM
Anomaly Detection• 8y+
Recommendations
  • Build LLM evaluation and benchmark adapters, verification scripts and grading harnesses that need careful reproduction of upstream metrics and Docker-matching runtimes.
  • Implement production document-intelligence services that integrate OCR, LLM two-pass extraction and FastAPI-based inference endpoints.
  • Develop Docker-based calibration and resource-limited evaluation tooling for benchmarking reference solvers and speed-sensitive tasks.
  • Avoid hiring for deep research/model-training roles; do not assign large-scale custom training or research-grade model architecture work without additional ML-training evidence.
Repositories
The developer's experience in this domain has been verified based on AI analysis of the following repositories:
Junior Data Scientist Confidence: Medium Data Engineer
LLM-backed data engineer (senior-level, focused on document-to-structured-data pipelines) specializing in PDF/OCR lab-report extraction and pragmatic LLM heuristics. The strongest proven skill is robust document extraction and LLM orchestration, demonstrated by PDFRead/backend/app/services/lab_openai_extract.py together with the pydantic schema in PDFRead/backend/app/services/lab_schema.py. There is little public evidence of formal statistical analysis, production monitoring/observability, large-scale distributed data engineering (big data), or comprehensive automated test coverage.
Statistical Rigor
1/10
Correct use of statistics
Minimal statistical rigor is present; the work focuses on robust extraction heuristics and LLM-driven parsing but contains no statistical tests, uncertainty quantification beyond simple confidence fields, multiple-comparison handling, or causal analysis.
Evidence
PDFRead/backend/app/services/lab_openai_extract.py: _schema_prompt includes confidence fields but no hypothesis testing or uncertainty propagation
PDFRead/backend/app/services/lab_schema.py: FieldValue.confidence exists but there are no statistical tests or significance checks
Data Wrangling & Cleaning
7/10
Preparing and cleaning data
Strong data wrangling and cleaning practices for document extraction are evident: careful regexes, header/context selection, fallback heuristics, two-pass LLM strategy, and safe file/job handling.
Evidence
PDFRead/backend/app/services/lab_openai_extract.py: _make_pages_text implements header selection, row-detection regex, context radius and max_rows_total safeguards
PDFRead/backend/app/services/lab_schema.py: extract_lab_schema_heuristic implements line-level parsing, normalization, evidence snippets and conservative confidence assignments
PDFRead/backend/app/services/jobs.py: start_job uses threading.Lock, persists job records and writes result files for durable pipeline behavior
Exploratory Analysis & Visualization
1/10
Exploring and visualizing data
Exploratory analysis and visualization are minimal to non-existent; the code returns structured extraction results and warnings but does not include analytic charts, interpretative write-ups, or data storytelling artifacts.
Evidence
PDFRead/backend/app/routers/extraction.py: _flatten_v2 returns a flattened payload for a frontend but contains no EDA/visualization steps
PDFRead/README.md: project describes business intent (lab result classification) but there are no notebooks or visual analysis artifacts
Predictive Modeling
2/10
Building models that predict
Predictive modeling activity is limited to LLM-based information extraction and pragmatic heuristics; there is a thoughtful two-pass extraction and validation fallback but no cross-validation, model evaluation metrics, or calibration work typical of predictive modeling projects.
Evidence
PDFRead/backend/app/services/lab_openai_extract.py: extract_with_openai_two_pass implements a two-pass LLM extraction and merging logic
PDFRead/backend/app/services/lab_schema.py: models define status and confidence but no evaluation metrics or CV code
Business Insight & Impact
3/10
Turning analysis into business value
Some product-minded decisions link outputs to business needs (e.g., classifying biomarkers into optimal/normal/out_of_range and returning warnings), but explicit business-metric reasoning (cost-of-error, FP/FN tradeoffs, SLAs) is not present.
Evidence
PDFRead/backend/app/services/lab_openai_extract.py: _schema_prompt defines classification rules mapping to 'optimal'/'normal'/'out_of_range' tied to patient age/sex
PDFRead/README.md: describes mapping extracted biomarkers to classification using age/sex reference ranges
Reproducibility & Notebook Hygiene
4/10
Clean, repeatable analysis
Reproducibility basics are addressed (requirements.txt, dotenv usage, deterministic pipeline components and small test script), but there is limited evidence of pinned environments, CI, data versioning, or extensive automated tests.
Evidence
PDFRead/backend/requirements.txt: explicit dependencies for extraction stack
GPTBackend/setup_chatgpt.py: helper script to download OpenAPI schema and create assistant_instructions; multiple load_dotenv usages appear across services
Expertise
Analytics• Junior
Technologies
PostgreSQL
Google GenAI SDK
OpenAI SDK
Recommendations
  • Build production document ingestion pipelines that convert mixed-format clinical/enterprise PDFs into validated structured records (use the existing lab_openai_extract and page_pipeline as the core extraction stages).
  • Integrate the extraction pipeline with a small data warehouse and monitoring stack to capture drift and extraction failures and add automated evaluation tests for extraction accuracy on held-out documents.
  • Develop backend integrations and safe automation endpoints (email/Gmail workflows, app-control) where careful auth, audit logs and idempotency are required, reusing the Gmail/AppControl service patterns.
Repositories
The developer's experience in this domain has been verified based on AI analysis of the following repositories:
Junior Backend Developer Confidence: Medium API Engineer
Backend API engineer (Senior-level judgment, tier_score 3.8) who focuses on integration-heavy backend systems and AI/benchmark evaluation tooling. The strongest proven skill is building modular evaluation adapters and reliability-aware tooling as shown by the Harbor adapters and calibration/evaluation logic (for example harbor/adapters/algotune/src/algotune/utils.py and crmarena/template/tests/verify.py). There is limited public evidence of owning large-scale distributed production infrastructure, formal SRE observability stacks, or cloud IaC automation.
API Design
5/10
How well APIs are designed
Design of multiple HTTP/Realtime endpoints, consistent error handlers and server-side query appliers are present, but versioning/idempotency strategies and a formal API contract are not strongly evidenced.
Data Layer & Database
5/10
Working with databases
Clear data models and migration history exist, plus DB init and connection logic; explicit transaction/isolation strategies and advanced query tuning are present but not pervasive.
Scalability & Performance
4/10
Handling load and speed
Performance-aware code exists (Docker calibration, background tasks, async endpoints) and some rate limiting, but there is limited evidence of full caching invalidation strategies or measured tuning results.
Security & Auth
5/10
Protecting data and access
Authentication and input-boundary security are actively addressed (JWT, bcrypt, middleware, tests), though token lifecycle and revocation features are only partially shown.
Reliability & Observability
4/10
Stability and monitoring
Thoughtful error handling, timeouts and fallbacks are present, plus startup/shutdown hooks and logging. Full production-grade observability (structured correlation ids, metrics/alerts) is not strongly present in the public artifacts.
Expertise
Backend AI & LLM• Junior
Python• Junior
Node.js• Junior
PHP• Junior
Microservices & API Architecture• Junior
Messaging & Real-time• Junior
Technologies
Go• Junior
Java
PHP• Junior • 8y+
Laravel• 8y+
Recommendations
  • Build integration layers and evaluation harnesses that run models/tools in containerized sandboxes (LLM evaluation adapters, dockerized calibrations and verifiers).
  • Implement backend APIs and realtime features that integrate third-party messaging platforms (WhatsApp/Telegram/Slack) and surface them via FastAPI/Node endpoints.
  • Develop dataset ingestion and evaluation pipelines for ML benchmarks, including robust retries, timeouts and Docker resource management.
  • Harden production observability and SRE practices: add structured tracing/correlation ids, Prometheus metrics and documented alerting for long-running services.
Repositories
The developer's experience in this domain has been verified based on AI analysis of the following repositories: