Data Scientist
6+ years exp
5+ years ML exp
20+ projects
JavaScript
Python
SQL
Node JS
TypeScript
API Design: 5/10
System Architecture: 5/10
Reliability & Observability: 5/10
Active 3 days ago
+62 (821) 19658973
Invite to interview
Message
Download CVCV
Overview
Technical skills
Timeline
Roles
Overview
Senior backend engineer focused on AI-driven backend and image/video pipelines with strong production-minded FastAPI services and LLM integration. The strongest proven skill is building an end-to-end AI image generation and composition pipeline with validation - implemented in code/backend_v2/app/services/image_generator.py and the ImageCompositionService in code/backend/app/services/composition_service.py. There is limited evidence of formal infra, deployment automation, DB schema migrations, or large-scale distributed systems design in public code.
Technical skills
JavaScript• Middle • 6y+ • 5+ projects
Python• Staff • 5y+ • 20+ projects
SQL• Senior • 5y+ • 20+ projects
Node JS• Middle
TypeScript• Middle • 3 projects
Python
Aiohttp
Requests
Databases
FAISS
PostgreSQL• 3y+ • 5+ projects
Google BigQuery• 3y+ • 5+ projects
pgvector
MongoDB• 3 projects
Amazon Aurora• 3 projects
MySQL
AI/ML
Ray
Speech Recognition
Transfer Learning
Spark
OpenAI SDK
OpenCV
NumPy
Pillow
MediaPipe
Anomaly Detection• 5y+
Embeddings• 4y+
LLM• 4y+
RAG• 4y+
Time Series Forecasting• 3y+
AWS Bedrock
LangChain
Prompt Engineering
SciPy
NLP
Claude
Claude Code
Multimodal AI
Semantic Search
Cursor
LangGraph
MLFlow
PyTorch
Scikit-learn
TensorFlow
XGBoost
Frontend
PostCSS
Bootstrap• 6y+
JQuery• 6y+
Tailwind CSS• 6y+
Next.js
React.js
DevOps
Rest API
AWS• 4y+
AWS CDK• 4y+
AWS Lambda• 4y+
Docker• 4y+ • 20+ projects
Amazon EC2• 3y+ • 5+ projects
AWS Fargate• 3y+ • 5+ projects
Docker Compose• 3y+ • 20+ projects
Kubernetes
CI/CD• 5+ projects
Datadog• 3 projects
GCP
Vector
GitLab CI
Analytics
Tableau• 4y+ • 10+ projects
Design
Figma• 5+ projects
Timeline
AI Engineer
•
Middle
Apartment Specialists New Zealand
•
Full-Time
Built an AI pipeline for real-time transcription and insight extraction from buyer/agent conversations, using FastAPI-backed services with PostgreSQL storage and LangChain-based orchestration. Implemented an audio workflow for diarization, noise reduction, and speaker/segment classification. Developed a GPT-vision floor plan measurement tool and a role-based, multi-stage form approval workflow with real-time request status tracking.
FastAPI
PostgreSQL
LangChain
SciPy
Python
Next.js
React.js
Node JS
Senior AI Engineer
•
Senior
Rubythalib.ai
•
Freelance
Conducted research and internal sessions on LLM agent frameworks and production-oriented tool integrations. Prototyped an AI call center assistant using OpenAI Realtime-style streaming concepts with LangChain and supporting FastAPI endpoints for inference and monitoring hooks. Presented guidance on improving LLM robustness, latency, caching, and cost-effectiveness in deployments.
LangChain
FastAPI
Mentor Data Scientist and AI Engineer
•
Middle
Dibimbing.id
•
Freelance
Provided mentorship for cohorts in data science, machine learning, and deep learning with curriculum and learning materials development. Delivered training covering ETL/ELT, preprocessing, visualization, and modeling using Python and SQL. Led webinars and group project sessions to support applied skills and community learning over multiple cohorts.
Pythonsince 2022
SQL
Mentor Data Scientist and AI Engineer
•
Middle
Kelas.com
•
Freelance
Mentored and taught students across data science and ML fundamentals, including Python and SQL, with structured hands-on assignments. Delivered sessions on ETL/ELT, data cleaning, EDA, and supervised/unsupervised modeling. Supported learners on applied projects such as recommender systems, time-series forecasting, and transfer learning.
Python
SQL
Senior Data & AI Engineer
•
Senior
PT. Mastersystem Infotama Tbk
•
Full-Time
Designed and deployed an AWS-based chatbot using LangChain with a pgvector-backed retrieval layer and Telegram integration. Created GenAI strategies with Amazon Bedrock, exposing prompt-driven capabilities through FastAPI APIs. Delivered RAG prototypes (chatbot, document QA, OCR, and text-to-SQL) and built an LLM-assisted Kubernetes monitoring concept to speed up issue detection and reporting to internal teams via Telegram.
AWS Bedrock
LangChainsince 2024
pgvector
FastAPIsince 2024
Kubernetes
Prompt Engineering
RAG
SQL
Data Scientist & AI Engineer
•
Middle
PT. KitaLulus International
•
Full-Time
Built ETL/ELT pipelines with leakage prevention logic and implemented feature engineering and rule-based scoring for both research and production workflows. Developed an ML pipeline on AWS using CDK to provision orchestrated batch inference with Step Functions/Lambda/EventBridge patterns. Improved job-matching models using NLP approaches with embeddings and LLMs, and set up automated training and monitoring with Slack alerts plus Metabase dashboards.
AWS CDK
AWS Lambda
Embeddings
LLM
RAGsince 2022
Slack
Data Scientist
•
Middle
PT. Sharing Vision Indonesia
•
Full-Time
Validated anomalies in data visualizations used for monitoring across a large set of applications and stakeholders. Built scalable ETL/ELT pipelines using SQL tools and PySpark, and performed data mining over large customer datasets. Developed dashboards for operational metrics and improved anomaly detection by defining validation rules and automated checks across system logs.
SQLsince 2021
pySpark
Anomaly Detection
Mentor Data & Business Analytics
•
Middle
Ruangguru
•
Freelance
Created coding and learning materials for data processing and analytics using Python and SQL. Developed dashboard-oriented content using Tableau and supported learners with project reviews and guidance. Facilitated consulting sessions focused on career guidance and building portfolio-ready analytics projects.
Python
SQL
Tableau
Software Engineer - Frontend
•
Middle
PT. Anugerah Indonesia Lima
•
Full-Time
Collaborated with UI/UX teams to translate mockups into responsive web interfaces. Implemented frontend components with JavaScript and styling frameworks, including Bootstrap, Tailwind CSS, and jQuery, and connected screens to backend APIs. Delivered iterative UI integrations within sprint timelines and gathered feedback to refine usability across desktop and mobile.
JavaScript
Bootstrap
Tailwind CSS
JQuery
Institut Teknologi Bandung (ITB)
Bachelor's Degree •
Computational Physics
Middle Backend Developer
Confidence: High API Engineer
Senior backend engineer focused on AI-driven backend and image/video pipelines with strong production-minded FastAPI services and LLM integration. The strongest proven skill is building an end-to-end AI image generation and composition pipeline with validation - implemented in code/backend_v2/app/services/image_generator.py and the ImageCompositionService in code/backend/app/services/composition_service.py. There is limited evidence of formal infra, deployment automation, DB schema migrations, or large-scale distributed systems design in public code.
API Design
5/10
How well APIs are designed
Reasonable, production-oriented API design with FastAPI, response models, dependency injection and explicit error responses, but no visible API versioning strategy or advanced idempotency/pagination contracts.
Data Layer & Database
3/10
Working with databases
Data layer work exists (SQL agent, Postgres usage, FAISS vectorstore) and careful SQL/toolkit constraints, but there is no migration history, limited evidence of transaction/isolation management or hand-tuned SQL.
Scalability & Performance
4/10
Handling load and speed
Performance-aware pieces are present - async IO, aiohttp usage, FFmpeg GPU detection and CPU-fallback logic - but no explicit measured load-testing, caching-invalidation strategy or detailed connection pooling configs.
System Architecture
5/10
Overall system structure
Clear module separation and service decomposition (generators, composers, background remover, API endpoints, agent supervisor graph) with deliberate trade-offs, though not a multi-service distributed deployment.
Security & Auth
3/10
Protecting data and access
Some security awareness - Pydantic validation, API key dependency, rate limiter references - but risky patterns (fetching arbitrary external URLs for images/video) and limited evidence of thorough secrets handling or SSRF protections.
Reliability & Observability
5/10
Stability and monitoring
Good reliability and observability practices in code - structured logging, retries for OCR validation, global exception handlers and FFmpeg error fallbacks - but limited evidence of metrics/alerting, correlation ids, or circuit breakers.
Expertise
Backend AI & LLM• Middle
Databases & Vector Storage• Middle
Microservices & API Architecture• Middle
Python• Middle
Technologies
PostgreSQL• 3y+ • 5+ projects
FAISS
Requests
Aiohttp
Recommendations
- Lead development of an API service that exposes AI image-generation and composition as a scalable microservice (FastAPI + async generators + background workers).
- Implement a production deployment and observability stack - CI/CD, structured traces/correlation ids, Prometheus/Grafana metrics and alerting for the image-generation pipeline.
- Hardening for external inputs - add robust URL validation, SSRF protections, request timeouts and sandboxing for third-party downloads and subprocess calls.
- Expand DB and persistence maturity - add migrations, transactional boundaries, and explicit storage/retention policies for generated assets.
Repositories
The developer's experience in this domain has been verified based on AI analysis of the following repositories:
Intern AI/ML Engineer
Confidence: Low Generalist
Emerging developer at intern level focused on Python and AI-enabled tooling. No clear human-authored production code is present in the analyzed human-authored files to prove a specific technical strength. Public artifacts do not evidence end-to-end ML training, reproducible experiments, or MLOps practices.
Model Architecture & Training
How well models are designed and trained
Not evidenced in public code
Data Pipeline & Feature Engineering
How data is prepared for models
Not evidenced in public code
Experimentation & Evaluation
How results are measured and tested
Not evidenced in public code
MLOps & Deployment
How models are shipped to production
Not evidenced in public code
Computational Efficiency
How efficiently computing resources are used
Not evidenced in public code
Research Depth & Innovation
Depth of research and new ideas
Not evidenced in public code
Technologies
Rest API
LangChain
Spark
pgvector
OpenCV
AWS CDK• 4y+
Embeddings• 4y+
Prompt Engineering
SciPy
NLP
Speech Recognition
Transfer Learning
OpenAI SDK
AWS Bedrock
NumPy
AWS• 4y+
Kubernetes
LLM• 4y+
RAG• 4y+
MediaPipe
Ray
Pillow
AWS Lambda• 4y+
Anomaly Detection• 5y+
Time Series Forecasting• 3y+
Recommendations
- Build small, self-contained Python GUI tools that integrate a single LLM API and include unit tests and a README documenting architecture and limitations.
- Create a compact ML/LLM proof-of-concept with a clear experiment log, held-out validation, and simple reproducible training or fine-tuning steps.
- Contribute a focused, original module (for example: subtitle parsing or stable portrait conversion) with tests and CI to demonstrate ownership and depth.
- Document deployment and usage steps (requirements, config, secrets handling) and add basic error handling and logging for production readiness.
Repositories
The developer's experience in this domain has been verified based on AI analysis of the following repositories:
Intern Data Scientist
Confidence: Low Data Engineer
Junior data engineer focused on multimedia automation and AI integrations with a leaning toward desktop app tooling. The most tangible artifacts are a packaged multimedia short-clipper application and its core processing module (app.py and clipper_core.py) which show integration with yt-dlp/ffmpeg and OpenAI, but those files include low ownership signals and appear to be forked or heavily based on upstream work. There is no strong public evidence of statistical analysis rigor, production-grade pipeline design, test coverage, or original dataset collection in the available human-authored material.
Statistical Rigor
Correct use of statistics
Not evidenced in public code
Data Wrangling & Cleaning
Preparing and cleaning data
Not evidenced in public code
Exploratory Analysis & Visualization
Exploring and visualizing data
Not evidenced in public code
Predictive Modeling
Building models that predict
Not evidenced in public code
Business Insight & Impact
Turning analysis into business value
Not evidenced in public code
Reproducibility & Notebook Hygiene
Clean, repeatable analysis
Not evidenced in public code
Recommendations
- Maintain and extend desktop multimedia tooling that integrates video download, subtitle handling, and AI-powered highlight detection (app.py and clipper_core.py style workflows).
- Implement end-to-end reproducible pipelines for video metadata and transcript extraction (download -> subtitle extraction -> transcript store) with clear ownership, logging, and retries.
- Build small, testable components: unit tests for subtitle parsing and timestamp conversion, and CI to assert expected outputs for core functions.
- Document configuration and deployment steps and add example data / fixtures so future reviewers can validate changes without running the full GUI.
Repositories
The developer's experience in this domain has been verified based on AI analysis of the following repositories:
