Overview
Technical skills
Timeline
Roles

Overview

A practical data-engineering practitioner at a mid (middle) level who produces reliable preprocessing utilities and end-to-end experimental notebooks. The strongest proven skill is building data pipelines and preprocessing tooling, evidenced by the DataSplitter, schema helpers and serialization utilities in the cardiology project. What is not evidenced is rigorous statistical experimentation design, production-grade CI/CD for models, or sustained ownership of LLM/vector-db server infra (several components appear externally authored).

Technical skills

SQL
Python• Middle
Node JS• Middle
JavaScript
Python
Celery
Ruff
structlog
pySpark
Uvicorn
Requests
FastAPI
Pydantic
SQLAlchemy
Node JS
Express
Axios
Commander.js
Databases
FAISS
TimescaleDB
Weaviate
Qdrant
MongoDB
SQLite
pgvector
PostgreSQL
Redis
Supabase
AI/ML
AI Agents
CNN
CrewAI
Diffusion Models
DVC
Function Calling
Hallucination
LangChain
LangGraph
Multimodal AI
NLTK
NumPy
Pandas
Prompt Engineering
RAG
spaCy
Transformers
Vertex AI
YOLO
Deep Learning
Gradio
huggingface_hub
LLM Apps
Gaussian Splatting
OpenCV
PyTorch
Scikit-learn
Groq
Llama
Ollama
Streamlit
Claude
Computer Vision
Embeddings
Gemini
Grok
Langfuse
LLM
Model Context Protocol
NLP
DevOps
CI/CD
Docker
GCP
Git
Kubernetes
Rest API
Azure
Docker Compose
GitHub Actions
Nginx
Prometheus
Analytics
Matplotlib
Seaborn
Plotly
Cybersecurity
GDPR
Trivy
SonarQube
Frontend
React.js
Three.JS
Robotics
COLMAP
Open3D

Timeline

AI Engineer (LLM Agents Architecture, Production) Lead
McLeuker AI Internship
Jul 2026 to Present 1 Month Paris Remote/Hybrid

Audited an agent-memory architecture for LLM agents in production, including filesystem-backed memory, context compaction, and cross-session long-term retrieval using PostgreSQL with pgvector and hybrid search. Analyzed inference cost leaks and implemented prompt-caching stability across multi-provider routing. Designed an agentic loop with parallel tool dispatch, streaming via SSE, and retrieval patterns including typed relations.

Python
FastAPI
Supabase
PostgreSQL
pgvector
Embeddings
Claude
Grok
Gemini
Jan 2026 to Present 7 Months

Développé et déployé 22 modules métier : POS, carte NFC, WhatsApp bot darija, OCR de factures via GPT-4o Vision, dashboard analytique temps réel

FastAPI
React Native
PostgreSQL
TimescaleDB
Redis
Weaviate
Celery
AI Engineer (AI Agent for Real-Estate Qualification & Decisions) Middle
BAAM Construction Internship
Feb 2026 to Jun 2026 4 Months Marrakesh In office

Built and deployed a multi-tenant WhatsApp AI agent with a 6-phase state machine and Redis persistence to qualify leads for multiple real-estate projects. Reduced inference costs using a hybrid intent classifier and a PostgreSQL/pgvector RAG setup with embeddings and reranking. Implemented multilingual guardrails, LLM-as-judge evaluation with deterministic A/B testing, full observability, and GDPR-focused data handling.

Python
FastAPI
PostgreSQLsince 2026
pgvectorsince 2026
Redis
Embeddingssince 2026
Geminisince 2026
SQLAlchemy
Langfuse
Prometheus
Docker Compose
Pydantic
SonarQube
AI Engineer (Intelligent Multi-Document OCR) Middle
Exakis Nelite Internship
Jul 2025 to Oct 2025 3 Months Casablanca In office
Designed an end-to-end OCR architecture combining multimodal OCR and LLM-based extraction to handle multiple Moroccan administrative document types in several languages. Automated document classification with multiple domain prompts and produced structured JSON outputs validated with Pydantic, including regex fallbacks. Built a Streamlit UI with real-time KPIs and deployed the system using Docker Compose, Nginx, Azure App Service, and GitHub Actions CI/CD.
Python
Streamlit
Ollama
Groq
Llama
Pydanticsince 2025
SQLite
Plotly
Docker Composesince 2025
Nginx
Azure
GitHub Actions
Machine Learning Engineer (Dental 3D Classification) Middle
3D Smart Factory Internship
Jun 2024 to Sep 2024 3 Months Casablanca In office
Developed a 3D convolutional model to classify multiple dental categories from 3D meshes, including data augmentation, SMOTE, and k-fold cross-validation. Tracked experiments with MLflow and evaluated performance with reported accuracy and F1-score. Deployed the classifier as a REST API using FastAPI for production access.
Python
PyTorch
Open3D
FastAPIsince 2024
Scikit-learn
Computer Vision Engineer (Interactive 3D Reconstruction) Middle
HGS-Hightech Internship
Mar 2024 to May 2024 2 Months Remote/Hybrid
Implemented an interactive 3D reconstruction pipeline using 3D Gaussian Splatting from image sets, including efficient training for faster iteration. Integrated a 360° real-time viewer into the platform and exposed functionality through a REST API. Used traditional vision tooling and web rendering technologies to support fast user-facing visualization.
Pythonsince 2024
Gaussian Splatting
COLMAP
OpenCV
Three.JS
Rest API
Middle AI/ML Engineer Confidence: High LLM Engineer
A competent middle-level engineer focused on LLM-enabled systems and applied reinforcement learning. The strongest proven skill is building RAG/LLM stacks and API integrations, evidenced by a provider-factory pattern, FastAPI controllers and a working Qdrant provider implementation. Public code shows solid data engineering and prototyping but lacks structured experiment tracking, rigorous benchmarks, or production-grade monitoring and secrets hygiene.
Model Architecture & Training
4/10
How well models are designed and trained
Custom model code and training loops present (DQN implementations, agent replay logic, PPO/TF-Agents notebooks), but architectures are standard and lack rigorous ablation, distributed training or advanced optimization.
Data Pipeline & Feature Engineering
5/10
How data is prepared for models
Clear, pragmatic data engineering and feature handling - stratified splitting, Spark schemas, serialization utilities and TF-IDF + SMOTE preprocessing are implemented.
Experimentation & Evaluation
4/10
How results are measured and tested
Basic to moderate evaluation pipelines and visualization are present (metrics, confusion matrices, saved figures), but there is no evidence of structured experiment tracking (W&B/MLflow) or systematic ablation studies.
MLOps & Deployment
4/10
How models are shipped to production
Production-minded components exist - FastAPI routes, controllers, provider factories and a Qdrant provider implementation show real deployment and integration work for RAG systems, but full CI/CD, monitoring or drift detection are not evidenced.
Computational Efficiency
3/10
How efficiently computing resources are used
Some attention to efficiency - local quantized model loading (bitsandbytes 8-bit config) and Spark environment setup appear, but there is limited measured profiling, batching strategies, or quantization evaluation numbers.
Research Depth & Innovation
2/10
Depth of research and new ideas
Implements standard RL algorithms (Q-learning, SARSA, DQN) and applies RAG patterns, but there is no sign of novel algorithms, published benchmark reproduction or advanced research contributions.
Expertise
RAG• Middle
Conversational AI & Chatbots• Middle
Industries
Education• Middle
Health Care• Middle
Technologies
LLM Apps
SQL
PostgreSQL
Redis
Supabase
Weaviate
LangGraph
DVC
Rest API
LangChain
Claude
Grok
GCP
Docker Compose
pgvector
OpenCV
FAISS
Qdrant
TimescaleDB
Groq
Model Context Protocol
SQLAlchemy
YOLO
GitHub Actions
Vertex AI
Prometheus
Embeddings
Prompt Engineering
Multimodal AI
Function Calling
Diffusion Models
Computer Vision
AI Agents
NLP
Gradio
Langfuse
Ollama
Azure
spaCy
Llama
huggingface_hub
CI/CD
Transformers
NumPy
Git
SQLite
PyTorch
Docker
Kubernetes
CrewAI
Gemini
Nginx
LLM
RAG
Celery
NLTK
Ruff
structlog
Hallucination
CNN
Gaussian Splatting
Aider• mentioned only
NLP• mentioned only
Recommendations
  • Develop production RAG microservices and REST endpoints integrating Qdrant/FAISS with robust CI/CD, metrics and drift monitoring.
  • Own conversational AI features - build end-to-end chat stacks (retrieval, prompt templates, streaming responses) with test harnesses and cost/latency budgets.
  • Lead data engineering for healthcare ML pipelines - extend Spark streaming, schema validation, and privacy-preserving federated training with secure aggregation.
  • Add reproducible experiment tracking (W&B or MLflow), unit/integration tests for critical components, and documented evaluation scripts for benchmarks.
Repositories
The developer's experience in this domain has been verified based on AI analysis of the following repositories:
Middle Data Scientist Confidence: Medium Data Engineer
A practical data-engineering practitioner at a mid (middle) level who produces reliable preprocessing utilities and end-to-end experimental notebooks. The strongest proven skill is building data pipelines and preprocessing tooling, evidenced by the DataSplitter, schema helpers and serialization utilities in the cardiology project. What is not evidenced is rigorous statistical experimentation design, production-grade CI/CD for models, or sustained ownership of LLM/vector-db server infra (several components appear externally authored).
Statistical Rigor
2/10
Correct use of statistics
Basic evaluation metrics are present but statistical rigor is limited - no uncertainty quantification, no hypothesis testing or multiple-comparison controls, and notebooks show few written assumption checks.
Data Wrangling & Cleaning
6/10
Preparing and cleaning data
Clear, practical data-wrangling and pipeline utilities with schema definitions, stratified splits and robust serialization helpers - shows production-minded code for preprocessing and safe saving/loading.
Exploratory Analysis & Visualization
4/10
Exploring and visualizing data
Exploratory plots and visualizations are implemented (learning curves, confusion matrices, radar chart) but notebook narrative and written interpretation are sparse.
Predictive Modeling
4/10
Building models that predict
Implements non-trivial predictive artifacts (DQN agent, training loop, evaluation) and model saving/loading, but lacks robust baseline protocols, cross-validation schemes, calibration analysis and explicit leakage checks.
Business Insight & Impact
2/10
Turning analysis into business value
Some user-facing recommendations exist in the UI but there is little linking of results to business metrics, cost-of-error analysis or prioritized operational actions.
Reproducibility & Notebook Hygiene
3/10
Clean, repeatable analysis
Partial reproducibility: requirements files, seed constants and model save/load exist, plus Streamlit caching, but notebooks have low reasoning-to-code ratio and there is limited evidence of formal environment pinning, CI pipelines or data/versioning.
Expertise
Big Data• Middle
Streaming• Middle
Industries
Artificial Intelligence• Middle
Education• Middle
Health Care• Middle
Technologies
Deep Learning
Scikit-learn
Seaborn
Matplotlib
Plotly
Pandas
Streamlit
Aider• mentioned only
NLP• mentioned only
Recommendations
  • Develop production-ready streaming ETL pipelines (Spark Structured Streaming + Kafka) with monitoring, schema enforcement and end-to-end tests using the existing Spark environment code.
  • Harden reproducibility by adding pinned environment manifests (poetry/constraints), CI jobs that run core tests/notebooks, and dataset versioning (DVC or similar).
  • Extend model evaluation: add baselines, proper cross-validation or time-aware splits where required, calibration checks and uncertainty metrics to complement accuracy reports.
  • If building RAG/LLM services, consolidate ownership of the vector DB and API layers, add integration tests for provider factories and remove hard-coded secrets from notebooks.
Repositories
The developer's experience in this domain has been verified based on AI analysis of the following repositories:
Middle Backend Developer Confidence: Medium API Engineer
A pragmatic backend-focused engineer at a solid middle level who builds APIs and data-processing utilities end-to-end. The strongest proven skill is integrating LLM-driven workflows and a feedback loop (ArticleCrew + DataFlywheel in ennajari/article_generator) that connects generation, storage, and analytics. Public code shows reliable scripting, Spark glue, and file-based persistence, but lacks production-grade DB transactions, resilience patterns (retries/backoff), and formalized deployment/observability configurations.
API Design
4/10
How well APIs are designed
API surfaces show deliberate request/response models and input validation (FastAPI + pydantic); however there is no formal versioning, idempotency strategy, or documented error contract across services.
Data Layer & Database
3/10
Working with databases
Data layer is pragmatic and well-structured for file/JSON and CSV workflows with explicit save/load helpers and schema definitions, but lacks migrations, transactional boundaries, tuned queries, or DB schema evolution history.
Scalability & Performance
2/10
Handling load and speed
Some scalability signals (Spark environment setup, SPARK_CONFIG) and long timeouts in HTTP client code appear, but there are no robust caching/invalidation, queue-based decoupling, connection pooling, or measured performance optimizations.
System Architecture
3/10
Overall system structure
Modules are separated (crew, vectorstore, data flywheel, client/common), with clear responsibilities and a feedback loop design, but there is limited evidence of inter-service contracts, graceful degradation, or production service decomposition choices.
Security & Auth
4/10
Protecting data and access
Reasonable security basics are present - password hashing and token imports, input validation - but there are risky patterns (filesystem user store, lack of token lifecycle management, limited secrets handling).
Reliability & Observability
2/10
Stability and monitoring
Some observability and basic logging exist (module-level logging in Python utilities), but there is no consistent structured logging, correlation ids, retries/backoff, timeouts everywhere, or clear graceful-shutdown instrumentation.
Expertise
Backend AI & LLM• Middle
Python• Middle
Node.js• Middle
Databases & Vector Storage• Middle
Industries
Health Care• Middle
Technologies
Python• Middle
Node JS• Middle
MongoDB
Express
FastAPI
pySpark
Pydantic
Uvicorn
Requests
Axios
Docker• mentioned only
Recommendations
  • Develop REST/async APIs that wrap LLM workflows and data pipelines - implementing idempotency, request timeouts, and structured logging.
  • Build data preprocessing and ingestion components for ML pipelines (Spark jobs, schema enforcement, stratified splits) and unit/integration tests for them.
  • Integrate a production datastore and migration strategy (replace filesystem JSON with a DB, add migration history and transactional update paths).
  • Implement reliability features for LLM orchestration: request retries with jitter, circuit-breaker guards, and graceful shutdown for model-loading endpoints.
Repositories
The developer's experience in this domain has been verified based on AI analysis of the following repositories: