Overview
Technical skills
Timeline
Audited an agent-memory architecture for LLM agents in production, including filesystem-backed memory, context compaction, and cross-session long-term retrieval using PostgreSQL with pgvector and hybrid search. Analyzed inference cost leaks and implemented prompt-caching stability across multi-provider routing. Designed an agentic loop with parallel tool dispatch, streaming via SSE, and retrieval patterns including typed relations.
Développé et déployé 22 modules métier : POS, carte NFC, WhatsApp bot darija, OCR de factures via GPT-4o Vision, dashboard analytique temps réel
Built and deployed a multi-tenant WhatsApp AI agent with a 6-phase state machine and Redis persistence to qualify leads for multiple real-estate projects. Reduced inference costs using a hybrid intent classifier and a PostgreSQL/pgvector RAG setup with embeddings and reranking. Implemented multilingual guardrails, LLM-as-judge evaluation with deterministic A/B testing, full observability, and GDPR-focused data handling.
SQL
PostgreSQL
Redis
Supabase
Weaviate
LangGraph
DVC
Rest API
LangChain
Claude
Grok
GCP
Docker Compose
pgvector
OpenCV
FAISS
Qdrant
TimescaleDB
Groq
Model Context Protocol
SQLAlchemy
YOLO
GitHub Actions
Vertex AI
Prometheus
Embeddings
Prompt Engineering
Multimodal AI
Function Calling
Diffusion Models
Computer Vision
AI Agents
NLP
Gradio
Langfuse
Ollama
Azure
spaCy
Llama
huggingface_hub
CI/CD
Transformers
NumPy
Git
SQLite
PyTorch
Docker
Kubernetes
CrewAI
Gemini
Nginx
LLM
RAG
Celery
NLTK
Ruff
structlog
Hallucination
CNN
Gaussian Splatting
- Develop production RAG microservices and REST endpoints integrating Qdrant/FAISS with robust CI/CD, metrics and drift monitoring.
- Own conversational AI features - build end-to-end chat stacks (retrieval, prompt templates, streaming responses) with test harnesses and cost/latency budgets.
- Lead data engineering for healthcare ML pipelines - extend Spark streaming, schema validation, and privacy-preserving federated training with secure aggregation.
- Add reproducible experiment tracking (W&B or MLflow), unit/integration tests for critical components, and documented evaluation scripts for benchmarks.
Scikit-learn
Seaborn
Matplotlib
Pandas
Streamlit
- Develop production-ready streaming ETL pipelines (Spark Structured Streaming + Kafka) with monitoring, schema enforcement and end-to-end tests using the existing Spark environment code.
- Harden reproducibility by adding pinned environment manifests (poetry/constraints), CI jobs that run core tests/notebooks, and dataset versioning (DVC or similar).
- Extend model evaluation: add baselines, proper cross-validation or time-aware splits where required, calibration checks and uncertainty metrics to complement accuracy reports.
- If building RAG/LLM services, consolidate ownership of the vector DB and API layers, add integration tests for provider factories and remove hard-coded secrets from notebooks.
Python• Middle
Node JS• Middle
MongoDB
Express
FastAPI
pySpark
Pydantic
Uvicorn
Requests
Axios
- Develop REST/async APIs that wrap LLM workflows and data pipelines - implementing idempotency, request timeouts, and structured logging.
- Build data preprocessing and ingestion components for ML pipelines (Spark jobs, schema enforcement, stratified splits) and unit/integration tests for them.
- Integrate a production datastore and migration strategy (replace filesystem JSON with a DB, add migration history and transactional update paths).
- Implement reliability features for LLM orchestration: request retries with jitter, circuit-breaker guards, and graceful shutdown for model-loading endpoints.
