Overview
Technical skills
Timeline
Developed an ML system to personalize digital credit offers using demographic, transactional, and product-level features. Built a synthetic-data pipeline for more than 100,000 customer personas with realistic demographic, spending, and credit features. Built clustering and PyTorch autoencoder prototypes for more than five customer archetypes with a six-person engineering team.
Built a leakage-aware Python platform for ETF ingestion, portfolio-risk metrics, GMM/HMM/KMeans regime models, stress tests, and cost-aware backtests. Prevented look-ahead bias with chronological splits, train-only scaling, shifted signals, and future-mutation tests; shipped 322 tests and a nine-page dashboard.
Built a Rust matching engine with price-time priority, partial fills, deterministic replay, portfolio P&L, pre-trade risk limits, and a kill switch; CI passes 247 tests. Benchmarked the 10,000-event core path at 125 ns p50 and approximately 4.6 million events per second on an Apple M4 Pro versus approximately 1.2 million events per second for a naive Python baseline.
Built a full-stack platform for versioned datasets, prompts, and models; deterministic and LLM-based graders; evaluation dashboards; and CI quality gates. Caught a controlled 20-case RAG regression: pass rate fell from 95% to 85%, failures rose from 1 to 3, and estimated cost more than doubled. Shipped CI with 221 backend tests, a frontend build, and an eval gate enforcing pass rate, score, cost, and p95 latency.
Gunicorn
- Design and implement deterministic evaluation pipelines and CI quality gates for model and prompt changes, including runner, graders, and provable metrics.
- Build backend APIs and integrations for LLM-backed workflows that require typed contracts, fallback behavior, and robust persisted failure modes.
- Implement schema evolution and migration-aware data models for analytics and evaluation workloads, including Alembic migrations and DB-backed test fixtures.
SQL
C++
MATLAB
Rest API
Terraform
GCP
XGBoost
GitHub Actions
Vercel
Scikit-learn
Google GenAI SDK
CI/CD
Pandas
NumPy
Django
Git
PyTorch
AWS
Docker
LLM
RAG
PyTorch C++
Ruff
Streamlit
Pydantic
Google Cloud Run
Data Augmentation
- Design and implement EvalOps pipelines, graders, and CI eval gates for LLM systems (quality gates, mock profiles, deterministic graders).
- Build production-ready LLM inference layers and provider adapters with robust error handling and cost/latency accounting.
- Develop backend services for LLM-powered applications with strong test coverage, schema validation, and migration/seed tooling.
- Implement deterministic dataset ingestion, seed management, and reproducible evaluation artifacts for model comparison and regression testing.
Python• Senior
PostgreSQL
SQLAlchemy
FastAPI
OpenAI SDK
Gemini
HTTPX
Alembic
- Develop evaluation pipelines and CI gates that run deterministic and mock-backed LLM tests (implement and maintain the eval-gate and mock providers).
- Build and maintain ETL/data ingestion tooling with strict validation and seed-data workflows (extend the JSONL importer and seed loader pipelines).
- Design backend services for LLM evaluation, grading, and analytics with durable DB schemas and migration strategies (work on providers, graders, and analytics endpoints).
