AI Engineer
5+ years exp
Python
Java
MLOps & Deployment: 5/10
Data Pipeline & Feature Engineering: 4/10
Active 11 days ago
Invite to interview
Download CVCV
Overview
Technical skills
Timeline
Roles
Overview
LLM-focused backend engineer (senior) experienced building agentic multi-node pipelines for production voice-first workflows; strongest at integrating hosted LLMs, observability, and external API fallbacks. The most proven skill is architecting and operating an async LangGraph-based multi-agent orchestration with concrete artifacts in agent_graph.py and a comprehensive e2e QA runner in e2e_test.py. There is little or no evidence of custom model training, research-grade experiments, or GPU/quantization engineering in public code.
Phone
Technical skills
Languages
2
Python
Java
Python
6
Pydantic
SQLAlchemy
Asyncio
HTTPX
Alembic
Uvicorn
AI/ML
24
LLM
Gemini
LangChain
RAG
Model Context Protocol
Prompt Engineering
LangSmith
NLP
Text-to-Speech
Amazon SageMaker
AWS Bedrock
Claude
CrewAI
LangGraph
Llama
OpenCV
PyTorch
TensorFlow
Transformers
Agentic Workflows
Multi-Agent Systems
Tool Use
PEFT
Qwen
DevOps
4
GCP
Google Cloud Run
AWS
Kubernetes
Other
27
PostgreSQL
SQLite
Anthropic
Hugging Face
Ollama
OpenAI
GitHub
Containers
WebSockets
Grafana
Databases
Apache Kafka
Weights & Biases
Context Engineering
Docker
Sentry
Groq
Prometheus
Redpanda
AI Agents
Loki
Deepgram
Fine-tuning
Google ADK
Multimodal AI
OCR
Sentiment Analysis
Timeline
AI Engineer
•
Middle
Cyble
•
Full-Time
Built LLM-driven alert processing flows combining deterministic tagging with LLM reasoning to surface relevant events and reduce operational noise. Worked on Kafka/Redpanda queue monitoring, including consumer lag and queue health, and supported reliability improvements in production. Managed Kubernetes deployments by tuning CPU/memory, scaling replicas, and debugging distributed issues using application logs and traffic visibility. Implemented multimodal vision workflows for brand/logo misuse detection with OCR and automated image preprocessing, improving observability and debugging speed.
Apache Kafka
Redpanda
Kubernetes
Prometheus
Grafana
Loki
Sentry
OCR
Model Context Protocol
Senior LLM Engineer
•
Senior
Infocusp Innovations
•
Full-Time
Developed agentic and multi-agent workflows focused on improving reasoning quality, output fidelity, and response latency. Implemented RAG and context engineering approaches to enhance accuracy while maintaining reliability in distributed environments. Built supporting components for code sandboxing, LLM evaluation, and observability, and incorporated MCP and tool-use patterns in production-oriented research.
Agentic Workflows
Multi-Agent Systems
RAG
Context Engineering
Model Context Protocol
Tool Use
Data Scientist
•
Middle
SquadStack
•
Full-Time
Worked on LLM and agent quality evaluation, using quality and business metrics to identify improvement areas and reduce manual coordination effort. Built solutions involving RAG, prompt engineering, and NLP for knowledge and generation tasks. Contributed to voice-AI workflows including speech-to-text and text-to-speech components, supporting downstream NLP-driven applications.
RAG
Prompt Engineering
NLP
Text-to-Speech
LLM
Senior AI/ML Engineer
Confidence: Medium LLM Engineer
LLM-focused backend engineer (senior) experienced building agentic multi-node pipelines for production voice-first workflows; strongest at integrating hosted LLMs, observability, and external API fallbacks. The most proven skill is architecting and operating an async LangGraph-based multi-agent orchestration with concrete artifacts in agent_graph.py and a comprehensive e2e QA runner in e2e_test.py. There is little or no evidence of custom model training, research-grade experiments, or GPU/quantization engineering in public code.
Model Architecture & Training
2/10
How well models are designed and trained
Limited model engineering evidence; the code integrates hosted LLMs (Gemini) with structured output use but contains no model training, custom layers, or tuning pipelines.
Evidence
welfareflow-india/agent_graph.py: _generate_llm_reasoning — async Gemini Flash invocation with low-temperature settings
welfareflow-india/agent_graph.py: voice_intent_agent_node — llm_structured = _llm.with_structured_output(ExtractedCitizenProfile) and use of .ainvoke()
Data Pipeline & Feature Engineering
4/10
How data is prepared for models
Solid data-prep and domain-specific feature handling for names and documents, including custom phonetic normalisation and token-aligned similarity scoring.
Evidence
welfareflow-india/agent_graph.py: preprocess_indian_name — transliteration vowel/consonant normalisation and honorific stripping
welfareflow-india/agent_graph.py: _best_token_alignment_score and compute_name_match_score — token-aware Jaro-Winkler alignment used in document_audit_node
welfareflow-india/document_audit_node (agent_graph.py) — end-to-end OCR field mapping and cross-document name comparison
Experimentation & Evaluation
3/10
How results are measured and tested
Practical QA and reproducible acceptance testing for the pipeline exist, but no ML experiment tracking, ablation studies, or formal model evaluation harness is present.
Evidence
welfareflow-india/e2e_test.py: run_scenario / write_html_report — comprehensive end-to-end acceptance test scenarios and QA report
welfareflow-india/main.py: LangSmith integration and run_config metadata for pipeline runs (traceability at runtime)
MLOps & Deployment
5/10
How models are shipped to production
Strong production-oriented engineering: async FastAPI service patterns, background pipeline orchestration, observability/tracing hooks, DB session scoping, and multi-tier external API fallbacks.
Evidence
welfareflow-india/main.py: _lifespan, _run_pipeline, background tasks for SLA watchdog, and endpoint lifecycle handling
welfareflow-india/observability.py and agent_graph.py: LangSmith tracing (traceable decorators) and startup health checks
welfareflow-india/uipath_maestro.py: acquire_token token cache and tiered submission fallbacks
Computational Efficiency
2/10
How efficiently computing resources are used
Some asynchronous and efficiency-aware patterns (parallel OCR, retry/backoff, token caching) are used but there is no GPU/quantization/distributed training or low-level performance optimization.
Evidence
welfareflow-india/agent_graph.py: document_audit_node — parallel OCR via asyncio.create_task and asyncio.gather
welfareflow-india/agent_graph.py: _retry_async — exponential backoff wrapper used for external calls
welfareflow-india/uipath_maestro.py: in-memory token caching to avoid repeated auth round-trips
Research Depth & Innovation
3/10
Depth of research and new ideas
Some domain-driven algorithmic work (custom Jaro-Winkler and token-alignment strategy) shows thoughtful problem solving, but no novel research contributions or reproduced paper implementations are present.
Evidence
welfareflow-india/agent_graph.py: compute_jaro_winkler and _jaro_similarity — full Python implementation
welfareflow-india/agent_graph.py: _best_token_alignment_score — conservative assignment algorithm addressing token truncation and reordering
Expertise
AI Agents & Agentic Workflows• Senior
Industries
Government• Senior
Technologies
Databases
LangGraph
Weights & Biases
LangChain
Claude
Qwen
OpenCV
Groq
Model Context Protocol• since 2024
Loki• since 2026
Prometheus• since 2026
Fine-tuning
Prompt Engineering• since 2023
Multimodal AI
AI Agents
NLP• since 2023
LangSmith
Ollama
PEFT
Llama
AWS Bedrock
Transformers
TensorFlow
PyTorch
AWS
Docker
Kubernetes• since 2026
CrewAI
Gemini
Grafana• since 2026
LLM• since 2023
RAG• since 2023
Google ADK
Redpanda• since 2026
Sentiment Analysis
OpenAI
Anthropic
Hugging Face
Amazon SageMaker
GitHub
OCR• since 2026
Text-to-Speech• since 2023
Deepgram
Context Engineering• since 2025
Agentic Workflows• since 2024
Multi-Agent Systems• since 2024
Tool Use• since 2024
Deep Learning• mentioned only
LangGraph• mentioned only
Recommendations
- Develop production-grade multi-agent LLM systems that require robust orchestration, tracing, and sandbox fallbacks (voice + OCR + RPA integrations).
- Implement privacy-sensitive, compliance-first citizen-facing services that integrate external APIs and RPA (human-in-the-loop approval, consent revocation, masked PII vault).
- Build integration layers and reliability logic for third-party services (retry/backoff, tiered fallbacks, token caching, health probes) and lead backend service design for LLM orchestration.
Repositories
The developer's experience in this domain has been verified based on AI analysis of the following repositories:
Senior Data Scientist
Confidence: Medium Data Engineer
Senior-level engineer specializing in production-grade async multi-agent pipelines for government welfare automation. The strongest proven skill is pipeline orchestration and integration, as evidenced by the LangGraph-based agent_graph.py together with the multi-tier UiPath integration in uipath_maestro.py. What is not evidenced is any statistical modeling or model-training work, formal hypothesis testing, or extensive unit-test coverage beyond the end-to-end QA runner.
Statistical Rigor
1/10
Correct use of statistics
Minimal statistical rigor evidenced; there are QA checks and scenario pass/fail assertions but no hypothesis testing, uncertainty quantification, or experimental design.
Evidence
welfareflow-india/e2e_test.py: scenario definitions and PASS/FAIL verdicts used for QA rather than statistical analysis
Data Wrangling & Cleaning
8/10
Preparing and cleaning data
Strong data wrangling and cleaning work for identity verification: a bespoke Indian name normaliser, a token-aware Jaro-Winkler alignment implementation, robust OCR integration and sandbox/mock fallbacks.
Evidence
welfareflow-india/agent_graph.py: preprocess_indian_name, compute_jaro_winkler, compute_name_match_score, _best_token_alignment_score
welfareflow-india/agent_graph.py: _call_sarvam_vision with mock OCR fallbacks and parallel OCR task orchestration
Exploratory Analysis & Visualization
4/10
Exploring and visualizing data
Exploratory outputs and human-oriented reporting are present but not research-grade EDA; there is a readable QA HTML report and a simple impact dashboard linking pipeline outcomes to business metrics.
Evidence
welfareflow-india/e2e_test.py: write_html_report generates a self-contained QA acceptance report
welfareflow-india/main.py: impact_dashboard aggregates case metrics and computes unlocked benefits
Predictive Modeling
1/10
Building models that predict
Uses LLMs for structured extraction and plain-language reasoning but shows no model development, feature engineering or predictive model evaluation; the ML work is orchestration and inference-only.
Evidence
welfareflow-india/agent_graph.py: _llm with Gemini Flash and llm_structured.ainvoke for ExtractedCitizenProfile schema
Business Insight & Impact
6/10
Turning analysis into business value
Clear product and policy-aware framing for a government welfare use case, with DPDP consent handling, human-in-the-loop gating, and impact-oriented metrics; analysis is tied to operational business outcomes rather than abstract modeling.
Evidence
welfareflow-india/main.py: initialize_case includes consent, OTP checks and DPDP-aware revoke_consent logic
welfareflow-india/agent_graph.py: eligibility_router_node implements PM-Kisan and Ayushman Bharat business rules and emits citizen-facing messages
Reproducibility & Notebook Hygiene
5/10
Clean, repeatable analysis
Reproducibility practices are moderate: pinned requirements.txt, sandbox modes and a runnable end-to-end QA runner; there is no explicit data versioning or CI shown but the system is designed to run in zero-infra sandbox mode.
Evidence
welfareflow-india/requirements.txt: pinned dependencies for reproducible environment
welfareflow-india/e2e_test.py: sandbox-mode scenarios and a runnable QA acceptance test suite
Expertise
Data Science• Middle
Streaming• Middle
Industries
Government• Senior
Farming & Agriculture• Middle
Financial Services• Middle
Health Care• Middle
Technologies
PostgreSQL
SQLite
Asyncio
HTTPX
Uvicorn
Alembic
Deep Learning• mentioned only
LangGraph• mentioned only
Recommendations
- Develop real-time, event-driven backend services that integrate LLMs with external APIs and RPA systems, leveraging the existing LangGraph and UiPath integration patterns.
- Build identity verification and name-normalisation modules for government or financial services, extending the token-aware Jaro-Winkler approach in agent_graph.py.
- Implement human-in-the-loop workflows and compliance features for regulated domains, including consent management and auditable traceability.
- Create robust end-to-end QA harnesses and sandboxed developer workflows, expanding the e2e_test.py patterns into CI jobs and reproducible test fixtures.
Repositories
The developer's experience in this domain has been verified based on AI analysis of the following repositories:
Senior DevOps Engineer
Confidence: Medium Generalist
Senior-level LLM and backend engineer focused on real-time audio integrations and recommendation engines with solid concurrency and test practices. The strongest proven skill is building a production-grade real-time Twilio-to-Gemini voice bridge with custom mulaw resampling and session management (voicebot-resume/app/server.py). Public code does not show infrastructure-as-code, CI/CD pipelines, cloud cost engineering, or documented runbooks and incident postmortems.
CI/CD Pipelines
Automated build and deploy
Not evidenced in public code
Infrastructure as Code
Managing servers with code
Not evidenced in public code
Containerization & Orchestration
2/10
Working with containers
Basic containerization evidence (single Dockerfile) is present but lacks multi-stage builds, non-root user, resource rationale, or orchestration manifests.
Evidence
voicebot-resume/Dockerfile: simple single-stage Python image and uvicorn run command
Observability & Monitoring
2/10
Watching system health
Lightweight observability is present via structured debug logging and NDJSON debug logging, but no SLOs, alerting rules, or dashboards as code were found.
Evidence
voicebot-resume/dev/poc.py:debug_log - NDJSON debug logger writing hypothesis entries
voicebot-resume/app/server.py:log - consistent operational log messages across async tasks
Reliability & Incident Response
4/10
Keeping systems up
Good reliability practices in application code: async task separation, background tool execution to avoid blocking, silence monitors, blip-handling and graceful shutdowns are implemented, but there are no runbooks, incident postmortems, or deployment rollback strategies in code.
Evidence
voicebot-resume/app/server.py:media_stream - offloaded tool execution (execute_and_respond) and monitor_silence with explicit thresholds and session cancellation
voicebot-resume/dev/poc.py:receive_audio and monitor_silence - blip counters, fatal error detection, and _handle_sigint for graceful shutdown
Cloud & Cost Optimization
1/10
Smart use of the cloud
Cloud usage appears in application-level code (Google APIs and a Cloud Run comment), but there is no evidence of autoscaling strategy, IAM least-privilege design, spot/eviction handling, or cost-measurement artifacts.
Evidence
voicebot-resume/app/server.py:book_appointment - uses google.auth.default and googleapiclient (Calendar API)
voicebot-resume/Dockerfile: comment referencing Cloud Run default port 8080
Expertise
Observability & Monitoring• Middle
Site Reliability Engineering• Middle
Industries
Artificial Intelligence• Middle
Technologies
Containers
Python• since 2021 • Senior
GCP
SQLAlchemy
WebSockets
Pydantic
Google Cloud Run
Recommendations
- Develop real-time conversational voice agents and telephony integrations that require low-latency audio processing and robust session management.
- Build recommendation services or ranking systems where the candidate can extend the existing weighted scoring engine and productionize it with database schemas and tests.
- Harden production deployments by adding IaC (Terraform/CDK), CI/CD pipelines (parameterized reusable workflows), and secrets/workload-identity patterns for Google APIs.
- Create observability runbooks: define SLOs, alerting rules, dashboards as code, and incident playbooks tied to the existing silence/connection error handling.
Repositories
The developer's experience in this domain has been verified based on AI analysis of the following repositories:
