{"id":1227674,"url":"https://alion.io/job/consulting-pandits-ml-system-engineer","title":"ML System Engineer","company":{"id":3800739,"name":"Consulting Pandits","domain":"consultingpandits.com","url":"https://alion.io/company/consulting-pandits","size_band":null,"is_staffing_agency":true,"employer_type":"agency","is_intermediary":false,"listed_via":null,"ats_vendor":null,"truth_index":null},"role":"DevOps","role_family":"DevOps","seniority":"middle","employment_type":null,"work_mode":"on_site","remote_scope":null,"remote_scope_basis":null,"remote_working_hours":null,"hiring_geo_confidence":"structured","locations":["Pune, India"],"countries":["IN"],"hiring_countries":[],"hiring_countries_total":0,"salary":null,"salary_estimate":{"min_usd":11000,"max_usd":29000,"period":"year","method":"role_seniority_country_remote_cell","sample_n":15},"experience_years_min":4,"visa_sponsorship":false,"relocation_package":false,"has_equity":false,"technologies":[{"name":"Accelerate","optional":false},{"name":"Agentic Workflows","optional":false},{"name":"AI Agents","optional":false},{"name":"Anthropic","optional":false},{"name":"AutoGen","optional":false},{"name":"AWQ","optional":false},{"name":"CrewAI","optional":false},{"name":"Docker","optional":false},{"name":"Edge AI","optional":false},{"name":"Function Calling","optional":false},{"name":"GGUF","optional":false},{"name":"GPTQ","optional":false},{"name":"Hugging Face","optional":false},{"name":"KV Cache","optional":false},{"name":"LangGraph","optional":false},{"name":"Linux","optional":false},{"name":"LlamaIndex","optional":false},{"name":"LLM","optional":false},{"name":"Mistral","optional":false},{"name":"Multi-Agent Systems","optional":false},{"name":"Multimodal AI","optional":false},{"name":"OpenAI","optional":false},{"name":"PEFT","optional":false},{"name":"Prompt Engineering","optional":false},{"name":"Python","optional":false},{"name":"PyTorch","optional":false},{"name":"PyTorch C++","optional":false},{"name":"Quantization","optional":false},{"name":"RAG","optional":false},{"name":"SGLang","optional":false},{"name":"Structured Outputs","optional":false},{"name":"TensorRT-LLM","optional":false},{"name":"TGI","optional":false},{"name":"Tool Use","optional":false},{"name":"Transformers","optional":false},{"name":"vLLM","optional":false},{"name":"Arize Phoenix","optional":true},{"name":"C++","optional":true},{"name":"CUDA","optional":true},{"name":"CUDA Toolkit","optional":true},{"name":"Fine-tuning","optional":true},{"name":"Grafana","optional":true},{"name":"Human-in-the-Loop","optional":true},{"name":"Kubernetes","optional":true},{"name":"LangChain","optional":true},{"name":"LangSmith","optional":true},{"name":"llama.cpp","optional":true},{"name":"LoRA","optional":true},{"name":"Prometheus","optional":true},{"name":"QLoRA","optional":true},{"name":"Rust","optional":true},{"name":"TensorFlow","optional":true},{"name":"TensorFlow C++","optional":true},{"name":"TensorRT","optional":true},{"name":"Triton","optional":true},{"name":"Weights & Biases","optional":true}],"status":"live","first_seen_at":"2026-09-23T10:59:18Z","employer_posted_date":null,"last_verified_at":"2026-09-23T10:59:18Z","board_verified":false,"closed_at":null,"days_open":5,"trust":{"level":"not_scored","repost_count":null,"flags":[],"days_open":5},"description":"About the Role : \n\nWe are hiring an ML Systems Engineer to design and deliver cutting-edge AI solutions for enterprise clients at the frontier of agentic AI, inference engineering, and ML systems architecture. You will go beyond applied ML - dissecting how AI systems are built, optimized, and scaled - designing production-grade architectures spanning retrieval systems, inference pipelines, and agentic workflows. You will translate state-of-the-art capabilities into robust, performant solutions, operating at the intersection of ML research awareness and engineering discipline.\n\nKey Responsibilities : \n\n- Design and deliver production-grade AI systems for enterprise clients spanning agentic workflows, LLM inference pipelines, and retrieval-augmented architectures.\n\n- Lead ML systems architecture decisions - model serving topology, inference backend selection, KV cache management, batching strategies, and memory optimization - alongside ML performance engineering to profile bottlenecks, benchmark throughput/latency, and evaluate quantization strategies (GPTQ, AWQ, GGUF).\n\n- Architect RAG pipelines and agentic AI systems - from chunking, embedding, hybrid retrieval, and re-ranking through to multi-agent orchestration, tool use, and memory architectures.\n\n- Evaluate frontier model capabilities - reasoning models, multimodal systems, fine-tuned variants - and make principled architectural trade-off decisions for client contexts.\n\n- Build reusable accelerators, reference implementations, and evaluation/observability frameworks encoding best practices across engagements.\n\n- Contribute to technical solutioning - architecture designs, proof-of-concepts, and feasibility assessments - in client-facing contexts.\n\nTechnical Qualifications : \n\n- Python & ML ecosystem : Strong programming skills with production AI system experience; hands-on with the PyTorch ecosystem including Hugging Face Transformers, PEFT, Accelerate, and Datasets.\n\n- LLM inference & serving : Deep knowledge of KV cache mechanics, quantization, and batching; hands-on with at least one inference runtime (vLLM, TGI, TensorRT-LLM, SGLang, or similar).\n\n- Hands-on experience supporting AI/ML and LLM inference platforms at scale, including working with vLLM for high-performance LLM serving, optimization, and large-scale inference.\n\n- RAG & Agentic Systems : Experience designing retrieval architectures and building agentic systems using LangGraph, LlamaIndex Workflows, AutoGen, or CrewAI - including tool use, memory, and multi-agent coordination.\n\n- LLM APIs & prompt engineering : Strong grasp of structured output generation, function calling, and provider SDK usage across OpenAI, Anthropic, Mistral, Hugging Face, and similar.\n\n- Deployment fundamentals : Proficiency with Docker, containerization, and Linux environments for packaging, deploying, and debugging AI systems.\n\n- Comfortable leveraging AI-assisted tools for collaborative development, code generation, refactoring, and productivity enhancement.\n\nPreferred Qualifications : \n\n- Fine-tuning : Experience with LoRA/QLoRA, dataset curation, and instruction tuning; understanding of when fine-tuning is the right lever vs. prompting or RAG.\n\n- Low-level AI systems : Familiarity with CUDA, Triton, or similar GPU programming models; working knowledge of C++ or Rust.\n\n- Infrastructure & observability : Kubernetes for containerized AI workloads; experience with LangSmith, Arize, W&B, Phoenix, or Prometheus/Grafana for ML observability.\n\nWays to Stand Out : \n\n- You have built and deployed a production agentic system and can speak to the failure modes and design decisions that only emerge at runtime.\n\n- You have done inference optimization at a systems level - tuning serving infrastructure, implementing custom batching logic, or optimizing a quantization pipeline to hit real SLAs.\n\n- You have open-source contributions to prominent ML systems repositories - vLLM, SGLang, llama.cpp, TGI, LangChain, LlamaIndex, or similar - demonstrating work that holds up to community scrutiny.\n\n- You have designed custom LLM evaluation frameworks with structured regression harnesses, domain-specific evals, or human-in-the-loop feedback loops - beyond off-the-shelf metrics.\n\n- You bring a client-facing engineering mindset and can defend opinions on reasoning models, long-context retrieval, or inference hardware tradeoffs based on hands-on experience.\n\nSkills\nMachine Learning, Agentic AI, Artificial Intelligence, LLM, Python, Tensorflow","description_format":"text","description_chars":4493,"description_truncated":false,"requirements":{"experience_years_min":4,"management_years_min":null,"team_size_min":null,"manages_managers":false,"education":null,"security_clearance":false,"languages":[]},"benefits":[],"hiring_locations":[],"hiring_excludes":[],"relocation_offered":false,"industries":[],"lifecycle":[{"event":"open","at":"2026-09-25T13:06:44Z"}],"liveness":{"score":50,"band":"ok","label":"Likely open","p_open":1,"p_active":0.505,"p_room":1,"age_days":4,"expected_fill_days":16,"reasons":["seen:4","agency","velocity","win:early"],"computed_at":"2026-09-28T05:45:00Z"},"pay":null,"html_url":"https://alion.io/job/consulting-pandits-ml-system-engineer","json_url":"https://alion.io/job/consulting-pandits-ml-system-engineer.json","meta":{"generated_at":"2026-09-29T03:13:36Z","cache_seconds":300,"methodology":"https://alion.io/methodology","terms":"https://alion.io/terms","contact":"https://alion.io/contact","api":"https://alion.io/developers","usage":{"tier":"crawler","counted_by":"address","units_charged":1,"used_today":2951,"day_limit":5000,"remaining_today":2049,"minute_limit":60,"resets_at":"2026-09-30T00:00:00Z"}}}