{"id":2086664,"url":"https://alion.io/job/smartncode-aigenai-engineer-llm-integration","title":"AI/GenAI Engineer - LLM Integration","company":{"id":3912697,"name":"SMARTnCODE","domain":"smartncode.com","url":"https://alion.io/company/smartncode","size_band":null,"is_staffing_agency":false,"employer_type":"direct","is_intermediary":false,"listed_via":null,"ats_vendor":null,"truth_index":null},"role":"AI/ML","role_family":"AI/ML","seniority":"junior","employment_type":null,"work_mode":"on_site","remote_scope":null,"remote_scope_basis":null,"remote_working_hours":null,"hiring_geo_confidence":"structured","locations":["Hyderabad, India"],"countries":["IN"],"hiring_countries":[],"hiring_countries_total":0,"salary":null,"salary_estimate":{"min_usd":12000,"max_usd":29000,"period":"year","method":"role_seniority_country_remote_cell","sample_n":10},"experience_years_min":2,"visa_sponsorship":false,"relocation_package":false,"has_equity":false,"technologies":[{"name":"Amazon SageMaker","optional":false},{"name":"Anthropic","optional":false},{"name":"AWS","optional":false},{"name":"Azure","optional":false},{"name":"Chain-of-Thought","optional":false},{"name":"Claude","optional":false},{"name":"Falcon","optional":false},{"name":"FastAPI","optional":false},{"name":"Fine-tuning","optional":false},{"name":"Function Calling","optional":false},{"name":"Gemini","optional":false},{"name":"Hugging Face","optional":false},{"name":"LangChain","optional":false},{"name":"Llama","optional":false},{"name":"LlamaIndex","optional":false},{"name":"LLM","optional":false},{"name":"LLM Guardrails","optional":false},{"name":"LoRA","optional":false},{"name":"Mistral","optional":false},{"name":"MLFlow","optional":false},{"name":"Multimodal AI","optional":false},{"name":"OpenAI","optional":false},{"name":"PEFT","optional":false},{"name":"Pinecone","optional":false},{"name":"Prompt Engineering","optional":false},{"name":"Python","optional":false},{"name":"Qdrant","optional":false},{"name":"QLoRA","optional":false},{"name":"Quantization","optional":false},{"name":"RAG","optional":false},{"name":"Semantic Search","optional":false},{"name":"Semantic Search","optional":false},{"name":"TGI","optional":false},{"name":"Tool Use","optional":false},{"name":"Transformers","optional":false},{"name":"Vertex AI","optional":false},{"name":"vLLM","optional":false},{"name":"Weaviate","optional":false},{"name":"Weights & Biases","optional":false},{"name":"Cohere SDK","optional":true},{"name":"Docker","optional":true},{"name":"GCP","optional":true},{"name":"JavaScript","optional":true},{"name":"PostgreSQL","optional":true},{"name":"Redis","optional":true},{"name":"TypeScript","optional":true}],"status":"live","first_seen_at":"2026-10-08T12:20:10Z","employer_posted_date":null,"last_verified_at":"2026-10-08T12:20:10Z","board_verified":false,"closed_at":null,"days_open":2,"trust":{"level":"not_scored","repost_count":null,"flags":[],"days_open":2},"description":"Position : AI/GenAI Engineer (LLM Integration Specialist)\n\nExperience : 2 - 4 years\n\nCTC : 12 to 16 LPA\n\nLocation : Hyderabad\n\nAbout the Role : \n\nWe're building a high-performance chat application and looking for an AI/GenAI Engineer to lead the integration and optimization of Large Language Models (LLMs). You'll be responsible for connecting LLM APIs, implementing domain-specific fine-tuning strategies, prompt engineering, and ensuring optimal performance for production use.\n\nKey Responsibilities : \n\nLLM Integration & Architecture : \n\n- Integrate multiple LLM APIs (OpenAI, Anthropic Claude, Google Gemini, or open-source models)\n\n- Design and implement robust API wrapper services with retry logic, fallback mechanisms, and error handling\n\n- Implement streaming responses for real-time chat experience\n\n- Build rate limiting and quota management systems\n\n- Handle token counting, context window management, and cost optimization\n\nDomain Customization & Fine-tuning : \n\n- Develop domain-specific prompt engineering strategies\n\n- Implement RAG (Retrieval Augmented Generation) pipelines using vector databases\n\n- Fine-tune or adapt models for specific use cases using techniques like LoRA, prompt tuning\n\n- Create and maintain knowledge bases for domain-specific responses\n\nPerformance & Optimization : \n\n- Optimize API response times and reduce latency\n\n- Implement caching strategies for common queries\n\n- Monitor and optimize token usage to control costs\n\nSafety & Quality : \n\n- Implement content moderation and safety filters\n\n- Build guardrails to prevent prompt injection and jailbreaking\n\n- Develop evaluation frameworks to measure response quality\n\nTechnical Stack : \n\n- Languages : Python (primary), JavaScript/TypeScript (basic understanding)\n\n- LLM APIs : OpenAI, Anthropic, Google Gemini, Cohere\n\n- Frameworks : LangChain, LlamaIndex, FastAPI\n\n- Vector DBs : Pinecone, Weaviate, Qdrant, or ChromaDB\n\n- Infrastructure : Docker, Redis, PostgreSQL, Message Queues\n\n- Cloud : AWS/GCP/Azure\n\nInfrastructure :\n\n- Design scalable architecture for handling concurrent LLM requests\n\n- Implement queue systems for managing high-volume API calls\n\n- Set up monitoring and logging for LLM interactions\n\n- Work with DevOps to deploy models (if self-hosted)\n\nRequired Skills & Experience :\n\nMust Have :\n\n- Experience working with LLMs and GenAI technologies\n\n- Strong experience with OpenAI API, Anthropic Claude, or similar LLM APIs\n\n- Proficiency in Python (FastAPI, LangChain, LlamaIndex preferred)\n\n- Strong understanding of prompt engineering techniques and best practices\n\n- Experience with vector databases (Pinecone, Weaviate, Qdrant, ChromaDB)\n\n- Knowledge of RAG (Retrieval Augmented Generation) implementation\n\n- Understanding of transformer architecture and attention mechanisms\n\n- Experience with API integration, webhooks, and streaming responses\n\n- Strong problem-solving skills and ability to debug complex AI systems\n\n- Experience with LangChain, LlamaIndex, or similar LLM frameworks\n\n- Knowledge of fine-tuning techniques (LoRA, QLoRA, PEFT)\n\n- Experience with embedding models and semantic search\n\n- Familiarity with HuggingFace Transformers library\n\n- Experience deploying models using vLLM, TGI (Text Generation Inference)\n\n- Knowledge of function calling/tool use with LLMs\n\n- Experience with model evaluation metrics (BLEU, ROUGE, BERTScore)\n\n- Understanding of token economics and cost optimisation\n\n- Experience with open-source models (Llama, Mistral, Falcon)\n\n- Knowledge of model quantization and optimization techniques\n\n- Experience with multi-modal models (vision, audio)\n\n- Familiarity with MLOps practices and experiment tracking (Weights & Biases, MLflow)\n\n- Experience with AWS SageMaker, Google Vertex AI, or Azure ML\n\n- Understanding of chain-of-thought prompting, ReAct, agents\n\n- Experience building chatbots or conversational AI systems\n\n- Publications or contributions to AI/ML community\n\nSkills\nGenerative AI, LLM, Python, LangChain, Prompt Engineering, RAG, AI Integration, VectorDB, Cloud, Artificial Intelligence","description_format":"text","description_chars":4061,"description_truncated":false,"requirements":{"experience_years_min":2,"management_years_min":null,"team_size_min":null,"manages_managers":false,"education":null,"security_clearance":false,"languages":[]},"benefits":[],"hiring_locations":[],"hiring_excludes":[],"relocation_offered":false,"industries":[],"lifecycle":[{"event":"open","at":"2026-10-08T12:29:00Z"}],"visa":[],"liveness":{"score":86,"band":"hot","label":"Hiring now","p_open":1,"p_active":0.86,"p_room":1,"age_days":1,"expected_fill_days":16,"reasons":["seen:1","win:early","comp:junior"],"computed_at":"2026-10-10T05:45:15Z"},"pay":null,"html_url":"https://alion.io/job/smartncode-aigenai-engineer-llm-integration","json_url":"https://alion.io/job/smartncode-aigenai-engineer-llm-integration.json","meta":{"generated_at":"2026-10-11T05:13:40Z","cache_seconds":300,"methodology":"https://alion.io/methodology","terms":"https://alion.io/terms","contact":"https://alion.io/contact","api":"https://alion.io/developers","about":"Alion is a live layer of people, companies and AI agents: who they are, whether they are real and active right now, what they do and how to work with them, readable by people and by agents and paid per call.","catalog":"https://alion.io/catalog.json","usage":{"tier":"crawler","counted_by":"address","units_charged":1,"used_today":4742,"day_limit":5000,"remaining_today":258,"minute_limit":60,"resets_at":"2026-10-12T00:00:00Z"}},"offers":[{"id":"company.slices","title":"One company in depth, by slice","status":"live","price":{"credits":0.02,"usd":0.002,"plus_per_slice":{"credits":0.05,"usd":0.005}},"unit":"per company, plus each slice with data","note":"the employer in depth","call":{"mcp_tool":"get_company","arguments":{"id":3912697},"rest":"https://alion.io/mcp/rest/get_company?id=3912697"},"human":"https://alion.io/catalog?offer=company.slices&for=job%2Fsmartncode-aigenai-engineer-llm-integration"},{"id":"market.stats","title":"A market slice: pay, demand and time to fill","status":"live","price":{"credits":1,"usd":0.1},"unit":"per slice","note":"pay, demand and time to fill for this role and place","call":{"mcp_tool":"market_stats"},"human":"https://alion.io/catalog?offer=market.stats&for=job%2Fsmartncode-aigenai-engineer-llm-integration"},{"id":"job.search","title":"Open jobs by role, technology, place, pay and visa","status":"live","price":{"credits":0.02,"usd":0.002},"unit":"per posting in a list","note":"similar open postings","call":{"mcp_tool":"search_jobs"},"human":"https://alion.io/catalog?offer=job.search&for=job%2Fsmartncode-aigenai-engineer-llm-integration"},{"id":"company.verify","title":"Is this company real and active right now","status":"pilot","price":null,"unit":"per company","request":{"url":"https://alion.io/catalog/request","method":"POST","body":"{\"offer\": \"company.verify\", \"for\": \"job/smartncode-aigenai-engineer-llm-integration\", \"note\": \"what you need it for\"}"},"human":"https://alion.io/catalog?offer=company.verify&for=job%2Fsmartncode-aigenai-engineer-llm-integration"}]}