{"id":1229949,"url":"https://alion.io/job/square-root-consulting-ai-software-stack-validation-performance-engineer","title":"AI Software Stack Validation & Performance Engineer","company":{"id":3800022,"name":"Square Root Consulting","domain":"squarerootindia.com","url":"https://alion.io/company/square-root-consulting","size_band":null,"is_staffing_agency":false,"employer_type":"direct","is_intermediary":false,"listed_via":null,"ats_vendor":null,"truth_index":null},"role":"AI/ML","role_family":"AI/ML","seniority":"senior","employment_type":null,"work_mode":"on_site","remote_scope":null,"remote_scope_basis":null,"remote_working_hours":null,"hiring_geo_confidence":"structured","locations":["Bengaluru, India"],"countries":["IN"],"hiring_countries":[],"hiring_countries_total":0,"salary":null,"salary_estimate":{"min_usd":41000,"max_usd":104000,"period":"year","method":"role_seniority_country_cell","sample_n":10},"experience_years_min":8,"visa_sponsorship":false,"relocation_package":false,"has_equity":false,"technologies":[{"name":"C++","optional":false},{"name":"CUDA","optional":false},{"name":"CUDA Toolkit","optional":false},{"name":"Pytest","optional":false},{"name":"Python","optional":false},{"name":"PyTorch","optional":false},{"name":"PyTorch C++","optional":false},{"name":"Triton","optional":false},{"name":"vLLM","optional":false},{"name":"CI/CD","optional":true},{"name":"LLM","optional":true},{"name":"Machine Learning","optional":true}],"status":"live","first_seen_at":"2026-09-21T12:15:54Z","employer_posted_date":null,"last_verified_at":"2026-09-21T12:15:54Z","board_verified":false,"closed_at":null,"days_open":6,"trust":{"level":"not_scored","repost_count":null,"flags":[],"days_open":6},"description":"AI Software Stack Validation & Performance Engineer\n\nLocation : Bangalore\n\nAbout the Role :\n\nWe are building a next-generation AI software stack optimized for high-performance execution on custom AI hardware. We are looking for a strong AI/ML Systems Engineer who can use and validate the platform end-to-end - from ML frameworks and model serving to compilers, runtimes, and low-level kernels.\n\nThis is an engineering-focused role, not traditional QA. You will develop and run real-world AI workloads, benchmark performance, debug complex issues across the software stack, and work closely with compiler, runtime, framework, and kernel teams to improve the platform.\n\nKey Responsibilities :\n\n- Develop, port, run, and validate real-world AI/ML workloads end-to-end.\n\n- Build and benchmark workloads to identify functional issues, performance bottlenecks, regressions, and usability gaps.\n\n- Develop automated validation and regression test suites using Python, C++, PyTest, and related tools.\n\n- Work across PyTorch, model serving, compilers, runtimes, and AI kernels.\n\n- Write, modify, optimize, and benchmark Triton/CUDA kernels.\n\n- Debug issues across multiple layers, from Python/framework behavior to compiled kernels and runtime execution.\n\n- Perform performance profiling and benchmarking using relevant ML performance tools, including MLPerf.\n\n- Collaborate with compiler, runtime, framework, and kernel engineers to triage and resolve issues.\n\n- Build scalable automation and CI-integrated validation infrastructure.\n\n- Contribute to release quality, regression tracking, and performance analysis.\n\nRequired Qualifications :\n\n- 8+ years of overall industry experience, with 3+ years in AI/ML systems, software validation, or performance engineering.\n\n- Bachelor's or Master's degree in Computer Science, Electrical Engineering, or a related field.\n\n- Strong programming skills in Python and C++.\n\n- Hands-on experience with PyTorch and ML model development/debugging.\n\n- Experience with LLMs, inference, model serving, or vLLM.\n\n- Experience writing or modifying GPU/AI accelerator kernels using Triton, CUDA, or similar technologies.\n\n- Strong experience with PyTest or similar automated testing frameworks.\n\n- Understanding of compilers, runtimes, operators, kernel execution, and software stacks.\n\n- Strong analytical and debugging skills with the ability to trace issues across different layers of a complex system.\n\nGood to Have :\n\n- Experience enabling new ML models/architectures on GPUs or AI accelerators.\n\n- Experience with ML workload profiling and performance optimization.\n\n- Ability to understand compiler-generated code and IR-level representations.\n\n- Experience building CI/CD and automated validation infrastructure.\n\n- Exposure to GPU, custom accelerator, or heterogeneous computing platforms.\n\n- Experience with containers and orchestration for ML serving.\n\n- Experience with software release processes and quality metrics.\n\nKey Skills :\n\n- Python, C++, PyTorch, LLM, vLLM, Triton, CUDA, AI/ML Systems, GPU, AI Accelerators, Model Serving, MLPerf, Performance Engineering, Compilers, Runtime, PyTest, CI/CD, Kernel Development, ML Infrastructure\nSkills\nPython, C++, PyTorch, CUDA, PyTest, Artificial Intelligence, Machine Learning, LLM, Hardware Architecture","description_format":"text","description_chars":3288,"description_truncated":false,"requirements":{"experience_years_min":8,"management_years_min":null,"team_size_min":null,"manages_managers":false,"education":null,"security_clearance":false,"languages":[]},"benefits":[],"hiring_locations":[],"hiring_excludes":[],"relocation_offered":false,"industries":[],"lifecycle":[{"event":"open","at":"2026-09-25T14:00:00Z"}],"liveness":{"score":84,"band":"hot","label":"Hiring now","p_open":1,"p_active":0.843,"p_room":1,"age_days":5,"expected_fill_days":24,"reasons":["seen:5","velocity","win:early"],"computed_at":"2026-09-27T05:45:00Z"},"pay":null,"html_url":"https://alion.io/job/square-root-consulting-ai-software-stack-validation-performance-engineer","json_url":"https://alion.io/job/square-root-consulting-ai-software-stack-validation-performance-engineer.json","meta":{"generated_at":"2026-09-28T04:21:08Z","cache_seconds":300,"methodology":"https://alion.io/methodology","terms":"https://alion.io/terms","contact":"https://alion.io/contact","api":"https://alion.io/developers","usage":{"tier":"crawler","counted_by":"address","units_charged":1,"used_today":2616,"day_limit":5000,"remaining_today":2384,"minute_limit":60,"resets_at":"2026-09-29T00:00:00Z"}}}