{"id":1427065,"url":"https://alion.io/job/nference-site-reliability-engineer","title":"Site Reliability Engineer","company":{"id":34471,"name":"Nference","domain":"nference.com","url":"https://alion.io/company/nference","size_band":"51-200","is_staffing_agency":false,"employer_type":"direct","is_intermediary":false,"listed_via":null,"ats_vendor":"Keka","truth_index":null},"role":"DevOps","role_family":"DevOps","seniority":"junior","employment_type":"full_time","work_mode":"hybrid","remote_scope":null,"remote_scope_basis":null,"remote_working_hours":null,"hiring_geo_confidence":"explicit","locations":["Bengaluru, India"],"countries":["IN"],"hiring_countries":[],"hiring_countries_total":0,"salary":null,"salary_estimate":null,"experience_years_min":1,"visa_sponsorship":false,"relocation_package":false,"has_equity":false,"technologies":[{"name":"AWS","optional":false},{"name":"Azure","optional":false},{"name":"Bash","optional":false},{"name":"CI/CD","optional":false},{"name":"Datadog","optional":false},{"name":"DNS","optional":false},{"name":"Docker","optional":false},{"name":"GCP","optional":false},{"name":"Git","optional":false},{"name":"Grafana","optional":false},{"name":"Incident Management","optional":false},{"name":"Kubernetes","optional":false},{"name":"Linux","optional":false},{"name":"Machine Learning","optional":false},{"name":"Platform Engineering","optional":false},{"name":"Prometheus","optional":false},{"name":"Python","optional":false},{"name":"TCP/IP","optional":false},{"name":"Terraform","optional":false},{"name":"Ansible","optional":true},{"name":"Apache Kafka","optional":true},{"name":"Configuration Management","optional":true},{"name":"RabbitMQ","optional":true}],"status":"live","first_seen_at":"2026-09-28T12:55:14Z","employer_posted_date":"2026-09-28","last_verified_at":"2026-09-29T00:04:22Z","board_verified":true,"closed_at":null,"days_open":0,"trust":{"level":"ok","repost_count":null,"flags":[],"days_open":0},"description":"Site Reliability Engineer (SRE)\nLocation: Bangalore, India\nAbout nference\nnference is an AI-first healthcare technology company dedicated to accelerating biomedical discovery and transforming healthcare through cutting-edge artificial intelligence and large-scale computing. Our platforms help pharmaceutical companies, healthcare providers, and research organizations unlock insights from complex clinical, molecular, and imaging data to drive better decisions and improve patient outcomes.\nFounded with the vision of solving some of healthcare's most challenging problems, nference combines deep expertise in software engineering, machine learning, data science, and life sciences to build products that make a meaningful impact on global healthcare. Recognized by The Washington Post as the \"Google of Biomedicine,\" we continue to push the boundaries of innovation by partnering with leading healthcare institutions and industry pioneers.\nOur People\nOur people are our greatest strength. At nference, you'll work alongside exceptional software engineers, AI researchers, physicians, scientists, and product leaders who are passionate about solving meaningful, real-world problems.\nWe foster a culture built on ownership, curiosity, collaboration, and continuous learning. Engineers are encouraged to challenge assumptions, contribute ideas, influence technical direction, and take ownership of their work from concept to production. Whether you're improving platform reliability, automating cloud infrastructure, or enhancing developer productivity, you'll collaborate with some of the brightest minds in technology and healthcare.\nThe Opportunity\nWe're looking for a Site Reliability Engineer (SRE) who is passionate about building and maintaining reliable, scalable, and highly available cloud infrastructure. In this role, you'll help ensure the stability, performance, and operational excellence of the platforms that power our AI-driven healthcare products.\nYou'll work closely with Software Engineering, Platform Engineering, AI, Data Science, and Product teams to improve system reliability through automation, observability, and modern infrastructure practices. This role is ideal for someone who enjoys solving operational challenges, automating repetitive tasks, and continuously improving platform resilience and developer experience.\nWhat You'll Do\nMonitor the health, availability, and performance of cloud infrastructure, Kubernetes clusters, CI/CD systems, and platform services.\nAssist in production incident response by collecting logs, metrics, and diagnostic information to support rapid troubleshooting and recovery.\nConfigure, maintain, and improve monitoring dashboards, alerting systems, and observability platforms.\nDevelop automation scripts and operational tooling to eliminate manual processes and improve infrastructure efficiency.\nContribute to Infrastructure-as-Code (IaC) modules for provisioning and managing cloud resources.\nSupport the reliability, maintenance, and continuous improvement of CI/CD pipelines and deployment workflows.\nParticipate in production operations while learning and applying Site Reliability Engineering principles, including SLIs, SLOs, error budgets, and incident management.\nCollaborate with software engineers to improve system reliability, scalability, and deployment processes.\nTroubleshoot infrastructure, networking, and platform-related issues across cloud environments.\nMaintain operational documentation, runbooks, and knowledge repositories to improve incident response and operational consistency.\nContinuously identify opportunities to improve platform reliability, automation, and operational excellence.\nWhat We're Looking For\nRequired Qualifications\nBachelor's or Master's degree in Computer Science, Information Technology, Software Engineering, or a related technical discipline.\n1-3 years of professional experience in Site Reliability Engineering, DevOps, Cloud Infrastructure Engineering, or a related systems engineering role.\nStrong understanding of Linux systems administration, command-line tools, process management, system diagnostics, and troubleshooting.\nProficiency in at least one scripting or programming language such as Python, Bash, or Shell for automation and tooling.\nHands-on experience with at least one major cloud platform (AWS, GCP, or Azure) and familiarity with core infrastructure services including compute, networking, and storage.\nWorking knowledge of containerization and orchestration technologies such as Docker and Kubernetes.\nFamiliarity with Infrastructure-as-Code (IaC) concepts and tools such as Terraform or similar automation frameworks.\nExperience with monitoring, logging, and observability platforms such as Prometheus, Grafana, ELK Stack, Datadog, or similar tools.\nGood understanding of networking fundamentals including DNS, TCP/IP, load balancing, and distributed systems concepts.\nFamiliarity with CI/CD pipelines and modern software delivery practices.\nExperience using version control systems such as Git.\nStrong analytical, troubleshooting, and problem-solving skills.\nExcellent verbal and written communication skills.\nPreferred Qualifications\nExposure to incident management and production support in cloud-native environments.\nFamiliarity with reliability engineering concepts including SLIs, SLOs, and error budgets.\nExperience with configuration management or automation tools such as Ansible.\nKnowledge of messaging technologies such as Kafka or RabbitMQ.\nExposure to distributed systems and microservices architectures.\nExperience supporting highly available production environments.\nFamiliarity with security best practices for cloud infrastructure.\nContributions to open-source projects are a plus.\nWhy Join nference?\nAt nference, you'll work on technology that has the potential to transform healthcare and improve lives across the world. Every engineering challenge you solve contributes to advancing biomedical research, accelerating drug discovery, and enabling better clinical decisions.\nYou'll be part of a team that values innovation, ownership, and collaboration, where you'll have the opportunity to work with cutting-edge cloud technologies, modern infrastructure platforms, and automation tools while growing alongside some of the brightest minds in software engineering, artificial intelligence, and life sciences.\nIf you're excited about building reliable platforms that enable engineering teams to innovate faster while delivering meaningful real-world impact, nference is the place for you.\nBenefits & Perks\n Industry Prestige: Build your career at the \"Google of Biomedicine\" (as recognized by The Washington Post), working alongside exceptional software engineers, physicians, scientists, and researchers.\nCutting-Edge Innovation: Solve complex healthcare challenges using advanced AI, machine learning, and large-scale clinical, molecular, and imaging datasets.\nMeaningful Impact: Contribute to technologies that accelerate drug discovery and biomedical research, with opportunities to be recognized as a contributing author on high-impact scientific publications where applicable.\nGrowth & Flexibility: Thrive in a collaborative, innovation-driven culture with continuous learning opportunities and a hybrid work model for eligible employees after successful completion of the three-month probation period.\nWellness & Perks: Enjoy reimbursements for gym memberships, technology gadgets, high-speed internet, professional development, comprehensive health insurance, and complimentary breakfast, lunch, and snacks at our Bangalore office.\nEqual Opportunity Employer\nnference is committed to building a diverse, equitable, and inclusive workplace where everyone has the opportunity to thrive. We celebrate different perspectives, backgrounds, and experiences because they strengthen our teams and drive innovation. We are proud to be an equal opportunity employer and welcome applicants from all backgrounds to join us in building technology that transforms healthcare.","description_format":"text","description_chars":8026,"description_truncated":false,"requirements":{"experience_years_min":1,"management_years_min":null,"team_size_min":null,"manages_managers":false,"education":{"level":"bachelor","optional":false},"security_clearance":false,"languages":[]},"benefits":["Continuous learning","Gym membership","Health insurance","Hybrid work","Professional development"],"hiring_locations":[{"name":"India","iso":"IN","kind":"country"}],"hiring_excludes":[],"relocation_offered":false,"industries":["Medical AI","Health Data & Interoperability"],"lifecycle":[{"event":"open","at":"2026-09-29T00:04:22Z"}],"liveness":{"score":86,"band":"hot","label":"Hiring now","p_open":1,"p_active":0.86,"p_room":1,"age_days":0,"expected_fill_days":30,"reasons":["conf:2","win:early","comp:junior"],"computed_at":"2026-09-29T02:54:00Z"},"pay":null,"html_url":"https://alion.io/job/nference-site-reliability-engineer","json_url":"https://alion.io/job/nference-site-reliability-engineer.json","meta":{"generated_at":"2026-09-29T02:54:00Z","cache_seconds":300,"methodology":"https://alion.io/methodology","terms":"https://alion.io/terms","contact":"https://alion.io/contact","api":"https://alion.io/developers","usage":{"tier":"crawler","counted_by":"address","units_charged":1,"used_today":2570,"day_limit":5000,"remaining_today":2430,"minute_limit":60,"resets_at":"2026-09-30T00:00:00Z"}}}