{"id":1231467,"url":"https://alion.io/job/cloudraft-senior-sre","title":"Senior SRE","company":{"id":1842586,"name":"CloudRaft","domain":"cloudraft.io","url":"https://alion.io/company/cloudraft","size_band":null,"is_staffing_agency":false,"employer_type":"direct","is_intermediary":false,"listed_via":null,"ats_vendor":"Gem","truth_index":null},"role":"DevOps","role_family":"DevOps","seniority":"senior","employment_type":"full_time","work_mode":"remote","remote_scope":"stated_countries","remote_scope_basis":"posting_text","remote_working_hours":null,"hiring_geo_confidence":"structured","locations":[],"countries":[],"hiring_countries":["IN"],"hiring_countries_total":1,"salary":null,"salary_estimate":{"min_usd":29000,"max_usd":64000,"period":"year","method":"role_seniority_country_cell","sample_n":8},"experience_years_min":5,"visa_sponsorship":false,"relocation_package":false,"has_equity":false,"technologies":[{"name":"AI Agents","optional":false},{"name":"AIOps","optional":false},{"name":"Amazon EKS","optional":false},{"name":"ArgoCD","optional":false},{"name":"AWS","optional":false},{"name":"Azure","optional":false},{"name":"Azure AKS","optional":false},{"name":"CI/CD","optional":false},{"name":"GCP","optional":false},{"name":"GitHub Actions","optional":false},{"name":"GitLab CI","optional":false},{"name":"Go","optional":false},{"name":"Google GKE","optional":false},{"name":"Grafana","optional":false},{"name":"Istio","optional":false},{"name":"Jenkins","optional":false},{"name":"Kubernetes","optional":false},{"name":"Loki","optional":false},{"name":"Mimir","optional":false},{"name":"Node JS","optional":false},{"name":"OpenShift","optional":false},{"name":"OpenTelemetry","optional":false},{"name":"Platform Engineering","optional":false},{"name":"Prometheus","optional":false},{"name":"Pulumi","optional":false},{"name":"Python","optional":false},{"name":"Terraform","optional":false},{"name":"Thanos","optional":false},{"name":"JavaScript","optional":true}],"status":"live","first_seen_at":"2026-08-10T05:58:36Z","employer_posted_date":"2026-08-20","last_verified_at":"2026-09-25T22:43:32Z","board_verified":true,"closed_at":null,"days_open":46,"trust":{"level":"ok","repost_count":null,"flags":[],"days_open":46},"description":"About CloudRaft\nCloudRaft is a premier cloud-native consulting and engineering company that helps ambitious startups and digital-first organizations build, scale, and operate mission-critical platforms. We partner with innovators at the forefront of artificial intelligence, developer productivity, observability, digital commerce, and enterprise software-enabling them to accelerate growth with resilient, scalable, and production-ready cloud infrastructure.\nOur experience spans organizations developing AI safety and governance platforms, AI Cloud, AI agent ecosystems, developer tooling, observability solutions, digital health products, customer engagement platforms, and technology-driven franchise networks. By combining deep expertise in Platform Engineering, Kubernetes, DevOps, Observability, and Cloud Native technologies, CloudRaft helps high-growth companies move faster, operate more reliably, and focus on building category-defining products.\nJob Description\nWe are looking for passionate Site Reliability Engineers (SREs) to join our growing team. In this role, you will take end-to-end ownership of designing, building, operating, and scaling mission-critical infrastructure for our partners. You will be responsible for ensuring reliability, performance, security, and operational excellence while driving automation, improving system efficiency, and implementing innovative solutions. Working at the intersection of software engineering and operations, you will help create resilient platforms that enable fast-growing organizations to scale with confidence.\nResponsibilities\nManage and maintain Kubernetes clusters across cloud platforms, including OpenShift, Amazon EKS, Azure AKS, and Google GKE.\nImplement and manage CI/CD pipelines using tools such as Jenkins, GitHub Actions, Argo CD, or GitLab CI/CD.\nDesign and maintain observability stacks with tools including Prometheus, Grafana, Loki, OpenTelemetry, and related technologies. Be part of the team who support open source projects like Prometheus, Thanos, Mimir, CloudNativePG, Istio and more.\nOptimize system performance and resolve production issues. Be part of the on call roster to provide 24x7 coverage for the critical production systems.\nImplement SRE principles, including Service Level Indicators (SLIs) and Service Level Objectives (SLOs), to uphold system reliability.\nAutomate infrastructure and operational tasks using programming languages such as Go or Python, and Infrastructure as Code (IaC) tools like Terraform.\nApply agentic AIto automate the SDLC lifecycle, AIOps and automation.\nLearn about emerging technologies, including AI, GPU Infrastructure\nContribute to knowledge sharing through technical writing and presentations.\nQualifications\nBachelor’s degree in Computer Science, Information Technology, or a related field.\n5+ years of experience in SRE, Platform Engineering, or DevOps Engineer.\nStrong expertise in Kubernetes, cloud-native technologies, on-premise and major cloud platforms (AWS, Azure, GCP).\nProficiency in programming languages such as Python or Go or Node.js.\nFamiliarity with CI/CD tools and modern deployment practices.\nProficiency in one or more open source observability stacks and Infrastructure as Code (Terraform/Pulumi).\nCKA/CKAD Certified (Brownie points!)\nExcellent problem-solving abilities and communication skills.\nInclination toward open-source contributions is advantageous.\nBenefits :\n- Competitive salary\n- Premium health insurance and various health & wellness benefits from a leading insurance provider through Plum\n- Opportunity to work on the latest AI stack and GPU infrastructure\n- Collaborative and supportive work environment full of learning\n- Chance to take a front seat where you lead and deliver","description_format":"text","description_chars":3745,"description_truncated":false,"requirements":{"experience_years_min":5,"management_years_min":null,"team_size_min":null,"manages_managers":false,"education":{"level":"bachelor","optional":false},"security_clearance":false,"languages":[]},"benefits":["Health insurance"],"hiring_locations":[{"name":"India","iso":"IN","kind":"country"}],"hiring_excludes":[],"relocation_offered":false,"industries":["DevOps & Platform Engineering","Cloud Consulting & Migration","Observability & APM"],"lifecycle":[{"event":"open","at":"2026-09-25T14:55:02Z"}],"liveness":{"score":20,"band":"cold","label":"Long shot","p_open":1,"p_active":0.449,"p_room":0.45,"age_days":46,"expected_fill_days":25,"reasons":["conf:3","win:tail"],"computed_at":"2026-09-26T02:06:46Z"},"pay":null,"html_url":"https://alion.io/job/cloudraft-senior-sre","json_url":"https://alion.io/job/cloudraft-senior-sre.json","meta":{"generated_at":"2026-09-26T02:06:46Z","cache_seconds":300,"methodology":"https://alion.io/methodology","terms":"https://alion.io/terms","contact":"https://alion.io/contact","api":"https://alion.io/developers","usage":{"tier":"crawler","counted_by":"address","units_charged":1,"used_today":2115,"day_limit":5000,"remaining_today":2885,"minute_limit":60,"resets_at":"2026-09-27T00:00:00Z"}}}