{"id":1227954,"url":"https://alion.io/job/algoleap-technologies-junior-site-reliability-engineer-devops","title":"Junior Site Reliability Engineer - DevOps","company":{"id":3800168,"name":"Algoleap","domain":"algoleap.com","url":"https://alion.io/company/algoleap-technologies","size_band":"1001-5000","is_staffing_agency":false,"employer_type":"direct","is_intermediary":false,"listed_via":null,"ats_vendor":null,"truth_index":null},"role":"DevOps","role_family":"DevOps","seniority":"junior","employment_type":null,"work_mode":"on_site","remote_scope":null,"remote_scope_basis":null,"remote_working_hours":null,"hiring_geo_confidence":"structured","locations":["Hyderabad, India"],"countries":["IN"],"hiring_countries":[],"hiring_countries_total":0,"salary":null,"salary_estimate":null,"experience_years_min":2,"visa_sponsorship":false,"relocation_package":false,"has_equity":false,"technologies":[{"name":"Agile","optional":false},{"name":"AIOps","optional":false},{"name":"Amazon CloudWatch","optional":false},{"name":"Amazon EC2","optional":false},{"name":"Amazon ECS","optional":false},{"name":"Amazon EKS","optional":false},{"name":"Amazon S3","optional":false},{"name":"Amazon SageMaker","optional":false},{"name":"Anomaly Detection","optional":false},{"name":"AWS","optional":false},{"name":"AWS Glue","optional":false},{"name":"AWS Lambda","optional":false},{"name":"Azure","optional":false},{"name":"Azure DevOps","optional":false},{"name":"Bash","optional":false},{"name":"Blue-Green Deployment","optional":false},{"name":"CI/CD","optional":false},{"name":"Claude","optional":false},{"name":"Datadog","optional":false},{"name":"Docker","optional":false},{"name":"ElasticSearch","optional":false},{"name":"Fine-tuning","optional":false},{"name":"GitHub","optional":false},{"name":"IAM","optional":false},{"name":"Incident Management","optional":false},{"name":"Jenkins","optional":false},{"name":"Kibana","optional":false},{"name":"Kubeflow","optional":false},{"name":"Kubernetes","optional":false},{"name":"Linux","optional":false},{"name":"LLM","optional":false},{"name":"LLMOps","optional":false},{"name":"MLFlow","optional":false},{"name":"OpenAI","optional":false},{"name":"PowerShell","optional":false},{"name":"Prometheus","optional":false},{"name":"Python","optional":false},{"name":"Scrum","optional":false},{"name":"Self-Healing","optional":false},{"name":"Terraform","optional":false}],"status":"live","first_seen_at":"2026-09-22T10:51:45Z","employer_posted_date":null,"last_verified_at":"2026-09-22T10:51:45Z","board_verified":false,"closed_at":null,"days_open":8,"trust":{"level":"not_scored","repost_count":null,"flags":[],"days_open":8},"description":"About the job :\n\nAs a Junior SRE/DevOps Engineer, you will support the setup and maintenance of Infra, CI/CD and help sustain different key projects used throughout globally, working under the guidance of senior engineers.\n\nAs a member of our geographically distributed development team your communication and analytical skills are essential to the role.\n\nKey Responsibilities:\n\n- Assist in designing and maintaining cloud infrastructure that is secure, scalable, and highly available on AWS/Azure\n\n- Work collaboratively with software engineering to define infrastructure and deployment requirements\n\n- Provision, configure and maintain AWS cloud infrastructure defined as code.\n\n- Containerization using Docker and Kubernetes\n\n- Troubleshoot problems across a wide array of services and functional areas\n\n- Build and maintain operational tools for deployment, monitoring, and analysis of AWS infrastructure and systems\n\n- Assist with infrastructure cost analysis and support optimization efforts.\n\n- Support the development of self-healing and automated remediation mechanisms using AI/ML techniques\n\n- Assist in integrating AI/LLM capabilities into DevOps workflows (e.g., log analysis, automated RCA, deployment insights)\n\n- Support monitoring strategy enhancements by leveraging intelligent alerting, noise reduction, and pattern-based anomaly detection across logs, metrics, and traces.\n\n- Assist in building and maintaining MLOps pipelines for model training, deployment, and continuous improvement.\n\n- Collaborate with a global team of engineers in a highly agile DevOps environment, focused on efficient operation of daily activities, developer productivity and continuous improvement of the framework.\n\n- Support the development, implementation, and maintenance of CI/CD frameworks, and contribute to tools development for hybrid environments (Cloud, On premise) with a vision to achieve \"CI/CD\" objectives for large-scale integration of systems in order to reduce manual build and deploy efforts.\n\n- Work with geographically dispersed teams including multi-vendor into Scrum teams to meet \"CI/CD\"\n\nRequired Knowledge & Skills:\n\n- 2 - 4 years of experience building and maintaining AWS infrastructure (VPC, EC2, Security Groups, IAM, ECS/EKS, CloudFront, S3, RDS, SQS, SNS, Lambda Function, Batch jobs, AWS Glue)\n\n- Working understanding of how to secure AWS environments and meet compliance requirements\n\n- Working knowledge of deploying and managing infrastructure with Terraform.\n\n- Exposure to or working knowledge of LLMs (OpenAI, Azure OpenAI, Claude etc.)\n\n- Basic awareness of LLMOps concepts (prompt management, model evaluation, versioning, fine-tuning lifecycle)\n\n- Familiarity with MLOps tools such as MLflow, SageMaker, Kubeflow or equivalent.\n\n- Familiarity with AIOps platforms/tools for intelligent monitoring and incident management.\n\n- Ability to apply AI techniques to improve deployment speed, reliability, and monitoring effectiveness.\n\n- Working experience on windows & Linux based environments.\n\n- Experience with Docker, GitHub, Jenkins, Azure DevOps, ELK and deploying applications on AWS.\n\n- Good command on scripting languages like Python, Bash/Shell, Powershell etc\n\n- Knowledge in log analytics tools like Elastic search and Kibana.\n\n- Basic knowledge of Cloud Migration/Disaster Recovery/Blue Green Deployment implementation.\n\n- Good understanding about monitoring the services and alerting using Cloudwatch, Datadog, Prometheus or Azure monitor.\n\n- Good to hire a candidate with certification\n\nPersonal Attributes:\n\n- Very good communication skills.\n\n- Ability to easily fit into a distributed development team.\n\n- Customer service oriented.\n\n- Enthusiastic/High initiative.\n\n- Ability to manage timelines of multiple initiatives.\n\n- Very good attention to detail and the ability to always follow up\nSkills\nSite Reliability, DevOps, CI/CD, Docker, Kubernetes, AWS, MLOps, AWS Infrastructure, Elastic Kubernetes Service, IAC Terraform, Monitoring Tools","description_format":"text","description_chars":3992,"description_truncated":false,"requirements":{"experience_years_min":2,"management_years_min":null,"team_size_min":null,"manages_managers":false,"education":null,"security_clearance":false,"languages":[]},"benefits":[],"hiring_locations":[],"hiring_excludes":[],"relocation_offered":false,"industries":["Science & Engineering","Engineering Services"],"lifecycle":[{"event":"open","at":"2026-09-25T13:06:44Z"}],"liveness":{"score":79,"band":"hot","label":"Hiring now","p_open":1,"p_active":0.875,"p_room":0.9,"age_days":7,"expected_fill_days":19,"reasons":["seen:7","velocity","win:mid","comp:junior"],"computed_at":"2026-09-30T05:45:00Z"},"pay":null,"html_url":"https://alion.io/job/algoleap-technologies-junior-site-reliability-engineer-devops","json_url":"https://alion.io/job/algoleap-technologies-junior-site-reliability-engineer-devops.json","meta":{"generated_at":"2026-10-01T03:16:49Z","cache_seconds":300,"methodology":"https://alion.io/methodology","terms":"https://alion.io/terms","contact":"https://alion.io/contact","api":"https://alion.io/developers","usage":{"tier":"crawler","counted_by":"address","units_charged":1,"used_today":2761,"day_limit":5000,"remaining_today":2239,"minute_limit":60,"resets_at":"2026-10-02T00:00:00Z"}}}