This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Senior Linux Systems Engineer based in India.
This is a hands-on technical leadership opportunity within a Linux Operations environment supporting mission-critical production systems.
You will manage, troubleshoot, and optimize Linux infrastructure across hosted, cloud, and remote environments.
The role combines advanced incident response, root-cause analysis, automation, infrastructure management, and continuous improvement.
You will work across AWS, Azure, and GCP environments while helping maintain secure, highly available, and reliable services.
Beyond technical delivery, you will serve as a trusted escalation point and mentor for other engineers.
You will collaborate closely with internal teams and external stakeholders, providing clear technical direction during incidents and changes.
This role is well suited to an experienced Linux professional who enjoys solving complex infrastructure challenges and driving operational excellence.
Accountabilities:
- Take end-to-end ownership of high-priority incidents, major incidents, service requests, and operational tasks, resolving issues with minimal escalation.
- Lead advanced troubleshooting, root-cause analysis, stabilization, and recovery activities for critical Linux production environments.
- Implement sustainable fixes and continuous improvements to reduce recurring incidents, improve reliability, and minimize MTTR.
- Manage and optimize Linux servers across hosted infrastructure, public cloud environments, and remote infrastructure.
- Develop and maintain Ansible playbooks, Bash scripts, and other automation to eliminate repetitive manual tasks and improve operational efficiency.
- Support infrastructure-as-code, configuration management, deployment, orchestration, Docker, Kubernetes, and related automation initiatives.
- Execute infrastructure changes through established ServiceNow and change-management frameworks, including impact analysis, risk assessment, validation, and stakeholder coordination.
- Create, review, and maintain SOPs, operational runbooks, disaster recovery procedures, and technical documentation.
- Identify procedural gaps and drive standardization across operational teams while supporting audit and compliance requirements.
- Monitor system health, configure alerts and dashboards, and use observability data to proactively identify performance and reliability issues.
- Maintain security, patching, vulnerability remediation, high-availability, and compliance standards across supported infrastructure.
- Act as a technical escalation point and mentor for junior engineers, providing guidance during complex troubleshooting and knowledge-sharing activities.
- Communicate clearly with internal teams and external clients regarding incident status, resolution timelines, maintenance activities, and service performance.
- Support SLA commitments and participate in 24/7 rotational shifts and on-call support, as required.
- 8+ years of hands-on production Linux/Unix administration and support experience in enterprise environments.
- Strong expertise with RHEL, CentOS, and Ubuntu, including LTS releases and operating-system lifecycle management.
- Advanced troubleshooting capabilities across Linux kernels, networking, storage, system performance, and related infrastructure components.
- Practical experience with Linux patching, security updates, vulnerability remediation, and system hardening.
- Strong Bash/shell scripting skills and proven experience developing automation with Ansible.
- Proficiency with systemd, RPM/APT package management, kernel modules, SSH, sudo, LDAP, and Active Directory authentication.
- Solid knowledge of Linux security technologies and practices, including SELinux, AppArmor, firewalls, and file permissions.
- Production experience supporting cloud-based infrastructure, particularly AWS EC2, GCP, and Azure Virtual Machines.
- Understanding of managed services, IaaS, DNS, routing, load balancing, and core infrastructure concepts.
- Experience with monitoring and observability platforms such as Datadog, Azure Monitor, Prometheus, or comparable technologies.
- Ability to work with alerts, dashboards, logs, metrics, traces, and other observability data to diagnose and prevent operational issues.
- Experience with incident management, change management, risk assessment, service restoration, and continuous improvement.
- Strong technical documentation skills, including the ability to create clear SOPs, runbooks, and operational procedures.
- Excellent communication and stakeholder-management skills, with the ability to explain complex technical issues clearly.
- Strong prioritization and time-management abilities in a multitasking, high-pressure environment.
- A collaborative and mentoring mindset, with the confidence to guide engineers through complex technical challenges.
- High attention to detail and a strong commitment to reliability, quality, security, and compliance.
- Willingness and ability to participate in 24/7 rotational shifts and on-call support when required.
- Opportunity to work on mission-critical Linux infrastructure in complex enterprise and cloud environments.
- Exposure to AWS, Azure, and GCP alongside modern automation, monitoring, and infrastructure technologies.
- Hands-on experience with automation and operational improvement using tools such as Ansible, Bash, Docker, and Kubernetes.
- Technical leadership and mentoring opportunities within a collaborative operations environment.
- Opportunity to work on complex incidents, infrastructure optimization, security, and reliability initiatives.
- Exposure to structured IT service management, change management, SLA, compliance, and disaster recovery practices.
- Opportunity for continuous technical development across Linux, cloud, automation, and infrastructure operations.
- India-based opportunity with support for 24/7 rotational shifts and on-call operations as part of the role.
- Compensation, healthcare coverage, leave, and additional employment benefits are determined by the hiring company and will be communicated during the recruitment process.

