We are looking for a Night Operations Engineer to keep the platform healthy while the rest of the team is offline through the night and on public holidays. You are the first line of defense during off-hours: watching for failures, judging how serious they are, and getting the right person on the problem fast. Your core job is not to fix everything yourself. It is to detect, classify, and escalate a spot when a database is down, a service has stopped responding, or the UI has broken; decide how urgent it is; and contact the right owner to resolve it. Clear judgement and clear communication matter more than deep experience.
This is a strong opportunity for a fresher or junior engineer, and we make it workable for one by giving you a severity matrix, an escalation ladder, runbooks, and a standing when in doubt, escalate policy. You will never be expected to make a hard call alone with no support. If you're a night owl who takes pride in keeping things running, we'd love to have you.
The core responsibilities for the job include the following:
System Monitoring and Anomaly Detection:
- Monitor platform health, uptime dashboards, and alerting systems throughout the shift, identifying and flagging anomalies before they escalate.
- Watch for critical failures: databases going down, services becoming unreachable or stopping responding, and UI breakage on key user flows.
- Note: backend uptime checks will not catch a broken frontend that still returns HTTP 200 visual checks and periodic manual smoke tests of critical flows (login, checkout, core reads) will catch UI breakage.
Incident Triage and Severity Assessment:
- For every alert or report, assign a severity using the matrix below, deciding whether it warrants immediate escalation or a handoff note.
- Triage incoming issues, categorize, prioritize, and route to the correct engineering owner.
- Recognize and immediately escalate suspected security incidents (unusual traffic, authentication spikes, data-access anomalies). Escalate; do not investigate.
- Log every incident completely: symptoms, timeline, severity, steps taken, and resolution or escalation status.
Escalation and On-Call Coordination:
- Contact the right owner for the affected service using the service-to-owner map and escalation ladder below.
- If the primary owner does not respond within the target window, follow the ladder to the secondary, and then the on-call lead never leaves a critical issue without an owner.
- Escalate with full context: what is broken, since when, severity, user impact, and what you have already checked.
Handoff Both Directions:
- Receive a start-of-shift handoff: known issues, planned deploys, and maintenance windows, so you don't escalate a planned outage.
- Prepare an end-of-shift handoff for the day team: what happened, what was resolved, what is still open, and what needs immediate attention.
- Maintain and update runbooks, FAQs, and support documentation as issues are encountered and resolved.
Learning and Growth:
- Build a working understanding of the product architecture, key services, and common failure modes over time.
- Flag recurring issues and patterns to the engineering lead, contributing to root-cause elimination, not just ticket closure.
- Participate in knowledge-sharing sessions and training to grow technical depth.
Requirements:
- 0-2 years in a technical support, helpdesk, QA, or junior engineering role: freshers with strong fundamentals are welcome.
- Basic understanding of how web applications work: client-server architecture, HTTP requests, APIs, and databases at a conceptual level.
- Sound judgement under uncertainty, willing to follow a severity matrix and escalate when unsure, rather than guessing alone.
- Comfortable with ticketing systems (Jira, Zendesk, Freshdesk, or similar) and communication tools (Slack, email).
- Able to read and interpret basic logs, error messages, and system alertsyou don't need to fix every issue, but you need to understand what you're looking at.
- Strong written communication: your incident logs and handoff notes must be complete enough that the day team never has to chase you.
- Reliable and punctual: the night shift requires consistent attendance and the discipline to work independently.
- Genuinely curious: wants to understand the why, not just close the ticket.
- Degree or diploma in computer science, IT, or a related field, or an equivalent self-taught technical background.
Good to Have:
- Familiarity with Linux command-line basics: navigating directories, reading logs, and checking running processes.
- Exposure to cloud platforms (AWS, GCP, or Azure) services, regions, and uptime concepts.
- Experience with monitoring tools (Grafana, Datadog, New Relic, or similar).
- Exposure to paging / on-call tools (PagerDuty, Opsgenie, or similar).
- Basic scripting in Python or Bash for automating repetitive support tasks.
- Prior experience working night shifts or in a 24/7 operational environment.
Key Competencies:
- Reliability: you show up on time, every shift, and are ready to own the night.
- Calm under pressure: when an alert fires at 2 AM, you follow the process, not the panic.
- Sound judgement: you classify severity correctly and escalate up when unsure.
- Clear communication: your logs and handoffs are complete enough that no one has to chase you for context.
- Ownership: Open issues don't stay open without a clear owner; that owner is you until it's escalated or resolved.

