Your Role
Megaport is looking for an experienced, hands-on engineer in OSS and operations automation. This is a new role within Megaport’s operations department. The purpose of the role is to continue managing our monitoring and customer support systems, implement synergies with existing tools, and develop automation for day-to-day work. Your customers are the front-line support team, CS / TSE / NOC, Network Operations Engineering and Production Development Engineering teams.
To be successful in this role, you will need to be experienced in relevant technologies, have excellent interpersonal skills, be able to project-manage dependencies across technical departments and lead the implementation of operations tools/automation strategy.
We are looking for someone excited to lead the development of this capability. There is an abundance of high-value opportunities in this space for Megaport. Proven experience in designing, implementing and maintaining OSS tools is a must. The ideal candidate would have a genuine interest in analysing measurable operations work activity, exploring opportunities to develop systemic automations and managing the design and implementation of the identified opportunities.
What You’ll Be Doing
- Deploy, maintain, and optimize open-source monitoring stacks (Prometheus, Grafana, OpenNMS, Alerta) to modernize network assurance.
- Configure streaming telemetry, SNMP, and automated alert thresholds across cloud apps and network elements for rapid failure detection.
- Provide Level 3 troubleshooting and root-cause analysis (RCA) for complex system and network incidents.
- Build custom automations, scripts, and front-end user interfaces (Python, JavaScript, HTML, REST APIs, Message Bus) to streamline tools and eliminate manual operational bottlenecks.
- Maintain solution delivery end-to-end-from solution design and API integrations to QA testing and release management.
- Support deployment pipelines using GitHub-based CI/CD and cloud infrastructure on AWS.
- Maintain operational integrations within Salesforce Service Cloud and ticketing workflows to keep front-line support running smoothly.
- Administer Jira spaces, custom workflows, and native automations to streamline cross-team handoffs between NOC, CS, and Engineering teams.
- Build targeted operational dashboards (Power BI, Salesforce, Jira) to track SLAs and performance metrics.
- Partner with NOC, CS, and Engineering leads to translate operational pain points into clear technical requirements, user stories, and workflows.
- Author and update technical documentation, architecture diagrams, and operational playbooks in Confluence.
- Participate in scheduled on-call rotations to ensure high platform availability.
- Deploy, maintain, and optimize open-source monitoring stacks (Prometheus, Grafana, OpenNMS, Alerta) to modernize network assurance.
- Configure streaming telemetry, SNMP, and automated alert thresholds across cloud apps and network elements for rapid failure detection.
- Provide Level 3 troubleshooting and root-cause analysis (RCA) for complex system and network
1. OSS, Observability & Network Assurance
2. Software Engineering, Automation & DevOps
3. Business Systems & Ecosystem Integration (Salesforce & Atlassian)
4. Stakeholder Engagement & Process Improvement
What We Are Looking For
- 5+ years of experience in Telecom, NaaS, or ISP environments working with OSS/BSS and Fault Management systems.
- Proven hands-on experience with open-source observability tools (Prometheus, Grafana, OpenNMS, Alerta, or similar).
- Strong programming and scripting skills in Python, Shell, and front-end web basics (HTML, JavaScript), alongside solid Linux sysadmin and DevOps/CI/CD pipeline experience.
- Strong cross-team collaboration skills with a track record of driving operational improvements under pressure.
- Familiarity with Salesforce Service Cloud or declarative/programmatic automations.
- Experience configuring Jira spaces, custom workflows, or Atlassian automations.
- AWS cloud infrastructure and REST API / Webhook integration experience.
This is a broad, high-impact role with wide visibility across our operations ecosystem. We know no candidate will check every single box on this list! If you have strong hands-on experience in OSS/Fault Management, Python/DevOps, and Telecom systems, but are keen to expand your skills into our Salesforce and Atlassian environments, we want to hear from you.
Must-Haves:
Nice-to-Haves (Bonus points, not dealbreakers):
Working Conditions, Location and Hours
- Full-time office-based role, at our Gurugram office
- The working day is 8 hours, and the working week is 40 hours.
- 24/7 on-call availability for emergencies
- You will get 2 days weekly off (on-call availability for emergencies)
- 90-day notice period for resignation after the probation period
- Working exclusively for Extreme Infocom Pvt. Ltd. and not for any other companies
What We Offer
- 18 days per year's privilege/earned leave
- Family health insurance according to company policy
- A motivated team combining industry experts and emerging talent.
- Recognition programs - including Legend and Kudos Awards.
- Health & wellness programs and mental well-being support.
Subject to the internal policies of Megaport, which may be updated from time to time:

