{"id":1625561,"url":"https://alion.io/job/clearcaptions-sr-technology-operations-reliability-manager-100-remote","title":"Sr. Technology Operations & Reliability Manager (100% Remote)","company":{"id":2107958,"name":"ClearCaptions","domain":"clearcaptions.com","url":"https://alion.io/company/clearcaptions","size_band":"501-1000","is_staffing_agency":false,"employer_type":"direct","is_intermediary":false,"listed_via":null,"ats_vendor":"Dayforce","truth_index":null},"role":"Operations","role_family":"Operations","seniority":"senior","employment_type":null,"work_mode":"remote","remote_scope":"stated_countries","remote_scope_basis":"inferred_company_offices","remote_working_hours":null,"hiring_geo_confidence":"inferred","locations":[],"countries":[],"hiring_countries":["US"],"hiring_countries_total":1,"salary":null,"salary_estimate":{"min_usd":82000,"max_usd":171000,"period":"year","method":"global_role_seniority_cell","sample_n":2298},"experience_years_min":8,"visa_sponsorship":false,"relocation_package":false,"has_equity":false,"technologies":[{"name":"AWS","optional":false},{"name":"CI/CD","optional":false},{"name":"Incident Management","optional":false},{"name":"SLI/SLO/SLA","optional":false},{"name":"CloudFormation","optional":true},{"name":"Datadog","optional":true},{"name":"GitHub Actions","optional":true},{"name":"GitLab CI","optional":true},{"name":"Grafana","optional":true},{"name":"Jenkins","optional":true},{"name":"Kubernetes","optional":true},{"name":"Linux","optional":true},{"name":"Microsoft Office","optional":true},{"name":"Microsoft Teams","optional":true},{"name":"New Relic","optional":true},{"name":"Prometheus","optional":true},{"name":"Slack","optional":true},{"name":"Terraform","optional":true}],"status":"live","first_seen_at":"2026-06-25T08:00:00Z","employer_posted_date":"2026-10-01","last_verified_at":"2026-10-02T09:58:21Z","board_verified":true,"closed_at":null,"days_open":99,"trust":{"level":"stale","repost_count":0,"flags":["stale"],"days_open":98},"description":"Summary:\n\nThe Sr. Technology Operations & Reliability Manager is responsible for leading and managing the DevOps and 24/7 Site Reliability Engineering (SRE) teams to ensure the performance, reliability, and scalability of the ClearCaptions call platform. The role requires a strategic and hands-on approach to optimizing cloud infrastructure, implementing automation, and enforcing best practices to maintain high availability of enterprise applications and databases. The Sr. Technology Operations & Reliability Manager collaborates with stakeholders across departments to ensure a resilient, secure, and highly efficient technology environment.\n\nThis role requires practical experience supporting highly available telecommunications, VoIP, or real-time communications platforms. Responsibilities include understanding the operational impact of SIP signaling, media quality, carrier interconnects, routing, SBCs, and telecom-related incident response, while leading teams responsible for platform reliability, observability, escalation management, and continuous improvement.\n\nWhat This Role Does:\nLeads and manages the DevOps and 24/7 SRE teams responsible for CI/CD pipelines, cloud infrastructure, observability, and automation.\nDevelops and enforces best practices for site reliability, incident management, and infrastructure optimization.\nEnsures high availability, performance, and scalability of the call platform while minimizing service disruptions.\nLeads operational reliability for ClearCaptions’ real-time communications platform, including VoIP, PSTN-connected services, SIP-based call flows, carrier interconnects, routing, and media quality.\nPartners with telecom engineering, carriers, vendors, SRE, Security, Product, and Engineering teams to improve resiliency, reduce incident frequency, and ensure rapid recovery from service-impacting events.\nOversees telecom-related incident response, escalation, RCA, and corrective/preventive action plans, with measurable improvements in availability, MTTR, call quality, and customer impact reduction.\nEnsures operational readiness for telecom platform changes, including runbooks, monitoring, alerting, change validation, failover testing, and disaster recovery planning.\nDrives observability and SLO development for call-platform services, including signaling, media, carrier connectivity, routing performance, and service health.\nEstablishes monitoring and alerting strategies using industry-leading observability tools to proactively detect and resolve issues.\nOversees infrastructure management, configuration, scaling, and deployment of enterprise server-based computing systems in AWS.\nImplements and manages infrastructure as code (IaC) solutions to enhance automation and cloud efficiency.\nDrives cloud-native architecture improvements to support business growth and operational excellence.\nEnsures compliance with security policies, internal controls, and cloud governance frameworks.\nRecruits, mentors, and develops DevOps and SRE team members, fostering a culture of continuous learning and innovation.\nCollaborates with cross-functional teams, including Security, Product, and Engineering, to align technical strategy with business objectives.\nLeads tactical responses to cybersecurity incidents and implements preventative security measures.\nDevelops strategic initiatives to improve operational efficiency, system resilience, and cost optimization.\nPerforms other duties as assigned.\n\nWhat You Will Bring:\nBachelor’s degree in computer science, engineering, or a related field. Equivalent combinations of education and relevant experience will be considered.\nA minimum of eight (8) years of experience in IT operations, DevOps, SRE, or telecommunications operations, including leadership responsibility for mission-critical cloud, VoIP, UCaaS, contact center, or real-time communications platforms.\nA minimum of three (3) years of experience leading people or teams, including accountability for performance, development, and delivery of results aligned with organizational objectives.\nTelecom, VoIP, or real-time communications platform experience strongly preferred; experience with SIP, SBCs, carrier interconnects, and VoIP/PSTN interoperability is highly desirable.\nWorking knowledge of SIP, RTP, SBCs, carrier interconnects, call routing, VoIP/PSTN interoperability, and telecom incident troubleshooting.\nExperience leading operational response for mission-critical communications platforms, including major incidents, carrier/vendor escalations, RCAs, and service reliability improvement plans.\nFamiliarity with telecom observability practices, including dashboards, alerting, SLOs, runbooks, packet/call-flow analysis, and quality/performance baselining.\nAWS Certified Solutions Architect – Associate certification required; higher-level AWS certifications preferred.\nExpertise in cloud computing, Kubernetes, and infrastructure automation tools (Terraform, CloudFormation, or similar).\nExperience managing 24/7 mission-critical applications and implementing SRE principles.\nProficiency in observability and monitoring tools such as Prometheus, Grafana, Datadog, or New Relic.\nHands-on experience with CI/CD tools such as Jenkins, GitHub Actions, or GitLab CI/CD.\nStrong knowledge of networking, Linux administration, and troubleshooting complex system performance issues.\nExperience optimizing enterprise databases and cloud-based storage solutions.\nDemonstrated ability to lead through influence, including mentoring others, leading initiatives or workstreams, contributing to best practices, and driving cross-functional outcomes.\nStrong analytical, decision-making, and problem-solving skills.\nExcellent leadership, communication, and stakeholder management skills.\nAbility to thrive in a fast-paced, high-growth environment while ensuring operational stability.\nAbility to work collaboratively with colleagues and staff to create a high-quality, results-driven, team-oriented environment.\nDemonstrated ability to use discretion, make sound decisions, and maintain confidentiality.\nWillingness to work flexible hours, participate in on-call rotations when necessary, and travel up to 10%, which may include occasional overnight travel.\nProficiency in Microsoft Office Suite and modern communication tools for virtual teams (e.g., Microsoft Teams, Slack).\n\nPhysical Demands:\n\nIn accordance with the Americans with Disabilities Act (ADA) and applicable state and local laws, the Company will provide reasonable accommodations to qualified individuals with documented disabilities to enable them to perform the essential functions of the job, unless such accommodations would impose an undue hardship. Employees seeking accommodation should contact the People Department to initiate the interactive process.\n\nEmployees may experience the following physical demands for extended periods of time: \nSitting, standing, and walking (90-100%).\nKeyboarding (95-100%).\nViewing computer monitor, tablet, and cell phone requiring close vision (70-90%).\nMay require occasional lifting and racking of equipment (up to 50 lbs.) in data center environments.\n\nWork Environment:\n\n100% Remote with Travel: Work environment is primarily indoors (home office, customer or vendor sites, or other business meeting venues); travel may involve exposure to varying weather and temperature conditions, as well as driving and traffic hazards. Travel is required, approximately 10%, and may include overnight and out-of-state trips.","description_format":"text","description_chars":7491,"description_truncated":false,"requirements":{"experience_years_min":8,"management_years_min":null,"team_size_min":null,"manages_managers":false,"education":{"level":"bachelor","optional":false},"security_clearance":false,"languages":[]},"benefits":["Continuous learning","Home office"],"hiring_locations":[{"name":"United States","iso":"US","kind":"country"}],"hiring_excludes":[],"relocation_offered":false,"industries":["Health Care","Hearing Care & Audiology"],"lifecycle":[{"event":"open","at":"2026-10-01T20:57:47Z"}],"liveness":{"score":14,"band":"cold","label":"Long shot","p_open":1,"p_active":0.486,"p_room":0.28,"age_days":98,"expected_fill_days":40,"reasons":["conf:8","win:tail","crowd:"],"computed_at":"2026-10-02T05:45:00Z"},"pay":null,"html_url":"https://alion.io/job/clearcaptions-sr-technology-operations-reliability-manager-100-remote","json_url":"https://alion.io/job/clearcaptions-sr-technology-operations-reliability-manager-100-remote.json","meta":{"generated_at":"2026-10-03T01:40:47Z","cache_seconds":300,"methodology":"https://alion.io/methodology","terms":"https://alion.io/terms","contact":"https://alion.io/contact","api":"https://alion.io/developers","usage":{"tier":"crawler","counted_by":"address","units_charged":1,"used_today":1659,"day_limit":5000,"remaining_today":3341,"minute_limit":60,"resets_at":"2026-10-04T00:00:00Z"}}}