721,406open jobs
43,048companies
102,069added this week
Browse all
Salary
$125k – $135k per year
Location
Remote (United States)
Seniority
Senior · 5+ years exp
Employment
Full-Time

Confirmed on the employer's own hiring board on Sep 24, 2026. First seen by Alion on Sep 23, 2026. Skylo Technologies scores A on the Alion truth index.

Overview
Company
Impact
Profile match
Skylo is a non-terrestrial network (NTN) service provider that enables direct-to-device satellite connectivity for cellular devices, wearables, and IoT equipment. By leveraging standard 3GPP satellite protocols, the company connects existing cellular chipsets directly to satellite networks without requiring specialized hardware modifications. Its platform allows telecom operators, device manufacturers, and enterprise customers to deliver continuous cellular coverage and critical messaging services in remote, off-grid regions globally.

The world still has coverage blind spots. You could help eliminate them at Skylo.

Skylo has pioneered a standards-based approach to satellite connectivity. We connect smartphones and IoT devices directly to satellites. No special hardware, no entirely new networks. Just billions of existing devices, suddenly reachable anywhere on Earth. We're not building toward this future. We're already in it.

Our direct-to-device service is live on millions of activated devices across five continents, covering more than 72 million square kilometers, in partnership with leading satellite operators, mobile network operators, Tier-1 chipset makers, and OEMs worldwide. And we're just getting started.

At the heart of it all is Skylo's commercial NTN vRAN: a 3GPP standards-based, cloud-native platform that seamlessly bridges terrestrial and satellite networks. It's the infrastructure that makes true anywhere, anytime connectivity possible.

When you join Skylo, you'll work at the intersection of three markets reshaping how the world stays connected: mass-market consumer devices, automotive, and industrial IoT. Enabling people outdoors and critical workflows in the world's most remote places.

This is a rare chance to work on technology that matters, at a company that's already proving it works

ABOUT SKYLO

Skylo is a global Non-Terrestrial Network (NTN) service provider based in Mountain View, CA, offering a service that allows smartphone and IoT cellular devices to connect directly over existing satellites.

Skylo's direct-to-device service is live on millions of activated devices across five continents, with more than 60 million square kilometers of coverage, in partnership with multiple satellite operators, mobile network operators (MNOs), Tier-1 chipset makers, and OEMs. Devices connected over satellite are managed and served by Skylo's commercial NTN vRAN - a 3GPP standards-based, cloud-native base station and core. Skylo provides an anywhere, anytime connectivity solution that seamlessly roams between terrestrial and satellite networks. Our focus is on enabling connected services across three main verticals: mass-market consumer devices, automotive, and industrial IoT.

HOW YOU WILL IMPACT SKYLO

As a Senior, Cloud Infrastructure and Networking, in the Global Product Support & Customer Success organization, you are the Cloud Infrastructure domain authority within Skylo's production NTN network. Everything runs on the infrastructure you keep healthy - RAN NFs, Core NFs, OSS, BSS, and the observability pipeline itself. When a GKE node fails, when ArgoCD drifts, when a Persistent Volume Claim goes unavailable, when a PostgreSQL replica falls behind, when Prometheus WAL corrupts - you own the response.

You operate across Skylo's full hybrid cloud estate: GCP public cloud (GKE clusters, Pub/Sub pipelines, Cloud SQL) and on-premise private cloud infrastructure (bare-metal Kubernetes, hyperconverged compute, software-defined storage). You own 24x7 platform health, the observability pipeline (Prometheus, VictoriaMetrics, Grafana, OpenTelemetry), persistent storage operations (PostgreSQL, Redis), and the operational interface with Network Implementation for all GitOps-driven infrastructure changes.

KEY RESPONSIBILITIES

Cloud Infrastructure Operations & Health Ownership

  • Own 24x7 cloud infrastructure health across Skylo's hybrid production environment: GKE cluster node status, namespace and pod health, Persistent Volume Claim availability, network policies, and multi-cluster federation across Skylo's GCP footprint.

  • Own on-premise Kubernetes cluster health: bare-metal node availability, container runtime stability, CNI networking, persistent storage arrays (Ceph/Rook or equivalent), and hyperconverged compute platform operations (Harvester, KubeVirt, or KVM).

  • Monitor and triage infrastructure alarms using OSS dashboards, Grafana/VictoriaMetrics telemetry, GCP Cloud Monitoring, and Loki log correlation - distinguish transient platform events from systemic infrastructure degradation.

  • Execute and own Cloud Infra runbooks for P2-P4 fault categories: GKE node recovery, pod eviction and rescheduling, PVC repair, database failover execution, Prometheus WAL corruption recovery, ArgoCD drift remediation, and certificate rotation - without requiring engineering involvement for covered fault classes.

  • Own BSS-IIS GKE cluster monitoring and infrastructure health; maintain runbooks that reflect current cluster topology after every infrastructure change.

Observability Pipeline & Data Platform Operations

  • Own the observability pipeline end-to-end: Prometheus scrape target integrity, VictoriaMetrics retention and query performance, Grafana dashboard coverage and accuracy, OpenTelemetry collector health, and alert routing via Pub/Sub to the OSS.

  • Maintain database reliability: PostgreSQL streaming replication health, backup and restore procedures, failover testing, query performance monitoring; Redis cluster operations, eviction policy management, and persistence configuration.

  • Ensure log aggregation pipeline health (Loki or ELK): ingestion rates, retention policies, query performance, and completeness - the observability stack must be operational before the network events it monitors can be triaged.

  • Partner with NI (Network Implementation & Infrastructure) on all planned infrastructure changes: receive advance notice, validate post-deployment observability, and sign off on operational readiness before the change window closes.

Cloud Infrastructure Incident Diagnosis & Escalation Authority

  • Serve as the L3 escalation authority for all Cloud Infra incidents: take ownership from the Incident Manager, diagnose at the Kubernetes, storage, network, and database layer using kubectl, GCP console, node logs, and infrastructure telemetry, and deliver a resolution or a decision-grade root cause.

  • Lead Cloud Infra troubleshooting bridges: command the technical investigation for GKE node failures, cluster upgrade failures, storage outages, PubSub pipeline disruptions, database failover events, and ArgoCD sync failures - drive to resolution or clear engineering handoff.

  • Diagnose and resolve infrastructure failure modes: node NotReady conditions, pod CrashLoopBackOff chains, PVC mount failures, CSI driver errors, network policy misconfigurations, Helm release drift, etcd latency spikes, and cross-cluster federation breaks.

  • Participate in the global 24x7 on-call rotation as the Cloud Infra domain escalation tier - reachable within defined SLA windows for Sev 1 events; function as the technical decision-maker, not the first responder.

SLO Engineering & Reliability

  • Define and maintain SLOs for all Cloud Infra components: GKE control plane availability, database query latency, storage IOPS, message pipeline throughput, and observability stack uptime - tied directly to network SLA commitments to MNO partners.

  • Own error budget tracking and the process for trading error budget against deployment velocity; escalate when error budget burn rate requires engineering intervention or deployment freezes.

  • Drive toil reduction: identify and eliminate manual Cloud Infra procedures; own the roadmap to automated cluster recovery, rolling restarts, storage repair, and certificate rotation in partnership with Ops Platform Engineering.

  • Lead capacity planning for compute, storage, and network resources across public and private cloud - forecast growth based on subscriber projections and new MNO partner onboarding.

Root Cause Analysis & Post-Incident Ownership

  • Own Cloud Infra RCA end-to-end: lead the investigation, document the complete causal chain from infrastructure trigger through upstream NF impact, and deliver systemic action items with owners, timelines, and measurable success criteria.

  • Deliver Initial RCA documentation within defined SLA windows; identify systemic infrastructure failure patterns - cluster upgrade regressions, storage controller bugs, network policy drift, resource exhaustion trends - and translate them into engineering requirements.

  • Contribute to the weekly and monthly Network Performance Report: infrastructure availability, database latency trends, storage IOPS, observability pipeline health, and SLA deviation analysis.

Runbook Authorship & Operational Standards

  • Author, own, and maintain all Cloud Infra runbooks and SOPs: GKE node recovery, database failover, storage expansion, Prometheus WAL repair, ArgoCD rollback, certificate rotation, and cluster upgrade procedures - every procedure tested before production reliance.

  • Define the diagnostic decision tree for each known infrastructure fault class: entry condition, triage steps, isolation method, resolution action, and escalation criteria - written at the level where a Senior NRE can execute independently.

  • Validate and sign off on operational readiness for all infrastructure changes: DCI/DCE build-outs, GKE cluster expansions, Kubernetes version upgrades, and new on-premise hardware deployments.

Cross-Functional Collaboration & Team Development

  • Partner with OPE as the Cloud Infra domain's primary automation consumer: define Kubernetes event schemas, alert-to-action contracts, and closed-loop policy requirements for infrastructure auto-remediation.

  • Represent Cloud Infra in NI architecture reviews: define observability and operational readiness requirements for all infrastructure expansions and GitOps pipeline changes before go-live.

  • Collaborate with Core NRE and RAN NRE on infrastructure-layer issues affecting network functions: pod scheduling, PVC availability, network policy changes, and platform upgrade impacts on NF workloads.

  • Partner with security teams on infrastructure hardening: patch compliance, RBAC policies, network segmentation, container image scanning, and runtime security monitoring across public and private cloud.

  • Surface toil and automation opportunities to the Service Assurance & Automation team - document the procedure, frequency, and MTTR cost as structured input to the automation backlog.

  • Mentor Senior NREs in Cloud Infra domain depth: Kubernetes troubleshooting patterns, storage operations, database reliability, observability pipeline internals, and escalation judgment.

GitOps, IaC & Platform Engineering Interface

  • Own operational oversight of GitOps tooling in production: ArgoCD sync health, Helm chart version management, drift detection, and rollback execution for multi-cluster deployments.

  • Review and validate Infrastructure as Code (Terraform, Ansible) changes that impact production - ensure operational impact is assessed and observability is in place before merge.

  • Engage Skylo's Platform Engineering and NI teams with full operational context when issues exceed operational resolution authority - deliver a structured problem statement, infrastructure telemetry bundle, and a clear question rather than a vague escalation.

REQUIRED QUALIFICATIONS

  • 5+ years of infrastructure engineering, Site Reliability Engineering, or cloud operations in a production 24x7 environment - with direct on-call ownership for Kubernetes-at-scale environments.

  • Deep Kubernetes expertise: multi-cluster operations (GKE or EKS), node pool management, RBAC, network policies, persistent storage (PVC, CSI drivers), CRD/operator patterns, and production cluster upgrade procedures.

  • Hybrid cloud operations: hands-on experience operating both public cloud (GCP or AWS) and on-premise/private cloud infrastructure (bare-metal Kubernetes, KVM, or hyperconverged platforms).

  • Production observability stack ownership: Prometheus (federation, remote write, WAL management), Grafana, VictoriaMetrics, OpenTelemetry, and alerting pipeline design with Pub/Sub or equivalent.

  • Database reliability: PostgreSQL streaming replication, backup/restore, failover procedures, and performance tuning; Redis cluster operations and persistence management.

  • GitOps tooling in production: ArgoCD or Flux CD for multi-cluster operations; Helm chart authorship and version management; Terraform or Ansible for infrastructure provisioning.

  • SRE fundamentals: SLO/SLI/SLA definition, error budget management, toil measurement, capacity planning, and on-call rotation design.

  • Container and Linux internals: container runtime debugging, kernel-level performance analysis, storage subsystem troubleshooting, and network packet flow understanding.

  • Runbook authorship: ability to write infrastructure diagnostic procedures at the level where a less-experienced engineer can execute them independently under incident pressure.

  • Strong written and verbal communication: capable of delivering RCA documents, engineering escalations with structured problem statements, and MNO-facing infrastructure summaries.

PREFERRED QUALIFICATIONS

  • Experience operating cloud infrastructure for telecom or NTN workloads: 5G Core NF hosting, vRAN compute requirements, or satellite ground segment infrastructure.

  • Software-defined storage expertise: Ceph, Rook, or equivalent distributed storage systems at production scale.

  • Private cloud platform experience: KubeVirt, Harvester, or OpenStack for VM-container convergence on bare-metal infrastructure.

  • Networking depth: BGP routing, VXLAN overlays, EVPN fabrics, software-defined networking, and hardware load balancer operations.

  • Strong development background in Go or Python for building custom automation tooling, Kubernetes operators, or infrastructure lifecycle integrations.

  • FinOps experience: cloud cost optimization, resource lifecycle automation, and capacity right-sizing across public cloud footprints.

  • Certifications: CKA, CKS, AWS Solutions Architect Professional, or Red Hat Certified Architect.

WHAT WE OFFER

With employees working across three continents, Skylo is proud to be an equal opportunity employer dedicated to building an inclusive and diverse workforce. Our worldwide culture encourages a flexible approach to work, and we offer an attractive range of benefits:

  • Competitive compensation packages including a stock option-based equity program

  • Comprehensive benefits including medical, dental, vision, and retirement plan

  • Monthly allowances for wellness and education reimbursement

  • A generous time-off policy, holidays, and the opportunity to temporarily work abroad

  • A once-in-a-career opportunity to operate the world's first commercial, live direct-to-device satellite network

  • Access to a world-class team across software, hardware, chipsets, telecom, satellite, and network virtualization

  • Open, transparent, inclusive culture that blends Silicon Valley, Nordic, and South Asia characteristics

The estimated base salary range for this role is $125,000 - $135,000. Within the range, individual pay is determined by work location and additional factors, including job-related skills, experience, and relevant training.

In addition to base salary, Skylo offers a competitive benefits package including health insurance, retirement plans, paid time off, and equity options.

EEO STATEMENT

Skylo is an equal-opportunity employer and we celebrate diversity. We do not discriminate based on race, religion, color, ancestry, national origin, caste, sex, sexual orientation, parent or caregiver status, political affiliation, gender, gender identity or expression, age, disability, medical condition, pregnancy, genetic makeup, marital status, or military service consistent with applicable federal, state, and local laws.

We are also committed to providing reasonable accommodations for qualified individuals with disabilities and disabled veterans in our job application procedures. Please let us know if you need assistance or accommodation due to a disability.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
721,406 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
In your city
In office • Bachelor's Degree
Python
SQL
Databases
PostgreSQL
Oracle
Analytics
Tableau
Power BI
Informatica
Apply
$30k – $35k per year • Remote/Hybrid • Bachelor's Degree • Milan
Python
JavaScript
TypeScript
SQL
Python
SQLAlchemy
FastAPI
Databases
Redis
AI/ML
Fine-tuning
LLM
Red Teaming
Machine Learning
Frontend
Tailwind CSS
React.js
DevOps
Rest API
CI/CD
Git
Cybersecurity
OWASP
Apply
$77k – $169k per year (Estimated) • In office • Full-Time • 2+ years exp • Bachelor's Degree • Overland Park
SQL
DevOps
Windows Server
Apply
$184k – $288k per year • In office • Full-Time • 8+ years exp • Bachelor's Degree • Santa Clara
Python
DevOps
Terraform
GitOps
ArgoCD
Kubernetes
Linux
Apply
$135k – $325k per year (Estimated) • In office • Full-Time • 2+ years exp • Tel Aviv
Python
SQL
AI/ML
Model Context Protocol
Function Calling
AI Agents
LLM
RAG
InfiniBand
Human-in-the-Loop
Context Engineering
LLM Guardrails
Agentic Workflows
Multi-Agent Systems
Tool Use
DevOps
SLURM
CI/CD
Git
Docker
Kubernetes
HPC
Linux
Apply
$135k – $145k per year • Remote/Hybrid • Full-Time • Mountain View
Python
Ruby
Bash
DevOps
gRPC
Jenkins
Linux
Unix
Management
Confluence
QA
Postman
Robot Framework
Apply
Senior IT Engineer 3 days ago
$115k – $125k per year • In office • Full-Time • 5+ years exp • Bachelor's Degree • Mountain View
DevOps
GCP
AWS
GitHub
IAM
Linux
Windows
DNS
DHCP
VPN
Wi-Fi
Cybersecurity
Okta
ISO 27001
SOC 2
Zero Trust
Management
Slack
Confluence
Jira
Google Workspace
ITIL
Service Desk
Apply
Remote/Hybrid • Full-Time • Mountain View
Marketing
LinkedIn
Apply
$160k – $175k per year • In office • Full-Time • 6+ years exp • New York
AI/ML
Claude
ChatGPT
Midjourney
Perplexity
Design
Canva
Apply
$148k – $306k per year (Estimated) • Equity • Remote/Hybrid • Full-Time • 5+ years exp • Mountain View
Apply
See all jobs
This is one of many
721,406 more open roles from verified company boards, updated every day.