368,657open jobs
9,442companies
50,883added this week
Browse all
Salary
$78k – $195k per year (Estimated)
Location
Remote (Switzerland)
Seniority
Senior
Employment
Full-Time
Overview
Company
Impact
Profile match
Jobgether is an AI-powered job platform focused on remote and flexible work. It matches candidates with relevant roles using skills and preference-based algorithms, and also offers career coaching and job-search guidance.

This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Senior Observability & Telemetry Engineer - Radian Arc based in Switzerland.

This is a high-impact engineering role focused on building the observability foundation for large-scale GPU cloud and edge infrastructure.

You will design and operate telemetry platforms that provide real-time visibility across distributed AI workloads, compute, storage, networking, and inference environments.

The role combines observability architecture, infrastructure telemetry, reliability engineering, and customer-facing performance insights.

You will work with high-cardinality metrics, logs, traces, and event data across both hyperscale environments and smaller edge deployments.

Your work will directly improve platform reliability, operational efficiency, performance, and incident response.

As a senior engineer, you will lead major initiatives, influence observability standards, and mentor engineers across multiple technical teams.

The environment is international, technically ambitious, and well suited to someone who enjoys solving complex distributed-systems challenges with significant autonomy.

Accountabilities

    • Design, implement, and operate scalable telemetry pipelines for metrics, logs, and traces across distributed GPU and edge infrastructure, supporting high-cardinality data from thousands of nodes and services.
    • Architect and maintain telemetry storage systems optimized for large-scale time-series and event data, while contributing to standards for instrumentation, logging, tracing, and SLO implementation.
    • Build comprehensive observability across compute, storage, networking, GPU clusters, inference workloads, and distributed training environments, identifying issues such as GPU throttling, network congestion, storage latency, and hardware degradation.
    • Develop dashboards, monitoring tools, and performance-analysis capabilities that provide internal teams and customers with actionable insights into workload health, GPU utilization, storage throughput, network latency, and inference performance.
    • Build and maintain network and infrastructure telemetry solutions using Python or Go, integrating data from technologies such as NVIDIA Cumulus Linux, VyOS, Citrix NetScaler/WAF, gNMI, SNMP, and streaming telemetry.
    • Develop advanced alerting, anomaly detection, reliability metrics, SLIs, and SLOs, integrating observability signals into operational workflows and incident management processes.
    • Collaborate with platform, networking, storage, compute, and operations teams to improve instrumentation, monitoring, incident response, and platform reliability.
    • Provide technical guidance and mentorship to engineers, promoting effective observability practices and consistent monitoring patterns across the organization.
    • Participate in on-call rotations supporting production observability and telemetry infrastructure.
    • Requirements

      • Proven experience operating observability systems and distributed infrastructure platforms at production scale, with strong expertise across metrics, logging, tracing, alerting, dashboards, and telemetry pipelines.
      • Strong programming skills in Go, Python, or Rust, with experience developing telemetry collectors, exporters, automation, or infrastructure tooling.
      • Hands-on experience with observability technologies such as Prometheus, OpenTelemetry, Grafana, distributed logging platforms, and large-scale telemetry databases such as ClickHouse or equivalent.
      • Experience working with large-scale GPU cloud, HPC, or AI infrastructure and monitoring distributed training or inference workloads.
      • Knowledge of GPU telemetry technologies such as NVIDIA DCGM, DCGM Exporter, NVML, GPU Operator telemetry, NVLink, and NVSwitch, as well as AI workload metrics including inference latency, throughput, NCCL health, synchronization latency, and storage I/O.
      • Strong understanding of networking and infrastructure telemetry, including experience with gNMI, SNMP, streaming telemetry, network flow telemetry, RDMA/RoCE, or comparable technologies.
      • Familiarity with cloud-native infrastructure, including Kubernetes, automation, CI/CD, and distributed systems.
      • Strong analytical and troubleshooting capabilities, with the ability to interpret complex telemetry signals, diagnose performance problems, identify systemic issues, and translate findings into actionable improvements.
      • Excellent collaboration and communication skills, with the ability to work effectively across infrastructure, networking, storage, compute, and operations teams.
      • A proactive, ownership-oriented mindset and the ability to lead complex observability initiatives in a fast-moving, technically sophisticated environment.
      • Benefits

        • Attractive compensation package reflecting your expertise, experience, transferable skills, and market conditions.
        • Permanent, full-time position with an EMEA-based remote work model.
        • Flexible and hybrid-friendly working environment designed to support international collaboration.
        • Opportunity to work on large-scale GPU, AI, cloud, networking, and edge infrastructure challenges.
        • Exposure to advanced observability, telemetry, reliability, and distributed-systems technologies.
        • Opportunity to contribute to major platform initiatives and influence observability standards across multiple engineering teams.
        • Mentorship and leadership opportunities, including the ability to guide engineers and promote best practices.
        • Career growth within a fast-growing international scale-up focused on innovative infrastructure solutions.
        • Inclusive and diverse working environment where qualified candidates are considered fairly and supported in their development.
Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
368,657 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
In your city
$89k – $212k per year (Estimated) • Remote/Hybrid • Full-Time • 6+ years exp • Rehovot
Bash
Python
DevOps
Amazon CloudWatch
Amazon EKS
AWS
CI/CD
Datadog
GitLab
GitLab CI
GitOps
Grafana
IAM
Kubernetes
PagerDuty
Pulumi
Terraform
Terragrunt
VictoriaMetrics
Apply
$40k – $86k per year (Estimated) • Remote/Hybrid • Full-Time • Élancourt
Bash
PowerShell
Python
SQL
Databases
MySQL
PostgreSQL
Redis
DevOps
Amazon CloudWatch
Ansible
AWS
Azure
CentOS Stream
CI/CD
Debian
Docker
GCP
GitLab
GitLab CI
Grafana
Hyper-V
IAM
Jenkins
Kubernetes
Nagios
Prometheus
Proxmox VE
Terraform
Ubuntu
VMWare
Windows Server
Zabbix
Apply
Data Engineer F/H 10 hours ago
$44k – $76k per year (Estimated) • Remote/Hybrid • Full-Time • Massy
SQL
DevOps
CI/CD
Management
Draw.io
Apply
$142k – $215k per year • In office • Full-Time • 12+ years exp • Princeton
JavaScript
TypeScript
Java
Java
Gradle
Hibernate
Maven
Spring Boot
Databases
Databricks
Snowflake
AI/ML
AI Agents
Claude
Claude Code
Frontend
Angular
React.js
Vue.js
DevOps
Amazon CloudWatch
Amazon ECS
Amazon EKS
Amazon EventBridge
Amazon S3
API Gateway
AWS
AWS Lambda
AWS Step Functions
Azure
CI/CD
IAM
Platform Engineering
Rest API
Kubernetes
Apply
$135k – $190k per year • In office • Full-Time • 12+ years exp • New York • Princeton
Databases
Apache Kafka
AI/ML
Hadoop
PyTorch
TensorFlow
DevOps
AWS
Azure
Azure DevOps
CI/CD
GCP
GitLab
GitLab CI
Jenkins
QA
Appium
Cypress
Playwright
Postman
Rest-Assured
Selenium
Apply
$152k – $229k per year • Remote • Full-Time
AI/ML
Human-in-the-Loop
Apply
$162k – $180k per year • Remote • Full-Time • 5+ years exp • Bachelor's Degree
Python
SQL
Databases
Snowflake
AI/ML
dbt
DevOps
AWS
Cybersecurity
HIPAA
Zero Trust
Analytics
ETL/ELT
Apply
$14k – $32k per year (Estimated) • Remote • Full-Time • 2+ years exp • Bachelor's Degree
DevOps
Incident Management
Management
ServiceNow
Apply
$29k – $60k per year (Estimated) • Remote • Full-Time • 8+ years exp
DevOps
Azure
Azure DevOps
Apply
$26k – $69k per year (Estimated) • Remote • Full-Time • 12+ years exp • Bachelor's Degree
Apply
See all jobs
This is one of many
368,657 more open roles from verified company boards, updated every day.