{"id":649708,"url":"https://alion.io/job/nscale-staff-observability-platform-engineer","title":"Staff Observability Platform Engineer","company":{"id":48155,"name":"Nscale","domain":"nscale.com","url":"https://alion.io/company/nscale","size_band":"501-1000","is_staffing_agency":false,"employer_type":"direct","is_intermediary":false,"listed_via":null,"ats_vendor":"Greenhouse","truth_index":{"grade":"A","score":92,"open_postings":41,"ghost_share":0,"stale_share":0.122,"repost_share":0,"time_to_fill_p50_days":63,"computed_at":"2026-09-28T05:45:00Z"}},"role":"DevOps","role_family":"DevOps","seniority":"staff","employment_type":null,"work_mode":"on_site","remote_scope":null,"remote_scope_basis":null,"remote_working_hours":null,"hiring_geo_confidence":"structured","locations":[],"countries":[],"hiring_countries":[],"hiring_countries_total":0,"salary":{"min":190000,"max":260000,"currency":"USD","period":"year","gross":null,"usd_annual":260000},"salary_estimate":null,"experience_years_min":6,"visa_sponsorship":false,"relocation_package":false,"has_equity":false,"technologies":[{"name":"Ansible","optional":false},{"name":"Apache Kafka","optional":false},{"name":"ClickHouse","optional":false},{"name":"Fluent Bit","optional":false},{"name":"Go","optional":false},{"name":"Grafana","optional":false},{"name":"HPC","optional":false},{"name":"Kubernetes","optional":false},{"name":"Loki","optional":false},{"name":"OpenTelemetry","optional":false},{"name":"Platform Engineering","optional":false},{"name":"Prometheus","optional":false},{"name":"Python","optional":false},{"name":"SLURM","optional":false},{"name":"Terraform","optional":false},{"name":"Thanos","optional":false},{"name":"Vector","optional":false},{"name":"VictoriaMetrics","optional":false}],"status":"live","first_seen_at":"2026-06-09T13:52:45Z","employer_posted_date":"2026-09-21","last_verified_at":"2026-09-28T22:39:29Z","board_verified":true,"closed_at":null,"days_open":111,"trust":{"level":"ok","repost_count":0,"flags":[],"days_open":110},"description":"About Nscale\nNscale is the GPU cloud engineered for AI. We provide cost-effective, high-performance infrastructure for AI start-ups and large enterprise customers. Nscale simplifies AI development while enabling superior results, supporting strategic business outcomes such as cost management, rapid innovation, and environmental responsibility.\nWe thrive on a culture of relentless innovation, ownership, and accountability, where every team member takes pride in their work and drives it with excellence and urgency. As an Nscaler, you'll build trust through openness and transparency while contributing to the technology that powers the future.\nAbout the Role\nAs a Staff Observability Platform Engineer, you'll play a critical role in building and evolving Nscale's observability platform, enabling deep visibility into GPU clusters, AI workloads, and the infrastructure that powers them.\nYou view observability as a product, not simply a collection of tools. You'll help define and implement scalable, reliable observability solutions that empower engineering teams to understand system behavior, diagnose issues quickly, and operate complex distributed systems with confidence.\nYou'll combine technical leadership with hands-on engineering, partnering across SRE, infrastructure, platform, and AI/ML teams to improve reliability, operational efficiency, and developer experience. You'll influence architectural decisions, establish engineering best practices, and help drive the evolution of observability capabilities across the organization.\nThis is a role for someone who enjoys solving difficult infrastructure problems, building platforms that scale, and helping engineering teams succeed through better visibility and operational insight.\nWhat You'll Do\nDesign, build, and evolve observability platforms across metrics, logs, traces, alerting, and telemetry pipelines.\n\nLead the implementation of scalable observability solutions that support Nscale's growing GPU and AI infrastructure.\n\nPartner with SRE, infrastructure, platform, and AI/ML teams to ensure observability is embedded throughout the software and infrastructure lifecycle.\n\nDrive improvements in monitoring coverage, alert quality, service health visibility, and incident response effectiveness.\n\nDevelop standards, frameworks, and reusable patterns that simplify observability adoption across engineering teams.\n\nIdentify reliability risks and operational blind spots, helping teams proactively address them before they impact customers.\n\nContribute to architectural decisions around telemetry collection, storage, retention, cardinality management, and performance optimization.\n\nLead technical initiatives and projects that improve platform scalability, reliability, and operational efficiency.\n\nMentor engineers and provide technical guidance through design reviews, code reviews, and knowledge sharing.\n\nParticipate in incident investigations and postmortems, translating operational learnings into durable platform improvements.\n\nEvaluate new observability technologies and practices, balancing innovation with operational simplicity and long-term maintainability.\n\nAbout You\n6+ years of experience in SRE, platform engineering, infrastructure engineering, observability engineering, or related disciplines.\n\nStrong experience building and operating observability platforms in cloud-native, distributed environments.\n\nDeep hands-on experience with several of the following technologies: Prometheus, Thanos, VictoriaMetrics, Grafana, Loki, Tempo, OpenTelemetry, ClickHouse, Elastic, or similar platforms.\n\nStrong software engineering skills with proficiency in Go, Python, or equivalent languages.\n\nExperience operating and troubleshooting Kubernetes-based platforms at scale.\n\nStrong understanding of monitoring, logging, tracing, telemetry pipelines, and modern observability practices.\n\nExperience designing systems with scalability, reliability, performance, and operational simplicity in mind.\n\nProficiency with Infrastructure-as-Code tools such as Terraform, Ansible, or equivalent.\n\nAbility to lead technical initiatives and influence engineering decisions across multiple teams.\n\nExcellent communication skills with the ability to explain technical tradeoffs and align stakeholders around pragmatic solutions.\n\nPreferred\nExperience operating observability systems in GPU, AI/ML, HPC, or large-scale compute environments.\n\nFamiliarity with Slurm, Kubernetes GPU scheduling, or AI infrastructure platforms.\n\nExperience with high-volume telemetry pipelines and streaming technologies such as Kafka, Vector, or Fluent Bit.\n\nKnowledge of observability challenges related to model training, inference workloads, GPU utilization, and distributed AI systems.\n\nExperience mentoring engineers and helping grow technical capability across teams.\n\nEqual Opportunities Statement\nWe strongly encourage applications from people of color, the LGBTQ+ community, people with disabilities, neurodivergent individuals, parents, carers, and people from lower socio-economic backgrounds.\nIf there's anything we can do to accommodate your specific situation, please let us know.\nNote: Responsibilities outlined are not exhaustive and may evolve as business needs change.\nThe range below reflects the base salary for the position. Actual compensation may vary based on job-related factors such as skill set, experience, education, and location. In addition to base salary, this role may be eligible for bonus, equity, and/or commission programs. Nscale may offer a competitive benefits package including medical, dental, vision, flexible paid time off, parental leave, and retirement plan participation.\nSalary Range\n$190,000—$260,000 USD\nFor information on how Nscale handles candidate personal data, please see our Employee & Candidate Privacy Notice: Here.\nNscale does not accept unsolicited candidate submissions from recruitment agencies.","description_format":"text","description_chars":5908,"description_truncated":false,"requirements":{"experience_years_min":6,"management_years_min":null,"team_size_min":null,"manages_managers":false,"education":null,"security_clearance":false,"languages":[]},"benefits":["Equity","Parental leave","Retirement plans"],"hiring_locations":[],"hiring_excludes":[],"relocation_offered":false,"industries":["Data Centers & Colocation","AI Compute & Inference"],"lifecycle":[{"event":"open","at":"2026-09-10T20:22:42Z"}],"liveness":{"score":33,"band":"fade","label":"Fading","p_open":1,"p_active":0.903,"p_room":0.36,"age_days":110,"expected_fill_days":63,"reasons":["conf:3","velocity","win:tail","crowd:brand"],"computed_at":"2026-09-28T05:45:00Z"},"pay":{"stated_usd_annual":260000,"is_top_pay":true},"html_url":"https://alion.io/job/nscale-staff-observability-platform-engineer","json_url":"https://alion.io/job/nscale-staff-observability-platform-engineer.json","meta":{"generated_at":"2026-09-29T03:01:09Z","cache_seconds":300,"methodology":"https://alion.io/methodology","terms":"https://alion.io/terms","contact":"https://alion.io/contact","api":"https://alion.io/developers","usage":{"tier":"crawler","counted_by":"address","units_charged":1,"used_today":2682,"day_limit":5000,"remaining_today":2318,"minute_limit":60,"resets_at":"2026-09-30T00:00:00Z"}}}