368,910open jobs
9,449companies
47,822added this week
Browse all
Salary
$130k – $150k per year
Location
In office (Bellevue)
Seniority
Senior · 7+ years exp
Employment
Full-Time
Overview
Company
Impact
Profile match
Viome was founded in 2016 with a mission to make illness optional by predicting and preventing chronic diseases through a deeper understanding of an individual's biology at a molecular level. Viome is the industry's only direct-to-consumer healthc...

At Viome, we are driven by a singular mission: to help people live a healthy, disease-free life. This mission guides our actions, fuels our passion, and shapes the impact we aim to have in the world. Our core values - Be Bold, Be Collaborative, Be Frugal, and Grow Continuously - underpin our approach to achieving this goal. If you are motivated by the idea of working in an environment that prioritizes bold innovation, teamwork, efficient resource use, and continuous learning, all towards promoting health and preventing disease, we warmly invite you to apply. Join us in our journey to transform lives and create a healthier future for all.

We are looking for a Senior Platform Operations Engineer, Infrastructure to take ownership of Viome’s Azure-based service platform with a clear mandate: make it dramatically simpler through robust Unix-based systems engineering.

This is a consolidation role for a seasoned systems administrator. The work involves migrating a suite of abstracted services onto a deliberately stable target platform: Linux VMs, systemd-supervised services, Apache, HAProxy, and Nginx routing. You must possess strong Unix knowledge to operate and eventually transform the current stack safely. We are looking for a traditional operations background where success is measured by the removal of unnecessary layers and a genuine bias toward simplicity and host-level stability.

Responsibilities

    Simplification & Migration (core mandate)

  • Plan and execute zero-downtime migrations of services from AKS/containers to VM-based hosting: systemd unit authoring and service supervision (restart policy, resource limits, sandboxing), Apache, HAProxy and/or nginx as reverse proxy and TLS terminator, certificate automation.

  • Build and own the deployment scripts for the target platform: build → test → archive → ship → symlink-flip → health-check → rollback, scripted in bash.

  • Preserve two non-negotiable invariants while simplifying everything else: immutable, commit-traceable artifacts and scripted rollback.

  • Inventory and retire stale infrastructure: dormant deployments, unused DNS records, orphaned firewall rules, unpinned image tags.

  • Maintain host hygiene for consolidated services: OS patching discipline, runtime vendoring, log rotation, centralized log aggregation.

  • Familiarity with UptimeKuma, Nagios or similar.

  • Operate and migrate data stores: PostgreSQL and/or MySQL.

  • Network and Infrastructure Operations

  • Operate and troubleshoot workloads during the transition, focusing on the networking layer, load balancing, and core platform services.

  • Administer the hub network: Azure Firewall rules, VPN gateways, VNet peering, public and private DNS zones and reason about a packet’s full path from public IP to service.

  • Support the existing release process and network and infrastructure operations until each service is migrated to the new VM-based standard.

  • Keep the observability stack healthy (OpenTelemetry, ELK, Grafana, uptime and cost monitoring) and carry its essentials forward to the simplified platform.

  • Coordinate cross-cloud dependencies with AWS: DNS/edge routing, queue consumers, and egress IP allowlists.

  • L2+ operations support; manage runbooks for external L0, L1 support.

  • External Integrations & Security

  • Own the external integrations most at risk during migration: e-commerce and subscription platforms, messaging/notification providers, and clinical/health-data partners - webhook delivery, signature verification, idempotency, and retry semantics.

  • Raise the security baseline as you consolidate: secrets management, webhook authentication, least-privilege network access, and data-retention hygiene. Findings from an internal review are ready for you to remediate; the instinct to spot and close this class of issue - and not create more - is part of the job.

Qualifications

    Required

  • 7+ years operating production Unix/Linux systems, with deep knowledge of systemd, process supervision; Apache, HAProxy

  • Strong shell plus one scripting language (bash/Python/PHP) with a track record of building deploy and rollback tooling, not just using it.

  • Demonstrated reverse-engineering ability: taking ownership of an undocumented production service and recovering its real dependencies and failure modes.

  • Message-queue literacy: Azure Service Bus, SQS, RabbitMQ, or equivalent - ordering, lock/ack semantics, dead-letter handling.

  • PostgreSQL operations: replicas/clones, connection proxying, network-restricted access, production diagnostics.

  • Working proficiency with Kubernetes and cloud networking - enough to operate private AKS clusters, an Istio-style ingress layer, and Azure firewall/DNS during the transition. We will onboard you on our specifics; deep specialization is not required.

  • Migration experience: consolidating or re-platforming production services with zero-downtime cutover and tested rollback.

  • A demonstrable record of reducing system surface area - services consolidated, infrastructure retired.

  • Clear written communication for runbooks, migration plans, and cross-functional coordination.

  • Strongly Preferred

  • Azure networking at landing-zone depth: hub-spoke VNets, Azure Firewall, Private Link/Private DNS, VPN gateways.

  • Administration of external-dns, cert-manager, and policy engines such as Kyverno.

  • ELK and OpenTelemetry pipeline operations.

  • Webhook-heavy integration experience (e-commerce platforms such as Shopify, subscription billing, marketing/notification platforms, or healthcare data exchanges).

  • Experience in a regulated or health-data environment: PHI handling, audit trails, least-privilege network design.

  • Terraform or equivalent IaC for cloud network and compute resources.

  • Enough AWS to manage cross-cloud seams (Route53, SQS, egress allowlists).

  • AI minded; proficient with Claude or similar.

What Success Looks Like

  • First 30 days: Trace and fix a production issue end-to-end (DNS → firewall → ingress → service → database) with guidance; produce a true-dependency inventory for one candidate service.

  • First 90 days: First service migrated off AKS to the VM/systemd platform with a passing rollback drill; firewall rules, DNS zones, and cluster add-ons documented and reproducible; stale-resource inventory complete and retirement underway.

  • First 6 months: Migration cadence established with multiple services consolidated; release and rollback drills routine on both platforms; measurable reduction in infrastructure footprint and spend; security-baseline remediations closed.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
368,910 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
Bellevue
$51k – $110k per year (Estimated) • In office • Full-Time • 12+ years exp • Bachelor's Degree • Bengaluru • Chennai • Hyderabad
Java
Java
Micronaut
Spring Boot
Databases
Apache Kafka
Azure Cosmos DB
PostgreSQL
Redis
Mobile
Reactive Programming
DevOps
Azure
Azure AKS
Azure DevOps
CI/CD
Docker
GitHub
GitHub Actions
Kubernetes
Apply
$47k – $102k per year (Estimated) • In office • Full-Time • 15+ years exp • Bachelor's Degree • Hyderabad
Java
Java
Micronaut
Spring Boot
Databases
Apache Kafka
Azure Cosmos DB
PostgreSQL
Redis
Mobile
Reactive Programming
DevOps
Azure
Azure AKS
Azure DevOps
CI/CD
Docker
GitHub
GitHub Actions
Kubernetes
Apply
$47k – $102k per year (Estimated) • In office • Full-Time • 12+ years exp • Bachelor's Degree • Hyderabad • Bengaluru • Coimbatore • Bhubaneswar
Java
Java
Micronaut
Spring Boot
Databases
Apache Kafka
Azure Cosmos DB
PostgreSQL
Redis
Mobile
Reactive Programming
DevOps
Azure
Azure AKS
Azure DevOps
CI/CD
Docker
GitHub
GitHub Actions
Kubernetes
Apply
$47k – $102k per year (Estimated) • In office • Full-Time • 12+ years exp • Bachelor's Degree • Nagpur • Bengaluru • Gurgaon • Chennai
Java
Java
Micronaut
Spring Boot
Databases
Apache Kafka
Azure Cosmos DB
PostgreSQL
Redis
AI/ML
AI Agents
Prompt Engineering
Mobile
Reactive Programming
DevOps
Azure
Azure AKS
Azure DevOps
CI/CD
Docker
GitHub
GitHub Actions
Kubernetes
Apply
$47k – $102k per year (Estimated) • In office • Full-Time • 15+ years exp • Bachelor's Degree • Hyderabad • Gurgaon
Java
Java
Spring Boot
Databases
Apache Kafka
Azure Cosmos DB
PostgreSQL
Redis
AI/ML
AI Agents
LLM
OpenAI
RAG
DevOps
Azure
Azure AKS
Azure DevOps
CI/CD
Docker
GitHub
GitHub Actions
Kubernetes
Platform Engineering
Apply
$150k – $180k per year • In office • Full-Time • 5+ years exp • Bellevue
Management
Jira
Marketing
Amplitude
Apply
Engineering Manager 9 hours ago
$161k – $310k per year (Estimated) • In office • Bachelor's Degree • Bellevue
AI/ML
Edge AI
DevOps
AWS
Azure
CI/CD
Docker
GCP
Kubernetes
Apply
$236k – $339k per year • In office • Full-Time • 10+ years exp • Bachelor's Degree • Bellevue
Java
Python
Databases
Snowflake
AI/ML
AI Agents
Feature Store
Apply
$215k – $260k per year • Equity • In office • Full-Time • 8+ years exp • Bachelor's Degree • San Francisco • Sunnyvale • Bellevue
C++
Go
DevOps
AWS
Azure
Cilium
eBPF
GCP
KVM
VMWare
Apply
$160k – $210k per year • Remote/Hybrid • Full-Time • 6+ years exp • Denver • Bellevue
C#
C++
Go
Python
SQL
Databases
Azure Cosmos DB
Azure SQL Database
DynamoDB
MySQL
AI/ML
Edge AI
DevOps
AWS
Azure
Azure AKS
CI/CD
GCP
Google GKE
Kubernetes
SLI/SLO/SLA
Analytics
Power BI
Management
UiPath
Apply
$159k – $285k per year (Estimated) • In office • Full-Time • 10+ years exp • Bellevue
AI/ML
AI Agents
Management
Smartsheet
Apply
See all jobs
This is one of many
368,910 more open roles from verified company boards, updated every day.