{"id":1785334,"url":"https://alion.io/job/virtasant-seniorstaff-platform-engineer","title":"Senior/Staff Platform Engineer","company":{"id":2124108,"name":"Virtasant","domain":"virtasant.com","url":"https://alion.io/company/virtasant","size_band":null,"is_staffing_agency":false,"employer_type":"direct","is_intermediary":false,"listed_via":null,"ats_vendor":"Teamtailor","truth_index":{"grade":"B","score":75,"open_postings":4,"ghost_share":0,"stale_share":1,"repost_share":0,"time_to_fill_p50_days":null,"computed_at":"2026-10-09T06:01:00Z"}},"role":"DevOps","role_family":"DevOps","seniority":"staff","employment_type":"full_time","work_mode":"on_site","remote_scope":null,"remote_scope_basis":null,"remote_working_hours":null,"hiring_geo_confidence":"explicit","locations":["Austin, United States"],"countries":["US"],"hiring_countries":[],"hiring_countries_total":0,"salary":null,"salary_estimate":{"min_usd":132000,"max_usd":258000,"period":"year","method":"role_seniority_country_remote_cell","sample_n":466},"experience_years_min":10,"visa_sponsorship":false,"relocation_package":false,"has_equity":false,"technologies":[{"name":"Ansible","optional":false},{"name":"AWS","optional":false},{"name":"CI/CD","optional":false},{"name":"Configuration Management","optional":false},{"name":"Datadog","optional":false},{"name":"DNS","optional":false},{"name":"Docker","optional":false},{"name":"GCP","optional":false},{"name":"Go","optional":false},{"name":"Grafana","optional":false},{"name":"IAM","optional":false},{"name":"Java","optional":false},{"name":"Kubernetes","optional":false},{"name":"Linux","optional":false},{"name":"OpenTelemetry","optional":false},{"name":"Platform Engineering","optional":false},{"name":"Prometheus","optional":false},{"name":"Puppet","optional":false},{"name":"Python","optional":false},{"name":"Service Mesh","optional":false},{"name":"Terraform","optional":false}],"status":"live","first_seen_at":"2026-10-02T15:15:46Z","employer_posted_date":"2026-10-02","last_verified_at":"2026-10-09T23:59:44Z","board_verified":true,"closed_at":null,"days_open":7,"trust":{"level":"ok","repost_count":null,"flags":[],"days_open":7},"description":"Senior/ Staff Platform Engineer\nType: Remote\nCoverage: Pacific Hours (8:00 AM - 5:00 PM PST and On-call every 4-5 weeks)\nJob Description:\nWe are looking for a Senior/Staff Platform Engineer to build, operate, and evolve large-scale production infrastructure. This is a hands-on platform and reliability engineering role for someone who has deep experience operating Kubernetes and cloud infrastructure, troubleshooting complex production systems, and building the automation and tooling that keeps those systems reliable.\nThe work spans Kubernetes, Linux, cloud infrastructure, networking, observability, CI/CD, reliability, and production operations. You will write production code and automation in Go, Python, or Java, but this is not primarily a software development role. We are looking for an engineer who understands the systems underneath the applications and can independently diagnose and solve infrastructure problems across multiple layers.\nThis is a highly autonomous role. You will work directly with technical stakeholders, own ambiguous infrastructure initiatives from design through production, and be trusted to drive technical decisions and critical issues without requiring constant direction.\nKey Responsibilities:\nPlatform and Kubernetes Engineering:\nDesign, build, operate, and improve production Kubernetes platforms.\n\nOwn platform-level concerns including cluster architecture, networking, workload isolation, resource management, security, upgrades, scaling, and reliability.\n\nTroubleshoot Kubernetes beyond the application layer, including networking/CNI, scheduling, node behaviour, resource constraints, controllers, and cluster-level failures.\n\nOperate and improve large-scale, highly available infrastructure across cloud, hybrid, virtualised, and/or bare-metal environments.\n\nDiagnose complex issues spanning Kubernetes, containers, Linux, networking, and underlying infrastructure.\n\nOptimise platform infrastructure for reliability, performance, scalability, and operational efficiency.\n\nSoftware Development and Automation:\nWrite, maintain, and improve production tooling and automation using Go, Python, or Java.\n\nBuild software and automation that improves platform operations, reliability, deployment, troubleshooting, and developer experience.\n\nRead, debug, and contribute to existing production codebases.\n\nDevelop internal services, APIs, integrations, and operational tooling where needed.\n\nAutomate repetitive operational processes and reduce manual intervention across the platform.\n\nApply sound software engineering practices, including testing, code review, maintainability, and documentation.\n\nReliability and Production Operations:\nOwn the reliability and operational health of critical production infrastructure.\n\nLead or contribute significantly to incident response for complex platform and infrastructure issues.\n\nInvestigate root causes and implement durable remediation rather than temporary fixes.\n\nDefine and improve SLOs, SLIs, alerting, and operational processes.\n\nTroubleshoot systems using logs, metrics, traces, profiling tools, and system-level diagnostics.\n\nDrive improvements in availability, performance, capacity, resilience, and operational readiness.\n\nContribute to disaster recovery planning, testing, and continuous improvement.\n\nInfrastructure as Code and Delivery:\nBuild and maintain infrastructure as code using Terraform and related automation technologies.\n\nCreate reusable infrastructure patterns and improve automation as the platform evolves.\n\nBuild and improve CI/CD and deployment workflows supporting large-scale engineering environments.\n\nBalance delivery speed with reliability, security, scalability, and operational requirements.\n\nWork across infrastructure provisioning, configuration management, deployment automation, and production operations.\n\nParticipate in planning and executing production cloud or infrastructure migrations, including dependency analysis, networking, cutover, rollback, and production validation.\n\nObservability:\nBuild and maintain production monitoring, metrics, dashboards, alerting, logging, and distributed tracing.\n\nImprove observability so engineers can identify and diagnose failures quickly.\n\nUse production telemetry to identify reliability, capacity, and performance problems before they become major incidents.\n\nContinuously improve incident detection and reduce time to diagnosis and recovery.\n\nCollaboration and Technical leadership:\nWork directly with customer and internal engineering teams to understand requirements, investigate problems, and drive technical solutions.\n\nCommunicate architecture, technical decisions, risks, trade-offs, and progress clearly to technical stakeholders.\n\nOwn complex infrastructure initiatives from initial problem definition through design, implementation, and production operation.\n\nContribute to architecture discussions, RFCs, design reviews, and technical direction.\n\nMentor other engineers and help improve engineering and operational practices across the team.\n\nOperate independently in ambiguous situations and take ownership when immediate technical or management direction is unavailable.\n\nQualifications:\nEducation and Experience:\n10+ years of professional experience in Platform Engineering, Site Reliability Engineering, Infrastructure Engineering, DevOps, or related fields; 10+ years is preferred for Staff-level candidates.\n\nSignificant hands-on experience operating complex production infrastructure and distributed systems.\n\nDemonstrated experience building and operating production Kubernetes platforms, not only deploying applications onto existing clusters.\n\nProduction programming experience with Go, Python, or Java.\n\nStrong experience with production reliability, incident response, troubleshooting, and operational ownership.\n\nExperience independently owning complex technical initiatives from an ambiguous starting point through production.\n\nExperience working directly with technical stakeholders or customers and communicating complex technical topics effectively.\n\nDegree in Computer Science, Engineering, or a related field, or equivalent practical experience.\n\nTechnical Skills:\nDeep understanding of production Kubernetes infrastructure, including cluster architecture, networking/CNI, NetworkPolicy, scheduling, resource management, nodes, security/RBAC, and cluster behaviour.\n\nStrong Linux fundamentals and hands-on production systems troubleshooting.\n\nStrong understanding of networking concepts including DNS, routing, load balancing, connectivity, and cloud/Kubernetes networking.\n\nProduction experience with at least one major cloud platform: AWS, GCP, or Alicloud.\n\nInfrastructure as code at scale using Terraform or equivalent tooling.\n\nConfiguration management and automation experience with technologies such as Ansible, Puppet, or similar.\n\nStrong production debugging and root-cause analysis skills across infrastructure and distributed systems.\n\nObservability experience using metrics, logs, traces, dashboards, and alerting platforms such as Prometheus, Grafana, OpenTelemetry, Datadog, or equivalent.\n\nExperience with CI/CD infrastructure and modern software delivery practices.\n\nDocker and container tooling as part of the production lifecycle.\n\nProduction coding and automation experience in Go, Python, or Java.\n\nUnderstanding of high availability, capacity planning, disaster recovery, and production resilience.\n\nPreferred:\nExperience planning and executing production cloud or infrastructure migrations, including cutover and rollback strategies.\n\nExperience operating large-scale or multi-cluster Kubernetes environments.\n\nExperience with hybrid cloud, on-premises, virtualised, or bare-metal infrastructure.\n\nExperience building or modifying Kubernetes controllers or operators.\n\nAdvanced Kubernetes networking, CNI, NetworkPolicy, service mesh, mTLS, or workload identity experience.\n\nMulti-cloud infrastructure experience.\n\nExperience designing and testing disaster recovery strategies.\n\nExperience with large-scale CI/CD or developer infrastructure.\n\nExperience with capacity planning and performance engineering.\n\nExperience building internal platform tooling or improving developer experience.\n\nExperience with security, infrastructure hardening, IAM, or compliance requirements.\n\nPrevious technical leadership, mentoring, or Staff/Principal-level engineering responsibilities.\n\nExperience working directly with external customers or stakeholders in a consulting or service-delivery environment.\n\nSoft Skills:\nExceptional written and verbal technical communication.\n\nStrong analytical, debugging, and problem-solving ability.\n\nHigh degree of ownership and ability to operate independently.\n\nComfortable making technical decisions and driving work forward in ambiguous environments.\n\nAble to communicate effectively with customers, engineers, and technical leadership.\n\nStrong technical judgment and ability to explain trade-offs rather than simply implement predefined solutions.\n\nAble to lead technically and influence others without requiring formal people-management authority.\n\nComfortable working within a distributed, highly technical team.\n\nPlease Note:\nThis is a hands-on platform and reliability engineering role. We are not looking for candidates whose experience has been limited to deploying applications onto Kubernetes, consuming managed cloud services, or provisioning infrastructure without owning its production operation.\n\nThe ideal candidate has personally built, operated, troubleshot, and improved production infrastructure and can clearly explain what they owned, how the underlying systems worked, and how they approached failures and architectural trade-offs.\n\nThis role requires significant autonomy and strong technical communication. Engineers should be comfortable working directly with stakeholders and driving complex technical issues without continuous oversight.\n\nProduction coding experience is required, but it may be in Go, Python, or Java. Go is not a requirement.\n\nCloud migration experience is strongly preferred but is not a knockout requirement.\n\nThis role is currently open to candidates based in Brazil, Mexico, or Canada.","description_format":"text","description_chars":10183,"description_truncated":false,"requirements":{"experience_years_min":10,"management_years_min":null,"team_size_min":null,"manages_managers":false,"education":{"level":"bachelor","optional":false},"security_clearance":false,"languages":[]},"benefits":[],"hiring_locations":[{"name":"Brazil","iso":"BR","kind":"country"},{"name":"Canada","iso":"CA","kind":"country"},{"name":"Mexico","iso":"MX","kind":"country"}],"hiring_excludes":[],"relocation_offered":false,"industries":["Artificial Intelligence","Books","Cloud Cost & Multi-Cloud Management"],"lifecycle":[{"event":"open","at":"2026-10-03T17:24:23Z"}],"visa":[],"liveness":{"score":87,"band":"hot","label":"Hiring now","p_open":1,"p_active":0.874,"p_room":1,"age_days":6,"expected_fill_days":30,"reasons":["conf:2","velocity","win:early"],"computed_at":"2026-10-09T06:01:00Z"},"pay":null,"html_url":"https://alion.io/job/virtasant-seniorstaff-platform-engineer","json_url":"https://alion.io/job/virtasant-seniorstaff-platform-engineer.json","meta":{"generated_at":"2026-10-10T02:40:15Z","cache_seconds":300,"methodology":"https://alion.io/methodology","terms":"https://alion.io/terms","contact":"https://alion.io/contact","api":"https://alion.io/developers","about":"Alion is a live layer of people, companies and AI agents: who they are, whether they are real and active right now, what they do and how to work with them, readable by people and by agents and paid per call.","catalog":"https://alion.io/catalog.json","usage":{"tier":"crawler","counted_by":"address","units_charged":1,"used_today":4905,"day_limit":5000,"remaining_today":95,"minute_limit":60,"resets_at":"2026-10-11T00:00:00Z"}},"offers":[{"id":"company.slices","title":"One company in depth, by slice","status":"live","price":{"credits":0.02,"usd":0.002,"plus_per_slice":{"credits":0.05,"usd":0.005}},"unit":"per company, plus each slice with data","note":"the employer in depth","call":{"mcp_tool":"get_company","arguments":{"id":2124108},"rest":"https://alion.io/mcp/rest/get_company?id=2124108"},"human":"https://alion.io/catalog?offer=company.slices&for=job%2Fvirtasant-seniorstaff-platform-engineer"},{"id":"market.stats","title":"A market slice: pay, demand and time to fill","status":"live","price":{"credits":1,"usd":0.1},"unit":"per slice","note":"pay, demand and time to fill for this role and place","call":{"mcp_tool":"market_stats"},"human":"https://alion.io/catalog?offer=market.stats&for=job%2Fvirtasant-seniorstaff-platform-engineer"},{"id":"job.search","title":"Open jobs by role, technology, place, pay and visa","status":"live","price":{"credits":0.02,"usd":0.002},"unit":"per posting in a list","note":"similar open postings","call":{"mcp_tool":"search_jobs"},"human":"https://alion.io/catalog?offer=job.search&for=job%2Fvirtasant-seniorstaff-platform-engineer"},{"id":"company.verify","title":"Is this company real and active right now","status":"pilot","price":null,"unit":"per company","request":{"url":"https://alion.io/catalog/request","method":"POST","body":"{\"offer\": \"company.verify\", \"for\": \"job/virtasant-seniorstaff-platform-engineer\", \"note\": \"what you need it for\"}"},"human":"https://alion.io/catalog?offer=company.verify&for=job%2Fvirtasant-seniorstaff-platform-engineer"}]}