{"id":1290002,"url":"https://alion.io/job/100ms-platform-engineer-core-infrastructure","title":"Platform Engineer — Core Infrastructure","company":{"id":1923312,"name":"100ms","domain":"100ms.live","url":"https://alion.io/company/100ms","size_band":"51-200","is_staffing_agency":false,"employer_type":"direct","is_intermediary":false,"listed_via":null,"ats_vendor":"Lever","truth_index":{"grade":"B","score":75,"open_postings":7,"ghost_share":0,"stale_share":1,"repost_share":0,"time_to_fill_p50_days":null,"computed_at":"2026-09-30T05:45:00Z"}},"role":"DevOps","role_family":"DevOps","seniority":"middle","employment_type":"full_time","work_mode":"on_site","remote_scope":null,"remote_scope_basis":null,"remote_working_hours":null,"hiring_geo_confidence":"structured","locations":["Bengaluru, India"],"countries":["IN"],"hiring_countries":[],"hiring_countries_total":0,"salary":{"min":1700000,"max":3500000,"currency":"INR","period":"year","gross":null,"usd_annual":36687},"salary_estimate":null,"experience_years_min":3,"visa_sponsorship":false,"relocation_package":false,"has_equity":false,"technologies":[{"name":"AI Agents","optional":false},{"name":"Alertmanager","optional":false},{"name":"ArgoCD","optional":false},{"name":"CI/CD","optional":false},{"name":"GCP","optional":false},{"name":"GitOps","optional":false},{"name":"Google GKE","optional":false},{"name":"Grafana","optional":false},{"name":"Helm","optional":false},{"name":"IAM","optional":false},{"name":"Kubernetes","optional":false},{"name":"Least Privilege","optional":false},{"name":"Linux","optional":false},{"name":"LLM","optional":false},{"name":"Loki","optional":false},{"name":"Prometheus","optional":false},{"name":"Rest API","optional":false},{"name":"Shift-Left","optional":false},{"name":"Shift-Left Security","optional":false},{"name":"Terraform","optional":false},{"name":"WebRTC","optional":false},{"name":"Falco","optional":true},{"name":"HashiCorp Vault","optional":true},{"name":"HIPAA","optional":true},{"name":"Platform Engineering","optional":true},{"name":"Sealed Secrets","optional":true},{"name":"Trivy","optional":true}],"status":"live","first_seen_at":"2026-04-27T16:09:22Z","employer_posted_date":"2026-04-27","last_verified_at":"2026-09-30T01:24:03Z","board_verified":true,"closed_at":null,"days_open":155,"trust":{"level":"stale","repost_count":0,"flags":["stale"],"days_open":155},"description":"What Will You Do\nOwn and operate production infrastructure across multiple GKE clusters supporting both real-time video workloads and AI agent pipelines - with HA, autoscaling, and full observability tuned to the demands of each.\nManage GitOps workflows using Argo CD for automated, version-controlled, and auditable deployments across both product lines.\n\nMaintain and optimize monitoring & alerting stacks using Open Source Monitoring Tools - with product-specific SLOs for low-latency video (jitter, packet loss, stream health) and AI workflow reliability (task throughput, failure rates, retry queues).\n\nImplement infrastructure as code using Terraform for GCP resources and helm chart for Kubernetes manifests, with a strong bias toward repeatability and auditability.\n\nSupport the unique infrastructure demands of real-time video - including media server scaling, WebRTC infrastructure, low-latency networking, and high-throughput data paths.\n\nSupport AI agent workloads - including LLM inference infrastructure, async task queues, and integration pipelines with external healthcare systems.\n\nLead or support incident response, cluster upgrades, and disaster recovery procedures across both platforms.\n\nOwn the security posture of our infrastructure - enforce least-privilege access controls, manage secrets hygiene, and drive security hardening across clusters and services.\n\nImplement and maintain compliance-aligned controls relevant to healthcare data environments (e.g., encryption at rest/in transit, audit logging, network segmentation).\n\nCollaborate with product and engineering teams to embed security early in the development lifecycle - shift-left on vulnerability scanning, dependency audits, and policy enforcement.\n\nWho Can Apply\nComputer Science / Engineering degree or equivalent practical experience.\n\nMinimum 3 years of hands-on experience with Kubernetes in a production environment.\n\nStrong knowledge of CI/CD pipelines and GitOps workflows using Argo CD or similar tools.\n\nProficient in infrastructure automation using Terraform and Helm.\n\nExperience in managing open source monitoring and logging stacks (Prometheus, Loki, Grafana, Alertmanager etc).\n\nWorking knowledge of cloud security principles - IAM, network policies, pod security, RBAC, and secrets management.\n\nComfortable with Linux systems, shell scripting, and basic networking - including an understanding of UDP/TCP behaviour relevant to real-time media or distributed systems.\n\nGood to Have\nPrior experience managing large-scale, multi-tenant or mixed-workload infrastructure.\n\nExposure to real-time media infrastructure - WebRTC, SFUs, TURN/STUN servers, or media server orchestration.\n\nHands-on experience with secrets management tools such as HashiCorp Vault or Sealed Secrets.\n\nFamiliarity with security scanning and policy tools (e.g., Trivy, OPA/Gatekeeper, Falco).\n\nExperience with GCP and GKE specifically.\n\nExposure to compliance frameworks relevant to healthcare or regulated industries (HIPAA awareness is a plus).\n\nExperience with AI/ML inference workloads or async pipeline infrastructure (queues, workers, orchestrators).\n\nExperience with open source contributions.\n\nStrong inclination to stay current with evolving infrastructure, security, and platform engineering practices - and a willingness to share ideas internally or externally.\n\nAbility to communicate fluently and clearly in English, written and spoken.\n\nWhy 100ms\nYou'll work on genuinely varied infrastructure - real-time video at scale and AI-driven healthcare automation are both hard problems with different constraints, and you'll own both.\n\nYou'll be part of a small, high-ownership team at a fast-growing, engineering-first startup with a meaningful mission - powering real-time experiences and helping patients access treatment faster.\n\nYou'll work alongside engineers with deep experience in distributed systems, real-time media, AI infrastructure, and platform engineering at scale.\n\nYou'll have the freedom to grow as an individual contributor or step into a team leadership role - with room to define your own goals and impact.\n\nSecurity and infrastructure are first-class concerns here, not support functions - your work directly shapes the trust and reliability our customers depend on.\n\nAdditional Information\nWe place a strong emphasis on in-office collaboration to maintain a tight feedback loop and a strong engineering culture.\n\nEmployees are expected to work from the office at least three days a week.\n\nWebsite\nhttps://www.100ms.ai/\n https://www.100ms.live","description_format":"text","description_chars":4545,"description_truncated":false,"requirements":{"experience_years_min":3,"management_years_min":null,"team_size_min":null,"manages_managers":false,"education":null,"security_clearance":false,"languages":[{"language":"English","level":"All levels","optional":false}]},"benefits":[],"hiring_locations":[],"hiring_excludes":[],"relocation_offered":false,"industries":["Software","Virtual Events","Photo & Video Editing Software"],"lifecycle":[{"event":"open","at":"2026-09-26T07:42:48Z"}],"liveness":{"score":8,"band":"cold","label":"Long shot","p_open":1,"p_active":0.283,"p_room":0.28,"age_days":155,"expected_fill_days":38,"reasons":["conf:4","win:tail","crowd:"],"computed_at":"2026-09-30T05:45:00Z"},"pay":{"stated_usd_annual":36687,"is_top_pay":false},"html_url":"https://alion.io/job/100ms-platform-engineer-core-infrastructure","json_url":"https://alion.io/job/100ms-platform-engineer-core-infrastructure.json","meta":{"generated_at":"2026-09-30T06:18:58Z","cache_seconds":300,"methodology":"https://alion.io/methodology","terms":"https://alion.io/terms","contact":"https://alion.io/contact","api":"https://alion.io/developers","usage":{"tier":"crawler","counted_by":"address","units_charged":1,"used_today":4413,"day_limit":5000,"remaining_today":587,"minute_limit":60,"resets_at":"2026-10-01T00:00:00Z"}}}