{"id":1138487,"url":"https://alion.io/job/astreya-devops-engineer-iv-infrastructure","title":"DevOps Engineer IV (Infrastructure)","company":{"id":1707,"name":"Astreya","domain":"astreya.com","url":"https://alion.io/company/astreya","size_band":"1001-5000","is_staffing_agency":false,"is_intermediary":false,"ats_vendor":"Workday","truth_index":{"grade":"B","score":77,"open_postings":13,"ghost_share":0,"stale_share":0.923,"repost_share":0,"time_to_fill_p50_days":28,"computed_at":"2026-09-23T05:45:00Z"}},"role":"DevOps","role_family":"DevOps","seniority":"senior","employment_type":"full_time","work_mode":"remote","remote_scope":"stated_countries","hiring_geo_confidence":"structured","locations":["India"],"countries":["IN"],"hiring_countries":["IN"],"hiring_countries_total":1,"salary":null,"salary_estimate":{"min_usd":24000,"max_usd":53000,"period":"year","method":"role_seniority_country_cell","sample_n":8},"experience_years_min":8,"visa_sponsorship":false,"relocation_package":false,"has_equity":false,"technologies":[{"name":"AWS","optional":false},{"name":"Azure","optional":false},{"name":"CI/CD","optional":false},{"name":"DNS","optional":false},{"name":"GCP","optional":false},{"name":"Go","optional":false},{"name":"Google GKE","optional":false},{"name":"Helm","optional":false},{"name":"IAM","optional":false},{"name":"Incident Management","optional":false},{"name":"Kubernetes","optional":false},{"name":"Least Privilege","optional":false},{"name":"LLM","optional":false},{"name":"Platform Engineering","optional":false},{"name":"PostgreSQL","optional":false},{"name":"Python","optional":false},{"name":"Redis","optional":false},{"name":"SQL","optional":false},{"name":"Terraform","optional":false},{"name":"Vertex AI","optional":false},{"name":"Azure AKS","optional":true},{"name":"ISO 27001","optional":true},{"name":"ITSM","optional":true},{"name":"Microsoft Entra ID","optional":true},{"name":"ServiceNow","optional":true},{"name":"SOC 2","optional":true},{"name":"Vault","optional":true}],"status":"live","first_seen_at":"2026-09-23T09:27:10Z","employer_posted_date":"2026-09-23","last_verified_at":"2026-09-23T13:45:13Z","board_verified":true,"closed_at":null,"days_open":0,"trust":{"level":"ok","repost_count":null,"flags":[],"days_open":0},"description":"Lead DevOps & Cloud Infrastructure Engineer\nWhat this Job Entails\nThis is a hands-on lead role owning the cloud infrastructure and DevOps function for Astreya's product portfolio. The person in this seat is accountable for how our products get built, deployed, secured and run - in our own environments and inside customer environments.\nGoogle Cloud is the primary platform; Azure is secondary. This is not a role that sets direction and hands the work to someone else: the expectation is that this person writes the Terraform, debugs the cluster, runs the deployment, and then mentors the engineers who will do it next time. They act as the subject matter expert for infrastructure across product, engineering and delivery teams, and as the technical owner in customer conversations about deployment, security and hosting models.\nScope\nOwns the infrastructure and deployment layer end to end, from CI pipeline to production runtime, across multiple products on independent release cadences\nResolves complex problems where the diagnosis requires in-depth evaluation of many interacting variables - networking, identity, cluster behaviour, managed service limits, cost\nExercises independent judgment in selecting platform services, patterns and tooling, and is expected to defend those choices technically and commercially\nOperates as a player-coach: a small team reports to this role, and the role remains individually productive\nYour Roles and Responsibilities\nCloud infrastructure (primary)\nDesign, build and operate our GCP footprint: GKE, Artifact Registry, Cloud SQL (Postgres), Cloud Storage, Memorystore (Redis), Vertex AI, Cloud Monitoring/Logging, Cloud DNS\nOwn network and perimeter design: VPC architecture, Private Service Connect, Cloud NAT, Cloud Armor, private connectivity to customer systems\nOwn secrets, keys and identity: Secret Manager, Workload Identity, IAM design and least-privilege service accounts, KMS, certificate lifecycle\nRun R&D on platform services we haven't used yet - evaluate the GCP service, identify the Azure and AWS equivalents, and produce a recommendation with a working proof of concept, not a slide\nManage cloud cost: attribution by product and environment, budget alerts, rightsizing, commitment planning\nDeployment into environments (a core reason this role exists)\nMake deploying our products into a new environment a repeatable, documented, low-drama exercise - both customer-VPC and Astreya-hosted models\nOwn the deployment artefacts: Helm charts, namespace and pod topology, environment configuration, migration and rollback paths\nWork directly with customer infrastructure and security teams during onboarding - answer their architecture and security questions, adapt to their constraints, and get us to production\nProduce and maintain the evidence infrastructure security reviews and audits ask for: architecture diagrams, data flow, access controls, encryption posture, audit logging, retention\nDevOps and platform engineering\nOwn CI/CD across products: build pipelines, artefact promotion, environment strategy, source control branching strategies across multiple release cadences\nWrite automation and internal tooling for provisioning and operating infrastructure - infrastructure as code (Terraform) as the default, scripts and services where it isn't\nBuild the observability layer: metrics, logs, traces, alerting, dashboards and SLOs that let us find problems before customers do\nImprove scalability, reliability, capacity and performance of the platform, including the infrastructure serving AI/LLM workloads\nHarden the platform: image scanning, dependency and vulnerability management, network policy, secrets hygiene, patching\nLeadership and process\nLead and mentor DevOps and infrastructure engineers; set standards and review their work\nWork with product owners and engineering leads to understand requirements, surface infrastructure bottlenecks early, and propose resolutions\nOwn incident response for platform issues and produce root cause analyses for outages that are honest about cause and specific about prevention\nDesign and implement improvements to existing support processes and tooling; introduce innovations and follow through on execution, not just proposal\nDocument decisions, runbooks and resolution history so the next person doesn't rediscover them\nOther duties as required. This list is not a comprehensive inventory of all responsibilities assigned to this position\nRequired Qualifications/Skills\nBachelor's degree (B.S/B.A) from a four-year college or university and 8+ years' related experience and/or training; or an equivalent combination of education and experience\nDeep, hands-on Google Cloud experience - has designed and run production workloads on GCP, not just passed a certification\nStrong Kubernetes experience in production: GKE preferred, including networking, ingress, autoscaling, resource management and debugging failures under load\nInfrastructure as code at a professional standard - Terraform, module design, state management, review discipline\nStrong scripting and coding ability (Python, Go, Bash or equivalent) and enough software development experience to work inside application repositories, not just around them\nPractical cloud networking and security depth: VPC design, private connectivity, IAM, secrets management, TLS\nCI/CD ownership across multiple products and release trains, with source control branching strategies\nMonitoring, alerting and incident management experience, including writing the RCA afterwards\nExperience with open-source tooling in large distributed systems\nGood communication skills - can hold a technical conversation with a customer's security team and a working conversation with a developer on the same day\nAbility to pick up unfamiliar technology quickly and carry several threads at once\nPreferred Qualifications\nWorking Azure knowledge (AKS, Key Vault, Entra ID, VNet) and the ability to translate a GCP design onto it; AWS exposure a bonus\nExperience deploying a product into customer-controlled cloud environments as a vendor\nExperience supporting AI/ML workloads in production - Vertex AI, model endpoints, GPU scheduling, inference cost control\nExposure to SOC 2, ISO 27001 or comparable audits from the infrastructure side\nExperience integrating with enterprise ITSM platforms (ServiceNow) and the connectivity that requires\nPrior experience as a first or early infrastructure hire on a product team\nRelevant certification (Google Professional Cloud Architect, Professional Cloud DevOps Engineer, CKA)\nPhysical Demand & Work Environment\nMust have the ability to perform office-related tasks which may include prolonged sitting or standing\nMust have the ability to move from place to place within an office environment\nMust be able to use a computer\nMust have the ability to communicate effectively\nSome positions may require occasional repetitive motion or movements of the wrists, hands, and/or fingers","description_format":"text","description_chars":6951,"description_truncated":false,"requirements":{"experience_years_min":8,"management_years_min":null,"team_size_min":null,"manages_managers":false,"education":{"level":"bachelor","optional":false},"security_clearance":false,"languages":[]},"benefits":[],"hiring_locations":[{"name":"India","iso":"IN","kind":"country"}],"hiring_excludes":[],"relocation_offered":false,"industries":["Artificial Intelligence","Information Technology","IT Outsourcing"],"lifecycle":[{"event":"open","at":"2026-09-23T09:27:10Z"}],"liveness":{"score":86,"band":"hot","label":"Hiring now","p_open":1,"p_active":0.86,"p_room":1,"age_days":0,"expected_fill_days":28,"reasons":["conf:1","win:early","comp:brand"],"computed_at":"2026-09-23T15:20:08Z"},"pay":null,"html_url":"https://alion.io/job/astreya-devops-engineer-iv-infrastructure","json_url":"https://alion.io/job/astreya-devops-engineer-iv-infrastructure.json","meta":{"generated_at":"2026-09-23T15:20:08Z","cache_seconds":300,"methodology":"https://alion.io/methodology","terms":"https://alion.io/terms","contact":"https://alion.io/contact","api":"https://alion.io/developers"}}