{"id":1253395,"url":"https://alion.io/job/opstree-gcp-devops-architect","title":"GCP DevOps Architect","company":{"id":2984556,"name":"Opstree","domain":"opstree.com","url":"https://alion.io/company/opstree","size_band":null,"is_staffing_agency":false,"employer_type":"direct","is_intermediary":false,"listed_via":null,"ats_vendor":null,"truth_index":null},"role":"DevOps","role_family":"DevOps","seniority":"staff","employment_type":null,"work_mode":"on_site","remote_scope":null,"remote_scope_basis":null,"remote_working_hours":null,"hiring_geo_confidence":"structured","locations":["Bengaluru, India"],"countries":["IN"],"hiring_countries":[],"hiring_countries_total":0,"salary":null,"salary_estimate":{"min_usd":26000,"max_usd":61000,"period":"year","method":"global_role_cell_scaled_by_country","sample_n":603},"experience_years_min":10,"visa_sponsorship":false,"relocation_package":false,"has_equity":false,"technologies":[{"name":"ArgoCD","optional":false},{"name":"CI/CD","optional":false},{"name":"Datadog","optional":false},{"name":"DNS","optional":false},{"name":"FinOps","optional":false},{"name":"GCP","optional":false},{"name":"GitHub","optional":false},{"name":"GitLab","optional":false},{"name":"GitOps","optional":false},{"name":"Google GKE","optional":false},{"name":"Grafana","optional":false},{"name":"IAM","optional":false},{"name":"Incident Management","optional":false},{"name":"Jenkins","optional":false},{"name":"Kubernetes","optional":false},{"name":"Least Privilege","optional":false},{"name":"Linux","optional":false},{"name":"OpenTelemetry","optional":false},{"name":"Platform Engineering","optional":false},{"name":"Progressive Delivery","optional":false},{"name":"Prometheus","optional":false},{"name":"Terraform","optional":false},{"name":"Apache Kafka","optional":true},{"name":"API Gateway","optional":true},{"name":"Istio","optional":true},{"name":"MySQL","optional":true},{"name":"PostgreSQL","optional":true},{"name":"Redis","optional":true},{"name":"Service Mesh","optional":true}],"status":"live","first_seen_at":"2026-09-10T06:06:27Z","employer_posted_date":null,"last_verified_at":"2026-09-10T06:06:27Z","board_verified":false,"closed_at":null,"days_open":21,"trust":{"level":"not_scored","repost_count":null,"flags":[],"days_open":21},"description":"Experience:\n\n10+ years in Cloud Infrastructure, DevOps, SRE, or Platform Engineering, with strong hands-on experience in Google Cloud Platform (GCP) and Google Kubernetes Engine (GKE).\n\nRole Overview:\n\nWe are looking for an experienced GCP DevOps Architect to design, implement, and operate highly scalable, secure, resilient, and cost-efficient cloud platforms on GCP.\n\nThe ideal candidate should have strong expertise in GCP, GKE, Kubernetes, DevOps, CI/CD, Infrastructure as Code, observability, security, and high-availability architecture, with experience building platforms capable of handling 5 - 10 million users / high-volume production traffic.\n\nThe candidate will be responsible for defining cloud architecture, DevOps standards, deployment strategies, reliability engineering practices, and platform automation.\n\nKey Responsibilities:\n\nGCP & Cloud Architecture:\n\n- Design highly available and scalable cloud architectures on GCP.\n\n- Architect platforms capable of supporting 5 - 10 million users and high-volume traffic.\n\n- Design multi-region / multi-zone architectures for high availability and disaster recovery.\n\n- Select and optimize appropriate GCP services based on scalability, reliability, performance, and cost.\n\n- Design networking architecture including VPC, subnets, Cloud Load Balancing, Cloud NAT, Cloud DNS, Private Service Connect, firewall policies, and hybrid connectivity.\n\n- Define cloud architecture standards, reference architectures, and engineering best practices.\n\nGKE & Kubernetes:\n\n- Strong hands-on experience designing and managing GKE production clusters.\n\n- Design scalable GKE architectures including regional and private GKE clusters, node pools, workload isolation, cluster autoscaling, HPA/VPA, Workload Identity, Ingress/Gateway architecture, network policies, and pod security.\n\n- Optimize Kubernetes workloads for performance, availability, and cost.\n\n- Define Kubernetes deployment, upgrade, backup, and disaster recovery strategies.\n\n- Troubleshoot complex production issues involving Kubernetes, networking, compute, and application performance.\n\nDevOps & CI/CD:\n\n- Design and implement enterprise-grade CI/CD pipelines.\n\n- Strong experience with tools such as GitHub/GitLab, Jenkins, Cloud Build, Argo CD, and/or other CI/CD platforms.\n\n- Implement GitOps-based deployment strategies where appropriate.\n\n- Automate build, test, security scanning, deployment, rollback, and release processes.\n\n- Implement progressive delivery strategies such as Blue/Green, Canary, and Rolling deployments.\n\n- Establish DevOps standards across development and operations teams.\n\nInfrastructure as Code:\n\n- Strong hands-on experience with Terraform.\n\n- Build reusable Terraform modules and infrastructure frameworks.\n\n- Automate provisioning and configuration of GCP infrastructure.\n\n- Implement infrastructure versioning, state management, policy controls, and automated validation.\n\nScalability & Reliability:\n\n- Architect systems for millions of users and high concurrent traffic.\n\n- Design autoscaling strategies across compute, Kubernetes, databases, and networking layers.\n\n- Implement SRE practices including SLIs/SLOs/SLAs, error budgets, capacity planning, reliability engineering, incident management, and performance engineering.\n\n- Conduct architecture reviews, scalability assessments, and production readiness reviews.\n\n- Design fault-tolerant systems with appropriate RTO/RPO targets.\n\nMonitoring & Observability:\n\n- Design comprehensive monitoring and observability solutions using Google Cloud Operations Suite, Prometheus, Grafana, OpenTelemetry, and Datadog.\n\n- Implement infrastructure, application, Kubernetes, and business-level monitoring.\n\n- Establish centralized logging, metrics, tracing, alerting, and dashboards.\n\n- Analyze production performance and identify bottlenecks.\n\nSecurity:\n\n- Implement GCP security best practices across infrastructure and Kubernetes.\n\n- Strong understanding of IAM, Service Accounts, Workload Identity, Secret Manager, KMS, VPC Service Controls, Organization Policies, Security Command Center, and container/image security.\n\n- Implement least-privilege access and secure CI/CD pipelines.\n\n- Integrate vulnerability scanning and security controls into the DevOps lifecycle.\n\nCost Optimization:\n\n- Monitor and optimize GCP infrastructure costs.\n\n- Optimize GKE compute, node pools, autoscaling, storage, networking, and logging costs.\n\n- Establish cloud FinOps practices and cost governance.\n\n- Identify opportunities for capacity optimization without compromising reliability.\n\nRequired Technical Skills:\n\n- Must Have: Strong GCP expertise, Strong GKE / Kubernetes expertise, Strong Terraform / IaC, Advanced CI/CD and DevOps, GCP networking and security, Linux and container technologies, Production experience with highly scalable systems, Experience supporting 5 - 10 million users or equivalent high-volume traffic, High Availability and Disaster Recovery architecture, Monitoring, logging, and observability, Strong troubleshooting and incident-management skills.\n\n- Good to Have: Google Cloud Professional Cloud Architect certification, Google Cloud Professional Cloud DevOps Engineer certification, Service mesh experience such as Istio, Argo CD / GitOps, Prometheus / Grafana / OpenTelemetry, Anthos / GKE Enterprise, FinOps experience, Experience with Kafka, Redis, PostgreSQL/MySQL, NoSQL platforms, Experience with API Gateway / Apigee, Experience with microservices architecture.\n\nLeadership & Soft Skills:\n\n- Strong architectural and problem-solving capabilities.\n\n- Ability to communicate complex cloud architecture to both technical and non-technical stakeholders.\n\n- Experience mentoring DevOps, SRE, and cloud engineering teams.\n\n- Ability to lead architecture decisions and establish engineering standards.\n\n- Strong ownership of production reliability and operational excellence.\n\n- Ability to work effectively with application, security, data, and infrastructure teams.\n\nSkills\nDevOps, Google Cloud Platform, Kubernetes, Terraform, Cloud Infrastructure, Linux, Jenkins, Monitoring Tools","description_format":"text","description_chars":6099,"description_truncated":false,"requirements":{"experience_years_min":10,"management_years_min":null,"team_size_min":null,"manages_managers":false,"education":null,"security_clearance":false,"languages":[]},"benefits":[],"hiring_locations":[],"hiring_excludes":[],"relocation_offered":false,"industries":[],"lifecycle":[{"event":"open","at":"2026-09-25T18:04:06Z"}],"liveness":{"score":56,"band":"ok","label":"Likely open","p_open":0.85,"p_active":0.735,"p_room":0.9,"age_days":20,"expected_fill_days":30,"reasons":["seen:20","velocity","win:mid"],"computed_at":"2026-10-01T05:45:00Z"},"pay":null,"html_url":"https://alion.io/job/opstree-gcp-devops-architect","json_url":"https://alion.io/job/opstree-gcp-devops-architect.json","meta":{"generated_at":"2026-10-01T21:09:00Z","cache_seconds":300,"methodology":"https://alion.io/methodology","terms":"https://alion.io/terms","contact":"https://alion.io/contact","api":"https://alion.io/developers","usage":{"tier":"crawler","counted_by":"address","units_charged":1,"used_today":3719,"day_limit":5000,"remaining_today":1281,"minute_limit":60,"resets_at":"2026-10-02T00:00:00Z"}}}