368,611open jobs
9,439companies
50,719added this week
Browse all
Salary
$112k – $217k per year (Estimated)
Location
Remote (United States)
Seniority
Senior · 5+ years exp
Employment
Full-Time
Overview
Company
Impact
Profile match
Jobgether is an AI-powered job platform focused on remote and flexible work. It matches candidates with relevant roles using skills and preference-based algorithms, and also offers career coaching and job-search guidance.

This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Senior DevOps Engineer based in United States.

Join a globally distributed infrastructure team responsible for powering highly available, production-critical financial technology.

You will design, automate, and operate cloud infrastructure on Google Cloud while helping engineering teams ship faster and more safely.

The role combines cloud architecture, Infrastructure-as-Code, Kubernetes, CI/CD, observability, networking, and reliability engineering.

You will take ownership of platform capabilities and build self-service solutions that reduce manual work and improve developer productivity.

You will also contribute to incident response, security, capacity planning, and continuous improvements across the infrastructure environment.

Working in an async-first international setting, you will have significant autonomy and a direct influence on technical decisions and platform evolution.

This is an opportunity to solve complex infrastructure challenges at scale while applying strong Platform-as-a-Product and SRE principles.

Accountabilities:

    • Design and evolve highly available cloud architecture on Google Cloud, including networking, interconnects, IAM, and resilient infrastructure topologies, using Terraform and GitOps practices.
    • Build and maintain secure CI/CD pipelines for Infrastructure-as-Code, incorporating automated planning and testing, code review, Policy-as-Code guardrails, drift detection, and controlled rollouts.
    • Develop a Platform-as-a-Product approach by creating self-service capabilities, reusable infrastructure patterns, and developer-friendly golden paths that enable teams to provision resources efficiently.
    • Operate and improve production Kubernetes/GKE environments, including workload deployment with Helm, scaling, networking, security, observability, and troubleshooting.
    • Strengthen platform observability across metrics, logs, traces, and alerting using technologies such as Prometheus, Thanos, Grafana, Loki, Tempo, and Alertmanager.
    • Operate infrastructure services and data platforms, including PostgreSQL and message brokers, while partnering with specialized SRE and database teams on complex operational challenges.
    • Participate in a global Follow-The-Sun on-call rotation, responding to alerts, supporting incidents, leading structured troubleshooting, coordinating escalations, and driving blameless post-mortems and follow-up actions.
    • Apply SRE principles such as SLIs, SLOs, error budgets, and capacity planning to improve reliability and operational maturity.
    • Design and troubleshoot cloud networking components including VPCs, routing, load balancing, DNS, TLS, firewalls, and interconnects.
    • Partner with security and engineering teams to implement secure-by-default infrastructure, least-privilege access, and effective infrastructure security practices.
    • Identify opportunities to eliminate manual toil through automation and continuously improve reliability, security, scalability, cost efficiency, and developer experience.
    • Mentor engineers and lead infrastructure initiatives that raise engineering standards and improve the overall platform.
    • Requirements:

      • 5+ years of professional experience in DevOps, SRE, Platform Engineering, Infrastructure Engineering, or a closely related discipline, with experience operating large-scale, highly available production systems.
      • Deep hands-on experience designing and operating cloud architecture on Google Cloud Platform, including landing zones, networking, IAM, and high-availability architectures.
      • Strong Terraform and Infrastructure-as-Code expertise, including structuring large codebases across multiple environments and applying GitOps and least-privilege principles.
      • Proven experience building CI/CD pipelines for Infrastructure-as-Code, including automated plan/apply workflows, code review, Policy-as-Code, drift detection, and safe deployment practices.
      • Significant production experience with Kubernetes, ideally Google Kubernetes Engine (GKE), and Helm-based workload deployment.
      • Strong understanding of cloud and L3/L4-L7 networking fundamentals, including VPCs, routing, load balancing, DNS, TLS, firewalls, and interconnects, with the ability to troubleshoot complex connectivity issues.
      • Practical experience with modern observability platforms covering metrics, logs, traces, and alerting, particularly Prometheus, Thanos, Grafana, Loki, Tempo, and Alertmanager.
      • Operator-level knowledge of production data stores and messaging systems such as PostgreSQL, RabbitMQ, or similar message brokers.
      • Solid understanding of SRE principles, including SLOs, error budgets, capacity planning, incident management, and root-cause analysis.
      • Strong scripting or programming capabilities in Python, Go, Shell, or equivalent technologies.
      • Demonstrated ability to independently troubleshoot complex production problems and drive issues through resolution.
      • Strong ownership, communication, documentation, and problem-solving skills, with the ability to work effectively in a globally distributed and async-first environment.
      • Willingness to participate in a Follow-The-Sun on-call rotation, including scheduled availability for urgent infrastructure incidents.
      • Experience with Policy-as-Code and IaC quality tools such as OPA/Conftest, Checkov, tflint, or Atlantis is a plus.
      • Experience managing Terraform state, module registries, and versioning at scale is a plus.
      • Experience building internal developer platforms or golden paths using tools such as Backstage or Tilt is a plus.
      • Working knowledge of Go, Linux, Docker/containerd, Ansible, Chef, or Puppet is advantageous.
      • Experience securing containers and Kubernetes/GKE environments is beneficial.
      • Knowledge of advanced Google Cloud security controls, regulated environments, SOC 2, secrets management, or audit logging is advantageous.
      • Familiarity with trading, brokerage, fintech, or other regulated and low-latency environments is a plus.
      • Relevant cloud certifications, particularly Google Cloud Professional certifications, are valued.
      • Benefits:

        • Competitive salary and stock options.
        • Health benefits.
        • One-time USD $500 home-office setup allowance for new hires.
        • USD $150 monthly stipend provided through a company expense card.
        • Fully remote work within the eligible location.
        • Opportunity to work with a globally distributed team across multiple regions.
        • High level of autonomy, ownership, and influence over infrastructure and platform strategy.
        • Opportunity to work on highly available, trading-critical systems and complex cloud infrastructure challenges.
        • A strong focus on developer productivity, automation, reliability, and Platform-as-a-Product practices.
        • Exposure to modern cloud, Kubernetes, Infrastructure-as-Code, observability, and SRE technologies.
        • Inclusive environment committed to building a diverse and collaborative workforce.
Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
368,611 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
In your city
$111k – $216k per year (Estimated) • Equity • Remote • Full-Time • 6+ years exp • Bachelor's Degree • United States
Python
Ruby
Databases
PostgreSQL
DevOps
AWS
Chef
CI/CD
Configuration Management
Datadog
Git
Jenkins
Nagios
Amazon CloudWatch
GitLab
Cybersecurity
Crowdstrike
FedRAMP
Analytics
Tableau
ETL/ELT
Apply
$33k – $78k per year (Estimated) • Equity • Remote • Full-Time • 8+ years exp • Bachelor's Degree • India
Apex
JavaScript
Python
TypeScript
Apex
Copado
Lightning Web Components
AI/ML
AutoGen
CrewAI
Fine-tuning
Hallucination
LangChain
LangGraph
LlamaIndex
LLM
RAG
Semantic Kernel
Semantic Search
Synthetic Data
Vertex AI
Agentforce
AWS Bedrock AgentCore
Semantic Search
AI Agents
Model Context Protocol
DevOps
AWS
CI/CD
GitHub Actions
Jenkins
Vector
GitHub
Cybersecurity
Crowdstrike
Management
Slack
Marketing
Salesforce
Apply
$18k – $51k per year (Estimated) • Remote • Moscow
Bash
Python
Databases
ClickHouse
PostgreSQL
AI/ML
Feature Store
Hadoop
LLM
DevOps
Ansible
CI/CD
Docker
GitLab
Kubernetes
Apply
$54k – $85k per year (Estimated) • In office • Full-Time • 5+ years exp • Gdańsk
Python
SQL
Python
Django
FastAPI
Flask
DevOps
Azure
Bitbucket
CI/CD
GitHub
GitLab
Apply
$29k – $70k per year (Estimated) • Remote/Hybrid • Full-Time • 11+ years exp • Bachelor's Degree • Hyderabad
Python
SQL
Databases
Databricks
Snowflake
DevOps
AWS
Azure
CI/CD
Apply
$152k – $229k per year • Remote • Full-Time
AI/ML
Human-in-the-Loop
Apply
$162k – $180k per year • Remote • Full-Time • 5+ years exp • Bachelor's Degree
Python
SQL
Databases
Snowflake
AI/ML
dbt
DevOps
AWS
Cybersecurity
HIPAA
Zero Trust
Analytics
ETL/ELT
Apply
$14k – $32k per year (Estimated) • Remote • Full-Time • 2+ years exp • Bachelor's Degree
DevOps
Incident Management
Management
ServiceNow
Apply
$29k – $60k per year (Estimated) • Remote • Full-Time • 8+ years exp
DevOps
Azure
Azure DevOps
Apply
$26k – $69k per year (Estimated) • Remote • Full-Time • 12+ years exp • Bachelor's Degree
Apply
See all jobs
This is one of many
368,611 more open roles from verified company boards, updated every day.