718,867open jobs
42,744companies
102,468added this week
Browse all
Salary
$62k – $138k per year (Estimated)
Location
Remote (LATAM)
Seniority
Architect · 5+ years exp
Employment
Contractor

Confirmed on the employer's own hiring board on Sep 23, 2026. First seen by Alion on Aug 27, 2026.

Overview
Company
Impact
Profile match
Hire top-tier Latin American engineering professionals fluent in English and based in US time zones. We handle onboarding, payroll, benefits, taxes, and compliance, so you can focus on thriving.

We're looking for a Platform Architect who can set the standard for how we build, ship, and operate reliable cloud platforms at scale. You sit at the intersection of platform engineering and SRE. You'll own the path from infrastructure design to reliable production services, bringing DevOps rigor to complex systems.

This is not a ticket-processing role, and it's not a research role. You'll tackle hard problems: platform reliability, scalability, cost efficiency, deployment automation, and workload operations. You'll have the scope to solve them properly. Senior professionals here identify problems before they're asked and raise the ceiling on what the platform can do.

What you will work on

  • Build and operate scalable backend and AI infrastructure, supporting real-time and batch workloads with a focus on performance, reliability, and multi-tenant architecture.
  • Design and maintain deployment workflows across services and environments, including versioning, staged rollouts, automated releases, monitoring, and safe rollback strategies.
  • Build and operate LLM and agentic systems in production, integrating model providers, APIs, gateways, tools, and external services while managing rate limits, reliability, guardrails, and graceful degradation.
  • Develop reusable services, APIs, automation, and data pipelines that support AI-powered products and internal platform capabilities.
  • Extend infrastructure-as-code across the platform using Terraform and reusable patterns to provision and manage cloud services consistently across projects and environments.
  • Maintain GitOps-based deployment workflows using tools such as ArgoCD, improving automation and consistency across environments and tenants.
  • Run distributed workloads on Kubernetes (GKE), managing scaling, workload placement, tenant isolation, service reliability, and infrastructure capacity.
  • Improve platform observability and reliability through metrics, logging, tracing, SLOs, alerting, incident response practices, and operational tooling.
  • Identify performance and infrastructure cost improvements across cloud services, compute resources, APIs, and AI workloads.
  • Use agentic coding and AI development tools to accelerate engineering work, including scaffolding services, generating and reviewing infrastructure and application code, debugging, and automating repetitive workflows.

What you won’t find here

A platform team that maintains the status quo. We're actively building: new scale requirements, new architectural domains, and an ML/AI footprint that's growing fast. Senior engineers here shape how the platform evolves, and the tools available to do it are better than they've ever been.

Requirements

Must have

  • 5+ years in platform engineering, SRE, or infrastructure, with meaningful time operating production systems at scale.
  • Strong SRE/DevOps foundation. You've owned reliability for production services, defined and measured SLOs, run post-mortems, and driven measurable improvements.
  • Deep Terraform expertise. You actively manage complex Terraform state, reusable modules, and multi-project configurations in production, with CI-driven plan/apply workflows.
  • Strong GitOps background (ArgoCD or Flux in production). You understand declarative infrastructure management at depth and have opinions on how to do it well.
  • Deep Kubernetes knowledge. You've operated clusters in production, dealt with real failure modes, and understand the system at the control plane level.
  • Strong cloud infrastructure background across at least one major public cloud (AWS, Azure, or GCP), including networking, compute, IAM, storage, and multi-account or multi-project design.
  • Hands-on experience building and operating CI/CD pipelines (GitHub Actions, Cloud Build, GitLab CI, or equivalent).
  • Automation-first thinking at a senior level. You implement systems that eliminate entire categories of manual work.
  • Active user of agentic coding tools. You know how to direct them effectively, review their output critically, and use them to multiply your output.
  • Strong communicator. You can articulate operational decisions, technical trade-offs, and incident summaries clearly to engineers and leadership alike.

Nice to have

  • MLOps experience, including hands-on experience deploying and operating ML or AI workloads in production.
  • Strong GCP experience, including VPC networking, Compute Engine, IAM, Cloud Storage, multi-project or organization design, and GKE (Standard and/or Autopilot).
  • Hands-on experience with GCP data services, especially BigQuery in production: partitioning and clustering, query cost and performance tuning, and dataset-level IAM. Familiarity with at least one of Dataflow, Pub/Sub, or Dataproc.
  • Experience with GPU/accelerator scheduling and node lifecycle management in production (e.g., GKE node auto-provisioning, GPU time-sharing, or equivalent).
  • Experience operating LLM inference at scale, managing provider quotas/throttling (TPS/TUPS), gateways, caching, and guardrails (e.g., Vertex AI, Gemini API, or equivalent).
  • Experience with ML pipeline and orchestration tooling such as Argo Workflows, Kubeflow, Cloud Composer/Airflow, Vertex AI Pipelines, or equivalent.
  • Experience with model registries, feature stores, and experiment tracking (e.g., MLflow, Feast, or equivalent).
  • Familiarity with model and data drift monitoring and ML-specific observability.
  • Background in FinOps: inference cost attribution, committed use discount (CUD) and reservation planning, and accelerator capacity forecasting.
  • Familiarity with data infrastructure such as object storage, CDC pipelines, or lakehouse patterns.
  • Experience with multi-tenant infrastructure: isolation patterns, noisy neighbor mitigation, and tenant lifecycle management.
  • Prior experience scaling ML or platform infrastructure at a startup moving toward enterprise-grade requirements.

Location: Remote in LATAM

Payment in USD

Working hours: EST time zone

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
718,867 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
In your city
SDET III 7 hours ago
In office • 6+ years exp • Bachelor's Degree • Gurgaon
Python
JavaScript
C#
AI/ML
Copilot
Prompt Engineering
AI Agents
Gemini
RAG
Agentic Workflows
DevOps
GCP
GitHub Actions
GitLab CI
Azure
CI/CD
Jenkins
AWS
Docker
Kubernetes
Shift-Left
Self-Healing
Pipeline as Code
GitLab
Cybersecurity
Shift-Left Security
QA
TestNG
Selenium
Cypress
Playwright
Appium
Rest-Assured
Pytest
Apply
$33k – $80k per year (Estimated) • Remote/Hybrid • Full-Time • Saint Petersburg
Python
Bash
Databases
PostgreSQL
DevOps
Terraform
Ansible
Zabbix
Prometheus
GitLab CI
CI/CD
Git
Docker
Kubernetes
Ubuntu
Nginx
Grafana
CentOS Stream
Linux
Astra Linux
RED OS
Apply
$31k – $84k per year • Equity 0.1–0.2% • Remote • Full-Time • 1+ year exp
AI/ML
AI Agents
DevOps
GitHub Actions
Marketing
LinkedIn
Apply
Business Analyst 7 hours ago
up to $22k per year • In office • 2+ years exp • Astana
SQL
Databases
PostgreSQL
MS SQL
AI/ML
AI Agents
DevOps
Rest API
SOAP
Management
Confluence
Jira
Scrum
UML
BPMN
QA
Postman
Apply
$29k – $61k per year (Estimated) • In office • Full-Time • 8+ years exp • Bengaluru
Python
DevOps
Terraform
Azure
CI/CD
GitOps
Kubernetes
Bicep
Apply
Remote • Contractor • 5+ years exp
JavaScript
Java
Kotlin
TypeScript
Objective-C
AI/ML
Cursor
Claude
AI Agents
Frontend
GraphQL
Next.js
React.js
Mobile
React Native
Apply
Remote • Contractor • 5+ years exp
JavaScript
Java
Kotlin
TypeScript
Objective-C
AI/ML
Cursor
Claude
AI Agents
Frontend
GraphQL
Next.js
React.js
Mobile
React Native
Apply
$71k – $147k per year (Estimated) • Remote • Contractor
JavaScript
TypeScript
Node JS
Node JS
Nest.JS
Databases
PostgreSQL
AI/ML
AI Agents
Frontend
GraphQL
Apply
$62k – $129k per year (Estimated) • Remote • Contractor
JavaScript
TypeScript
Node JS
Node JS
Nest.JS
Databases
PostgreSQL
AI/ML
AI Agents
Frontend
GraphQL
Apply
See all jobs
This is one of many
718,867 more open roles from verified company boards, updated every day.