This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Senior Platform Architect based in Brazil.
This is a senior platform engineering opportunity focused on designing and operating reliable, scalable cloud infrastructure at production scale.
You will work at the intersection of platform engineering, SRE, DevOps, and modern AI infrastructure.
The role covers Kubernetes, infrastructure as code, GitOps, CI/CD, observability, automation, and workload reliability.
You will also help operate LLM and agentic systems, supporting the rapid growth of AI-powered products and platform capabilities.
Senior ownership is expected, with the freedom to identify complex infrastructure challenges and solve them through durable, scalable solutions.
You will influence architectural standards while improving performance, reliability, operational efficiency, and cloud costs.
This is a remote opportunity in LATAM, with compensation paid in USD and working hours aligned with the EST time zone.
Accountabilities:
- Design, build, and operate scalable backend and AI infrastructure supporting real-time and batch workloads.
- Establish architectural patterns that prioritize performance, reliability, scalability, and multi-tenant operation.
- Design and maintain deployment workflows covering versioning, staged rollouts, automated releases, monitoring, and safe rollback strategies.
- Build and operate production LLM and agentic systems, integrating model providers, APIs, gateways, tools, and external services.
- Manage reliability considerations for AI workloads, including rate limits, quotas, guardrails, graceful degradation, and operational resilience.
- Develop reusable services, APIs, automation, and data pipelines supporting AI-powered products and internal platform capabilities.
- Extend infrastructure-as-code practices using Terraform, reusable modules, and consistent provisioning patterns.
- Maintain and improve GitOps-based deployment workflows using tools such as ArgoCD or Flux.
- Operate Kubernetes workloads in production, including scaling, workload placement, tenant isolation, reliability, and infrastructure capacity management.
- Improve observability through metrics, logging, tracing, SLOs, alerting, incident response, and operational tooling.
- Identify opportunities to improve infrastructure performance, cloud efficiency, and AI workload costs.
- Build automation that eliminates repetitive manual processes and improves engineering productivity.
- Use agentic coding and AI development tools to accelerate development, infrastructure work, debugging, code review, and workflow automation.
- Define operational standards and contribute to architectural decisions that support continued platform growth.
- 5+ years of experience in platform engineering, SRE, infrastructure, or a closely related discipline.
- Significant hands-on experience operating production systems at scale.
- Strong SRE and DevOps foundation, including ownership of reliability, SLOs, post-mortems, incident management, and measurable improvements.
- Deep hands-on expertise with Terraform, including complex state management, reusable modules, multi-project configurations, and CI-driven plan/apply workflows.
- Strong production experience with GitOps tools such as ArgoCD or Flux and a deep understanding of declarative infrastructure management.
- Advanced Kubernetes knowledge, including production cluster operations and troubleshooting real-world failure scenarios.
- Strong experience with at least one major public cloud platform, such as AWS, Azure, or GCP, including networking, compute, IAM, storage, and multi-account or multi-project architectures.
- Hands-on experience building and operating CI/CD pipelines using GitHub Actions, Cloud Build, GitLab CI, or equivalent technologies.
- Strong automation-first mindset, with experience designing systems that reduce manual operational work.
- Active experience using agentic coding and AI development tools, with the ability to direct them effectively and critically review their output.
- Excellent communication skills and the ability to clearly explain operational decisions, technical trade-offs, and incident outcomes.
- MLOps experience and hands-on experience deploying or operating ML/AI workloads in production is a strong advantage.
- Strong GCP experience, particularly with VPC networking, Compute Engine, IAM, Cloud Storage, multi-project or organization design, and GKE, is a plus.
- Experience with GCP data services such as BigQuery, Dataflow, Pub/Sub, or Dataproc is advantageous.
- Experience with GPU or accelerator scheduling and node lifecycle management is a plus.
- Experience operating LLM inference at scale, including quotas, throttling, gateways, caching, and guardrails, is advantageous.
- Familiarity with ML orchestration tools such as Argo Workflows, Kubeflow, Airflow, Vertex AI Pipelines, or equivalent technologies is a plus.
- Experience with model registries, feature stores, experiment tracking, ML observability, or data and model drift monitoring is beneficial.
- Background in FinOps, cloud cost attribution, committed-use discounts, reservations, or accelerator capacity planning is advantageous.
- Experience with multi-tenant infrastructure, including isolation, noisy-neighbor mitigation, and tenant lifecycle management, is a plus.
- Previous experience scaling platform or ML infrastructure within a growing startup or enterprise environment is beneficial.
- Fully remote opportunity within LATAM.
- Compensation paid in USD.
- Working hours aligned with the EST time zone.
- Opportunity to work on modern cloud platforms and production-scale infrastructure.
- Hands-on exposure to LLMs, agentic systems, and AI-powered workloads.
- Opportunity to shape platform architecture and engineering standards.
- High level of technical ownership and autonomy.
- Opportunity to work with Kubernetes, Terraform, GitOps, CI/CD, and modern SRE practices.
- Exposure to complex scalability, reliability, observability, and cloud cost challenges.
- Opportunity to use modern AI-assisted engineering and agentic coding tools.
- Environment focused on continuous platform evolution rather than maintaining the status quo.
As a Senior Platform Architect, you will define and improve how cloud platforms are designed, deployed, operated, and scaled. You will combine strong SRE and DevOps practices with hands-on infrastructure engineering to create reliable platforms capable of supporting complex backend, AI, and multi-tenant workloads.
Requirements:
The ideal candidate is a senior platform, SRE, or infrastructure engineer with substantial experience operating production systems at scale. You should combine deep technical expertise with strong architectural judgment, automation-first thinking, and the ability to communicate complex operational decisions clearly to both technical teams and leadership.
Benefits:

