Nosso cliente está conduzindo uma modernização massiva de seus pipelines de dados, migrando cargas de trabalho legadas em Databricks para uma nova arquitetura nativa no Google Cloud. Esta posição é central para essa transformação: o profissional atuará como um dos pilares hands-on da refatoração, transformando lógicas complexas de processamento distribuído em padrões modernos de ELT, sem comprometer a continuidade das integrações já existentes.
O Data Developer Senior atua na conversão direta de notebooks legados, na construção de pipelines de ingestão escaláveis e na garantia de que a arquitetura em camadas (Raw, Trusted Core e Gold) preserve os contratos de dados que sustentam sistemas e dashboards em produção. O papel combina profundidade técnica na reengenharia de código acoplado com colaboração próxima aos times de negócio para validar cada etapa da migração.
Responsibilities:
- Code Refactoring: Translates legacy Databricks notebook logic into modern ELT patterns, prioritizing declarative SQL-based transformation for standard cases and developing Python-based distributed processing routines for intricate logic (such as RDD-based operations).
- Data Contract Preservation: Ensures backward compatibility during migration by implementing trusted-core layering and reverse-view strategies, preventing breaking changes for downstream consuming systems and dashboards.
- Ingestion Pipeline Development: Builds scalable ingestion pipelines by parameterizing YAML configurations that feed an automated DAG generation pipeline, orchestrated through a workflow scheduler and consuming standardized processing templates.
- Data Quality Implementation: Applies synchronous and unit-level assertions within the transformation layers to prevent null values, enforce key uniqueness, and validate domain conformity.
- Migration Validation: Executes technical validation projects by reconciling data output between the legacy environment and the new platform, confirming record-level and field-level parity ("cara-a-crachá" checks).
- GitOps Development: Follows strict CI/CD workflows through feature branches, pull requests, and automated deployment pipelines to development and production environments.
- Cross-Team Alignment: Collaborates consultatively with business Data Stewards to align integration dependencies, negotiate refactoring scope, and validate migrated outputs side by side with stakeholders.
Requirements:
- Strong, hands-on technical depth in Python, PySpark, and advanced SQL for large-scale distributed data processing
- Robust experience migrating data lakes and refactoring legacy code, including reverse-engineering tightly coupled logic (such as Databricks or Azure Data Factory pipelines) into modular cloud-native components
- Practical mastery of the Google Cloud data engineering ecosystem, particularly BigQuery, Dataform, and Dataproc Serverless
- Solid understanding of Medallion Architecture (Raw/Bronze, Silver/Trusted Core, Gold) and analytical data modeling strategies, including Star Schema and One Big Table approaches
- Hands-on experience with modern GitOps development flows (code review, pull request submission) and orchestration automation using Airflow/Composer
- Maturity in data quality practices, with testing applied directly to transformation pipelines
- Consultative, collaborative communication style to align cross-team dependencies, negotiate refactoring scope, and validate outcomes alongside business Data Stewards
Nice to Have:
- Experience applying Generative AI agents (such as Gemini, Vertex AI, or Claude) to automate PySpark-to-SQL code refactoring, dependency graph analysis, or automated documentation generation
- Prior experience with Change Data Capture (CDC) architectures and event-driven ingestion patterns
- Strong performance tuning skills in BigQuery, including query optimization, partitioning strategies, and execution cost control
- Deep knowledge of the boundaries between Data Vault and Star Schema modeling approaches for trusted/silver layers

