1,356,297open jobs
79,166companies
206,537added this week
Browse all
Salary
$75k – $131k per year (gross)
Location
In office (Singapore)
Seniority
Senior · 5+ years exp
Employment
Full-Time

First seen by Alion on Oct 7, 2026.

Overview
Company
Impact
Profile match
At Arukah, we design, develop and sell institutional, registry-approved carbon credit and biofuels projects from large scale agriculture waste in the Global South for the world’s most discerning governments, companies and individuals.
Senior Data Platform Engineer

Platform, Data Engineering & Analytics

About Arukah

Arukah builds technology that connects biochar and biogas operations to measurable climate impact. Our systems bring together production records, sensors, laboratory results and supporting documents for digital measurement, reporting and verification (dMRV).

Our next challenge is to evolve the technology supporting a few plants into a platform for 10s of hundreds of plants, with an elite small engineering team. We need reusable systems, trustworthy data and thoughtful choices about what to build and what to obtain from managed services.

The opportunity

Own the data platform from source ingestion through operational analytics and registry reporting. You will assess the existing codebase, shape the architecture and deliver improvements incrementally while supporting current operations. The responsibilities include re-architecting where needed, not simply maintaining existing pipelines.

This role combines platform engineering, data engineering and analytics engineering. We are looking for strong platform judgment and hands-on Python and SQL skills, supported by solid analytical modeling. You should be comfortable choosing a managed capability, building a domain-specific component, or simplifying a workflow based on reliability, correctness, cost and the team's capacity to operate it.

You will use AI coding assistants as part of everyday engineering, with responsibility for the design, correctness and maintainability of everything delivered. You will also help operate AI-based document extraction and human review within the data workflow.

What you will own

  • Architecture for multiple plants. Design shared capabilities, facility configuration, access boundaries and data isolation. Establish a practical path from a few plants to many plants, using representative workloads and operational needs to guide decisions. Make each additional plant easier to onboard without creating separate code forks.

  • Managed-service and build-versus-buy decisions. Evaluate compute, storage, databases, orchestration, connectors, identity and monitoring. Consider engineering time, service limits, recurring cost, recovery, vendor dependencies and exit options. Prefer managed services when they meet requirements and reduce ongoing operational work.

  • Reliable ingestion and processing. Build reusable integrations for operational sheets, sensors, APIs, documents and laboratory results. Handle duplicates, late or missing data, schema changes, retries and historical reprocessing. Make source failures visible and recoverable.

  • Shared data models and analytics. Model plants, equipment, production batches, feedstock deliveries, samples and evidence with explicit grain, stable identifiers, units, time semantics and history. Create tested datasets and consistent metrics for plant operations, portfolio analysis and reporting.

  • Traceable calculations and evidence. Preserve source provenance, human corrections, calculation versions and approval states. Implement carbon and reporting rules with domain specialists, reconcile outputs and assess the effect of changes on historical results. Distinguish estimates, submitted evidence and accepted registry outcomes.

  • Improve identity, authorization, service permissions and secrets management across APIs, internal tools and automated workloads. Make sensitive data access and reviewer actions attributable to the right people and services.

  • Operate and improve our GCP and container infrastructure, including Cloud Run services, Dagster execution, PostgreSQL connectivity, storage and caching. Establish useful monitoring, alerts, recovery procedures and cost visibility.

  • Make development and releases repeatable through dependency management, automated tests, CI/CD, configuration management and infrastructure automation. Provide documented deployment and rollback paths.

  • Create operations and delivery that a small team can sustain. Establish automated checks, CI/CD, infrastructure configuration, useful monitoring, cost visibility and recovery procedures. Deliver migrations in bounded steps and document the system so another engineer can deploy, diagnose and recover it.

  • Effective AI-assisted engineering. Use coding assistants and agents for codebase exploration, implementation, refactoring, test development, documentation and debugging. Build repeatable ways to provide context, review changes and verify results while keeping ownership with the engineer.

What you bring

  • Experience owning production data systems through design, architectural change, migration and operation. You can explain the trade-offs and measurable outcomes of your decisions.

  • Strong Python and SQL, including maintainable interfaces, APIs, relational databases, warehouse transformations, numerical correctness and production debugging.

  • Strong data modeling: record grain, keys, temporal relationships, schema evolution, lineage, units and consistent metrics across heterogeneous sources.

  • Experience with orchestration, incremental processing, idempotency, backfills, reconciliation, monitoring and recovery. You understand how a successful pipeline can still produce incomplete or incorrect data.

  • Practical cloud engineering experience with managed compute, databases, object storage, event or queue services, identity and observability. GCP experience is valuable; comparable cloud experience is transferable.

  • Sound build-versus-buy judgment and experience making systems operable by a small team. You can justify both adopting a managed service and building a custom component.

  • Experience generalizing integrations or platform capabilities across multiple customers, sites, facilities or source systems, with configurable behavior and clear access boundaries.

  • The ability to prioritize, communicate architecture decisions clearly and work directly with plant operators, reviewers and domain specialists.

AI-assisted engineering requirements

We expect practical fluency with AI coding assistants and a disciplined approach to their limitations. You should be able to:

  • Turn an engineering task into a verifiable workflow. Supply relevant repository context, define constraints and acceptance criteria, break work into reviewable changes and use agents where they improve the outcome.

  • Evaluate generated code independently. Review architecture, dependencies, SQL, migrations and security implications. Test important behaviors and failure cases, inspect actual outputs and catch plausible but incorrect assumptions. Do not treat generated tests or an assistant's completion statement as proof of correctness.

  • Debug beyond the assistant's suggestions. Trace a problem through logs, data and source code; reproduce it; and explain the root cause. Maintain the Python, SQL and systems skills needed to resolve failures when AI assistance is ineffective.

  • Use tools with appropriate access. Protect credentials and sensitive operational data, follow approved tool and data-handling policies, and apply deliberate review to production changes, database writes and external submissions.

  • Make AI-assisted work maintainable. Produce focused diffs, meaningful tests, clear documentation and understandable code. Share reusable workflows with the team and assess their value through delivery time, defects, rework and maintenance effort.

We welcome experience with different coding assistants and agent tools; expertise in one vendor is not a prerequisite. This role does not require training foundation models. It does require the judgment to use AI effectively and remain accountable for the resulting system.

Useful additional experience

  • BigQuery or a comparable cloud warehouse; PostgreSQL, object storage, Pub/Sub or similar messaging, and managed application hosting.

  • Dagster, Airflow, Prefect or managed orchestration; dbt, Dataform or comparable transformation practices; Docker and infrastructure-as-code tools.

  • Industrial sensor data, intermittent connectivity, laboratory records, carbon accounting, lifecycle assessment or registry integrations.

  • OCR and LLM extraction pipelines, representative evaluation datasets, field-level accuracy measurement, model/prompt versioning and human review workflows.

  • Building practical analytical interfaces and operational dashboards, and helping nontechnical users interpret data-quality exceptions.

Our current stack is a starting point for informed decisions. You will help determine which components to retain, replace or obtain as managed services. Carbon-domain expertise is welcome and can be developed with our specialists.

How we will work together

You own the engineering implementation and technical reliability of the platform. Plant operations and reviewers own source-data verification; MRV specialists interpret and approve carbon methodology, eligibility rules and reporting decisions. You will make those responsibilities explicit in the system and work with their owners to resolve exceptions.

We will prioritize a bounded set of outcomes rather than expect every integration, dashboard and platform improvement at once. We will review engineering capacity as plant onboarding, support and reporting demand grows. Shared documentation, deployment practices and recovery exercises will help avoid dependence on one person.

Your first 90 days

We will agree scope and baseline measures together. Initial outcomes should include:

  • An architecture and migration plan: assess the existing platform, identify the highest-impact correctness and operational risks, and propose a staged path toward 10s of plants with assumptions to revisit before 100s. Include managed-service decisions and indicative costs.

  • One reusable source-to-reporting workflow: demonstrate it with representative data from more than one plant configuration, including data contracts, reconciliation, provenance, human review where needed and a safe historical reprocessing exercise.

  • A repeatable onboarding and operating path: document configuration, access, monitoring and recovery; measure manual engineering effort and demonstrate that another team member can follow the process.

  • A verified AI-assisted development workflow: show how repository context, implementation, review and meaningful checks fit together, with examples of catching and correcting assistant errors.

Success will be measured through data correctness, reporting traceability, reduced onboarding and operational effort, reliable recovery and sustainable cost. The first 90 days are a foundation for scaling, not a commitment to make the platform scale to 100s plants within that period.

Skills

Airflow, Google Cloud Platform, Platform Management, Data Engineering, Python, Data Architecture, Docker, Data Science, Dagster, BigQuery

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
1,356,297 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Data Science
Similar stack
Same company
Singapore
≈ $31k – $78k per year (Estimated) • In office • Full-Time • 5+ years exp • Bachelor's Degree • Sofia • Dublin
AI/ML
Machine Learning
Management
Agile
Apply
≈ $17k – $34k per year (Estimated) • In office • Full-Time • 5+ years exp • Bachelor's Degree • Mumbai
SQL
ABAP
Databases
Databricks
AI/ML
Hadoop
Airflow
DevOps
AWS
Analytics
Azure Data Factory
SAP BusinessObjects
Management
Agile
Apply
≈ $19k – $39k per year (Estimated) • In office • Full-Time • Bachelor's Degree • Bengaluru
Databases
Databricks
AI/ML
Hadoop
Airflow
DevOps
Azure DevOps
Azure
AWS
Analytics
Azure Data Factory
Management
Confluence
Jira
Agile
Scrum
Apply
≈ $47k – $108k per year (Estimated) • Remote (Italy) • Milan
Python
SQL
Databases
Db2
Databricks
Apache Kafka
Google BigQuery
BigQuery
AI/ML
Hadoop
Spark
Airflow
dbt
DALL-E
DevOps
Terraform
GCP
CI/CD
Git
Cybersecurity
GDPR
Analytics
ETL/ELT
Informatica
DataStage
Apply
Remote (location not specified) • Full-Time
C
C
MPI
AI/ML
CUDA Toolkit
Multimodal AI
PyTorch
CUDA
FSDP
NCCL
DevOps
Linux
Apply
≈ $7.5k – $15k per year (Estimated) • Hybrid • Full-Time • 4+ years exp • Bachelor's Degree • Chennai • Hyderabad
Python
SQL
Python
pySpark
bandit
AI/ML
Spark
Streamlit
Time Series Forecasting
Machine Learning
DevOps
GCP
AWS
Analytics
Tableau
Plotly
A/B Testing
Microsoft Excel
Apply
≈ $16k – $41k per year (Estimated) • In office • 5+ years exp • India
Python
JavaScript
TypeScript
SQL
Databases
PostgreSQL
Amazon Redshift
Amazon Aurora
AI/ML
Polars
Function Calling
AWS Bedrock
NumPy
LLM
Time Series Forecasting
LLM Guardrails
Tool Use
Frontend
GraphQL
D3.js
RxJS
Chart.js
Angular
React.js
Recharts
DevOps
Rest API
WebSockets
AWS
AWS Lambda
Amazon S3
IAM
Amazon CloudWatch
Amazon EventBridge
AWS Step Functions
API Gateway
Analytics
ETL/ELT
Apply
≈ $9k – $25k per year (Estimated) • In office • 3+ years exp • India
Python
JavaScript
SQL
Python
FastAPI
Databases
PostgreSQL
Frontend
React.js
React Query
AG Grid
DevOps
GitLab CI
CI/CD
AWS
Shift-Left
Cybersecurity
Shift-Left Security
QA
Cypress
Playwright
Gatling
Pytest
Vitest
k6
Apply
≈ $22k – $46k per year (Estimated) • In office • 5+ years exp • Bengaluru
Python
Python
Flask
FastAPI
Django
Databases
MySQL
Redis
RabbitMQ
Apache Kafka
Google BigQuery
BigQuery
AI/ML
LangChain
LlamaIndex
Embeddings
Scikit-learn
Prompt Engineering
AI Agents
TensorFlow
Pandas
NumPy
Keras
PyTorch
LLM
Context Engineering
Machine Learning
DevOps
GCP
GitHub Actions
CI/CD
AWS
Apply
≈ $21k – $46k per year (Estimated) • In office • 4+ years exp • Hyderabad
Python
JavaScript
TypeScript
Python
Flask
FastAPI
Django
AI/ML
LangGraph
LangChain
LlamaIndex
Fine-tuning
AI Agents
Semantic Kernel
LLM
RAG
Machine Learning
Frontend
Angular
React.js
DevOps
Rest API
Azure
CI/CD
AWS
Apply
≈ $62k – $125k per year (Estimated) • In office • Full-Time • PhD • Singapore
Java
C++
Scala
C++
PyTorch C++
Databases
Apache Iceberg
HBase
Delta Lake
Apache Hudi
AI/ML
Flink
PyTorch
Recommender Systems
Apply
≈ $57k – $104k per year (Estimated) • In office • 5+ years exp • Bachelor's Degree • Singapore
Databases
Google BigQuery
BigQuery
AI/ML
Machine Learning
DevOps
GCP
Kubernetes
Google GKE
Apply
≈ $55k – $116k per year (Estimated) • In office • 8+ years exp • Bachelor's Degree • Singapore
AI/ML
DeepSpeed
vLLM
Vertex AI
Fine-tuning
SGLang
TensorRT
TensorRT-LLM
Post-training
Pre-training
Megatron-LM
TPU
Edge AI
Machine Learning
DevOps
GCP
HPC
Apply
≈ $33k – $65k per year (Estimated) • In office • 2+ years exp • Bachelor's Degree • Singapore
Java
Kotlin
Dart
Mobile
Flutter
Apply
≈ $48k – $107k per year (Estimated) • In office • 5+ years exp • Bachelor's Degree • Singapore
Python
Go
Java
C++
Dart
AI/ML
AI Agents
Mobile
Flutter
Apply
See all jobs
This is one of many
1,356,297 more open roles from verified company boards, updated every day.