This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Data Platform Engineer - AI Platform based in the Canada.
This is a hands-on data engineering role focused on building and operating critical infrastructure within a regulated, high-availability government cloud environment. You’ll work on the serving layer that transforms complex datasets into fast, reliable answers for mission-critical investigations. The position combines distributed systems, database performance, data pipeline reliability, production operations, and AI-assisted engineering practices. You’ll take ownership of infrastructure where reliability, compliance, and operational resilience are essential. The role offers close collaboration with data platform, product, and forward-deployed engineering teams in a distributed environment. You’ll be empowered to make technical decisions close to the systems you operate while solving challenging problems at speed. This opportunity is ideal for an engineer who enjoys autonomy, continuous learning, and meaningful work with real-world impact.
Accountabilities
Optimize data-serving infrastructure: Own performance tuning for the StarRocks-backed serving layer, using query profiling and AI-assisted analysis to identify bottlenecks and resolve slow query patterns before they affect customers.
Build reliable data pipelines: Develop, maintain, and harden pipelines supporting government cloud investigations, ensuring data infrastructure remains dependable, scalable, and compliant.
Strengthen operational resilience: Become an independent operator of the serving layer, reducing single points of failure and improving incident response coverage and recovery times.
Lead production troubleshooting: Investigate infrastructure and data issues using logs, monitoring, AI-assisted debugging, and systematic root-cause analysis, turning complex incidents into sustainable fixes.
Support compliance and security: Deliver infrastructure changes that satisfy demanding government cloud and regulatory requirements, including backup, retention, audit, and data-management needs.
Own infrastructure improvements end-to-end: Take responsibility for technical initiatives from investigation and design through implementation, testing, deployment, documentation, and ongoing operation.
Collaborate across teams: Work closely with Data Platform, Forward Deployed Engineering, Product, and other technical teams to maintain alignment between government cloud capabilities and broader platform requirements.
Contribute to engineering standards: Participate in sprint planning, asynchronous operational updates, incident retrospectives, documentation, and runbook improvements to continuously strengthen platform reliability.
Apply AI effectively: Use AI tools to accelerate repeatable workflows, technical research, debugging, code review, documentation, and problem solving while maintaining strong engineering judgment and quality standards.
U.S. citizenship is required due to government cloud data-access requirements.
Distributed systems experience: Hands-on experience operating distributed OLAP, analytical database, or serving-layer technologies such as StarRocks, Trino, ClickHouse, or similar platforms.
Performance engineering: Proven experience with query tuning, database performance optimization, scalability, and troubleshooting production workloads at scale.
Data platform reliability: Experience owning data pipelines, production infrastructure, reliability initiatives, and incident response in operationally demanding environments.
Production ownership: Comfortable taking on-call responsibilities and independently investigating unfamiliar infrastructure with minimal supervision.
AI fluency: Experience using tools such as Claude, Cursor, or comparable AI assistants to accelerate debugging, code exploration, code review, documentation, and technical research.
Strong problem solving: Ability to investigate complex technical problems, identify root causes, make sound tradeoffs, and implement durable solutions rather than temporary fixes.
Autonomous mindset: Comfortable taking ownership of ambiguous problems, learning unfamiliar systems quickly, and driving projects from initial investigation through production deployment.
Communication and collaboration: Strong written and verbal communication skills, with the ability to work effectively in distributed teams and communicate technical issues clearly through synchronous and asynchronous channels.
Adaptability: Comfortable operating in a fast-moving environment where priorities can change quickly and where urgency, accountability, and measurable outcomes are highly valued.
Engineering judgment: Demonstrated ability to balance speed with reliability, security, maintainability, and operational quality in production systems.
Remote-first work environment within the Canada.
- Opportunity to work on technically challenging distributed data infrastructure supporting high-availability government cloud environments.
- Significant ownership and autonomy, with engineering decisions made close to the systems being operated.
- Exposure to AI-assisted engineering practices and the opportunity to develop advanced AI fluency in day-to-day technical work.
- Collaboration with a distributed, cross-functional engineering organization spanning multiple teams and technical disciplines.
- Opportunities to develop expertise in OLAP systems, data platforms, cloud infrastructure, reliability, performance engineering, and compliance-focused environments.
- A mission-driven environment focused on solving complex problems with meaningful real-world impact.
- Professional growth through hands-on ownership, technical experimentation, incident retrospectives, and continuous improvement.

