408,465open jobs
14,167companies
73,523added this week
Browse all
Salary
$190k – $230k per year
Location
In office (Seattle)
Employment
Full-Time
Overview
Company
Impact
Profile match
One platform to orchestrate durable AI workflows, real-time apps, and compute optimization. Your data never leaves your infrastructure.

About Us

At Union, we are solving one of the hardest challenges in AI infrastructure today: enabling high-velocity iteration while maintaining seamless production-readiness for AI workloads at scale.

Flyte, the open-source project we steward, has emerged as the modern standard for data and AI orchestration, and is trusted by leading technology organizations including LinkedIn, Stripe, and Wayve to run millions of mission-critical workflows on the platform. These workflows comprise data preparation, model training, and scaled inference spanning thousands of GPUs, all major clouds, and on-premise infrastructure.

We have a technical founding team who created Flyte while at Lyft, a deep bench of infrastructure experts from top companies, and have raised from top investors like NEA and Nava Ventures.

About the Role

We are hiring a Systems Development Engineer to improve the reliability, operability, and customer experience of our production platform. This is a 50/50 operations and engineering role. Part of your time will be spent investigating customer-impacting production issues, and part will be spent building the tools, automation, system design, and engineering practices that prevent those issues from recurring.

For this role, production is the customer. You will work from real operational signals: customer issues, incidents, on-call pages, recurring support patterns, and gaps in observability or automation. You will sit in engineering and partner with customer-facing teams to turn those signals into durable platform improvements.

This is not a traditional support role. It is a systems engineering role for someone who can debug deeply, communicate clearly, and write software that reduces operational load. The full engineering team backs you on on-call.

This role is hybrid, based out of our Seattle office.

What You'll Do

  • Investigate and resolve customer-impacting production issues across cloud infrastructure, workflow execution, access control, storage, networking, deployment systems, and observability.

  • Identify patterns in customer issues and convert them into automation, product improvements, runbooks, tests, or design changes.

  • Build internal tools and diagnostics that make production issues easier to detect, understand, and resolve.

  • Improve platform observability, including logs, metrics, dashboards, alerts, and customer-visible debugging information.

  • Participate in design and development so systems are easier to operate, debug, and support, and guide engineering teams toward durable fixes.

  • Define and uphold operational engineering practices: production readiness, alert quality, runbook discipline, observability standards, regression prevention, and code quality.

  • Drive measurable reductions in on-call pages, recurring customer issues, manual operational work, and time-to-resolution.

What We're Looking For

  • Strong software engineering skills in Python, Go, Java, or a similar language.

  • Experience debugging production systems across multiple layers of the stack.

  • Practical knowledge of Kubernetes, Linux, cloud infrastructure, distributed systems, networking, storage, and IAM.

  • Experience with infrastructure-as-code, deployment systems, CI/CD, observability, and operational automation.

  • Ability to move from ambiguous customer symptoms to clear technical diagnosis and durable remediation.

  • Strong judgment about when to fix directly, automate, escalate, redesign, or build a broader platform improvement.

  • Clear written and verbal communication, especially around root cause analysis, technical recommendations, runbooks, and design feedback.

  • A bias toward reducing toil through engineering rather than repeatedly solving the same issue by hand.

Preferred Experience

  • Operating customer-facing SaaS, cloud infrastructure, self-hosted or on-prem deployments, or workflow orchestration systems.

  • Batch workloads, autoscaling, capacity management, identity and access systems, storage systems, or platform observability.

  • Improving on-call health, reducing ticket volume, or building production diagnostics.

  • Working across support, customer success, product, and engineering teams.

Success Looks Like

  • Customer issues are diagnosed faster and recur less often.

  • Engineering teams receive actionable feedback from production and customer pain.

  • Common operational problems become automated workflows, better diagnostics, clearer runbooks, or product fixes.

  • On-call pages trend toward roughly one per month.

  • The platform becomes easier to operate, easier to debug, and safer to change.

  • Customers experience fewer production surprises and faster resolution when issues do happen.

Benefits & Belonging

At Union.ai we know that employees who feel their best can build amazing things and we are proud to offer best in class benefits that will continually evolve and grow as the needs of our employees do. Benefits may vary based on country.

  • Excellent medical - We pay 100% of your premiums and 90% for your dependents

  • Generous dental and vision plans- We pay 90% of the premiums for you and your dependents

  • Meaningful equity in the form of options - all employees are owners here

  • Unlimited time off + 12 company holidays

  • 401K match - Union.ai matches 100% of contributions up to the first 3%, and 50% up to 5%

  • 12 weeks paid parental leave for primary and secondary caregivers

  • Flexible work schedule (some restrictions apply)

  • For in office employees: Lunch provided onsite and well stocked kitchen with snacks and drinks.

We believe that our differences are what bring us together to achieve truly special outcomes. We strive to be inclusive and focus on building teams that embody that quality too. Union.ai is an equal-opportunity employer and we encourage you to apply, even if your experience doesn’t align exactly with our job description.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
408,465 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
Seattle
$148k – $265k per year • Equity • In office • Full-Time • 7+ years exp • United States
Java
Node JS
Python
TypeScript
JavaScript
Databases
Apache Kafka
DynamoDB
Kafka
DevOps
Amazon CloudWatch
Amazon EC2
Amazon EKS
Amazon Kinesis
Amazon S3
API Gateway
ArgoCD
AWS
AWS CDK
AWS Lambda
CI/CD
FinOps
GitHub
GitHub Actions
GitOps
Helm
Incident Management
Jenkins
Kubernetes
Platform Engineering
Terraform
Cybersecurity
SOC 2
Apply
In office • 12+ years exp • Bachelor's Degree
Python
TypeScript
JavaScript
Databases
BigQuery
Databricks
Google BigQuery
Frontend
Angular
React.js
DevOps
Amazon CloudWatch
AWS
CI/CD
Datadog
Docker
GCP
Grafana
Kubernetes
Terraform
Management
Confluence
Jira
Apply
$17k – $38k per year (Estimated) • In office • Full-Time • 6+ years exp • Bachelor's Degree • Gurgaon • Hyderabad
Java
Databases
MySQL
DevOps
Bitbucket
Git
IAM
Management
Confluence
Jira
ServiceNow
Apply
In office • Full-Time • 15+ years exp • Doha
DevOps
Azure
FinOps
GCP
Kubernetes
OpenShift
Service Mesh
Apply
In office • 3+ years exp
JavaScript
Python
SQL
Apex
Python
FastAPI
Apex
MuleSoft
Databases
Snowflake
Frontend
React.js
DevOps
Amazon ECS
AWS
CI/CD
GitHub
GitHub Actions
Terraform
Apply
$215k – $240k per year • In office • Full-Time • 5+ years exp • Seattle
Python
Rust
AI/ML
Flyte
OpenAI
DevOps
ArgoCD
AWS
Azure
Buildkite
CI/CD
Docker
GCP
Helm
Kubernetes
Terraform
Management
Stripe
Apply
$170k – $190k per year • In office • Full-Time • 3+ years exp • Bachelor's Degree • Seattle
TypeScript
JavaScript
AI/ML
Amazon SageMaker
Flyte
Ray
Frontend
Headless UI
Next.js
React.js
Tailwind CSS
Vite
Webpack
Zustand
DevOps
gRPC
Kubernetes
Management
Stripe
Apply
Founding CX Lead 9 hours ago
$100k – $140k per year • Equity 0.1–0.3% • In office • Full-Time • 3+ years exp • Seattle
SQL
DevOps
Datadog
Apply
$140k – $200k per year • In office • Full-Time • 1+ year exp • Seattle
Apply
$55k – $82k per year • In office • 5+ years exp • Bachelor's Degree • Seattle
Apply
$63k – $94k per year • In office • 2+ years exp • Bachelor's Degree • Seattle
Management
Outlook
Apply
$140k – $157k per year • In office • PhD • Seattle
Apply
See all jobs
This is one of many
408,465 more open roles from verified company boards, updated every day.