{"id":1385326,"url":"https://alion.io/job/bettermode-senior-platform-systems-engineer","title":"Senior Platform Systems Engineer","company":{"id":1818376,"name":"Bettermode","domain":"bettermode.com","url":"https://alion.io/company/bettermode-com","size_band":"11-50","is_staffing_agency":false,"employer_type":"direct","is_intermediary":false,"listed_via":null,"ats_vendor":"BambooHR","truth_index":{"grade":"B","score":75,"open_postings":6,"ghost_share":0,"stale_share":1,"repost_share":0,"time_to_fill_p50_days":null,"computed_at":"2026-10-01T05:45:00Z"}},"role":"Backend","role_family":"Backend","seniority":"senior","employment_type":"full_time","work_mode":"on_site","remote_scope":null,"remote_scope_basis":null,"remote_working_hours":null,"hiring_geo_confidence":"explicit","locations":["Toronto, Canada"],"countries":["CA"],"hiring_countries":[],"hiring_countries_total":0,"salary":{"min":160000,"max":180000,"currency":"AUD","period":"year","gross":null,"usd_annual":126485},"salary_estimate":null,"experience_years_min":null,"visa_sponsorship":false,"relocation_package":false,"has_equity":false,"technologies":[{"name":"Amazon Aurora","optional":false},{"name":"Amazon EKS","optional":false},{"name":"AWS","optional":false},{"name":"ClickHouse","optional":false},{"name":"Cloudflare","optional":false},{"name":"GDPR","optional":false},{"name":"GitOps","optional":false},{"name":"Go","optional":false},{"name":"gRPC","optional":false},{"name":"Helm","optional":false},{"name":"IAM","optional":false},{"name":"Kubernetes","optional":false},{"name":"Least Privilege","optional":false},{"name":"Linux","optional":false},{"name":"OpenTofu","optional":false},{"name":"OWASP","optional":false},{"name":"Platform Engineering","optional":false},{"name":"PostgreSQL","optional":false},{"name":"Rust","optional":false},{"name":"SOC 2","optional":false},{"name":"TCP/IP","optional":false},{"name":"Terraform","optional":false},{"name":"TypeScript","optional":false},{"name":"eBPF","optional":true},{"name":"KServe","optional":true},{"name":"Kubeflow","optional":true},{"name":"LLM","optional":true},{"name":"MLFlow","optional":true},{"name":"Monday.com","optional":true},{"name":"Webflow","optional":true}],"status":"live","first_seen_at":"2026-05-01T00:00:00Z","employer_posted_date":"2026-05-01","last_verified_at":"2026-09-30T16:25:31Z","board_verified":true,"closed_at":null,"days_open":153,"trust":{"level":"stale","repost_count":0,"flags":["stale"],"days_open":153},"description":"About Us\nAt Bettermode, we are redefining how businesses streamline customer experiences and foster strong relationships. Our platform empowers businesses to seamlessly craft powerful web apps with engagement tools in its core tailored to their unique needs.\nBacked by Silicon Valley investors and trusted by brands like Monday.com, Webflow, and Nubank, we’re proud to connect millions of end-users daily (check our Showcase page ).\nJoin us as we continue building tools that redefine customer engagement!\nBenefits\n At Bettermode, we’re dedicated to empowering our team to thrive-both professionally and personally. We offer location-based, competitive compensation that reflects your expertise and impact, with annual reviews so you can grow with us. Our culture is built on ownership and trust, giving you real influence over how we scale and succeed.\n From your very first day, you and your family are covered by comprehensive Canadian health benefits-dental and vision included-so you can focus on what matters most.\n Enjoy unlimited paid vacation days, paid parental leave to support your family, and bereavement leave should you need it.\n You’ll have all the equipment you need provided, or you can bring your own device and access our Device Upgrade Policy-an interest-free hardware stipend repayable via payroll deductions, allowing you to upgrade when you need.\n We want you to thrive in your work: every team member receives a monthly Tech & Appreciation Stipend-perfect for testing new software or tools and improving your workflows as you see fit.\n For in-person collaboration, our downtown Toronto office is less than a 15-minute walk from Union Station, with a free shuttle running throughout the day. The office features complimentary snacks, coffee, video games, and board games, as well as dedicated seating and a flexible environment that supports creativity, focus, and teamwork.\n Join a globally diverse and collaborative team where you’re supported to do your best work and have access to all the resources needed to succeed.\nAbout This Role\nEmployment Type:Full-time\nLocation:Canada\nLocation type:Remote or Hybrid (3 days at the office in Downtown Toronto - Monday, Tuesday and Wednesday, for employees residing within 40km of the company headquarters)\nTimezone: Eastern Standard Time\nThe Opportunity\nThis is not a generic DevOps role, not a narrow tool-operator role, and not a vendor-certified specialist role.\nYou will help shape foundational parts of Bettermode's platform across Kubernetes runtime architecture, Terraform-governed AWS and Cloudflare infrastructure, service-to-service networking, data-plane efficiency, OLAP analytics systems, databases such as Aurora PostgreSQL and MongoDB Atlas, cost visibility, security/compliance governance, and deployment architecture.\nThe role is intentionally broad across solutions architecture and systems programming: tune where appropriate, but build or redesign when necessary, with a strong emphasis on secure, recoverable, and well-governed platform foundations.\nOur operating model includes a production on-call program: engineers participate in an every-other-week rotation for P0 incident response, post-incident learning, and production ownership.\nResponsibilities\nDiagnose and remediate foundational platform problems across Kubernetes/EKS, Terraform-managed AWS and Cloudflare infrastructure, networking, observability, OLAP/data systems, security controls, and deployment architecture.\nOwn Kubernetes platform patterns and Terraform/OpenTofu workflows that make environments reproducible, reviewable, secure, and recoverable, including promotion, drift control, and policy-aware infrastructure changes.\nDesign AZ-aware and topology-aware improvements, starting with Aurora PostgreSQL routing/scalability and extending to other data-plane systems where traffic locality, availability, and cost matter.\nBuild cost and workload observability that attributes AWS infrastructure spend, network transfer, CPU/memory usage, and cross-AZ patterns to services, workloads, and teams.\nBuild production-grade platform components in Go, Rust, or TypeScript where appropriate, including Kubernetes controllers, Terraform plugins, telemetry collectors, bespoke proxies, and CLIs. Select the language based on SDK maturity, operational correctness, maintainability, and ecosystem fit.\nImplement platform security and compliance controls aligned with SOC 2, OWASP, GDPR, IAM least privilege, secrets handling, encryption, network segmentation, auditability, and data protection.\nSupport OLAP analytics infrastructure and the migration from Pinot to ClickHouse, with attention to ingestion topology, query performance, data correctness, cost, and operational safety.\nDesign safe rollout, resilience, and DR patterns, including canaries, bypass modes, fast rollback, degraded-mode operation, backup/restore workflows, failover procedures, RTO/RPO trade-offs, and incident playbooks.\nExample projects\nBuild an AZ-aware Aurora PostgreSQL routing/proxy component and topology-aware controls that reduce inter-AZ traffic, improve reader/writer behaviour, and provide predictable behaviour during scaling and failover.\nBuild workload-level cost and resource intelligence capabilities across AWS and Kubernetes, including attribution for network traffic, cross-AZ patterns, CPU/GPU/memory utilization, and other signals needed to improve infrastructure efficiency.\nRedesign fragile legacy deployment and infrastructure abstractions where the right answer is not more Helm or YAML, but stronger software-defined platform foundations.\nInvestigate and correct service-to-service transport pathologies, including HTTP/2 and gRPC enablement, multiplexing behaviour, socket starvation, connection skew, and load-distribution inefficiencies.\nWhat You Bring to the Team\nDeep production experience with Kubernetes/EKS, Terraform/OpenTofu, AWS, and Cloudflare, including secure deployment patterns, environment promotion, drift control, and operational ownership.\nStrong software engineering fundamentals building backend, infrastructure, or distributed systems in production, with systems instincts for concurrency, performance, failure modes, and operational correctness.\nProfessional experience with Go or Rust for production platform components such as Kubernetes controllers, Terraform providers, CLIs, proxies, telemetry collectors, reconcilers, or internal developer tooling; TypeScript is valuable for higher-level platform definition and GitOps tooling where appropriate.\nSolid understanding of Linux, TCP/IP, HTTP, HTTP/2, gRPC, connection behaviour under load, and service-to-service networking.\nStrong understanding of security and compliance-oriented platform engineering, including SOC 2 evidence, OWASP-aligned practices, GDPR-aware data handling, IAM boundaries, encryption, secrets management, and audit trails.\nPractical experience designing or operating DR capabilities, including backups, restore testing, RTO/RPO trade-offs, failover procedures, degraded-mode operation, and incident response playbooks.\nPractical familiarity with databases or OLAP/data infrastructure such as Aurora PostgreSQL, MongoDB Atlas, Pinot, ClickHouse, or similar systems, plus the ability to reason from first principles instead of relying only on vendor defaults.\nBonus if You\nHave built Kubernetes controllers, operators, reconcilers, or long-running platform agents in production.\nHave experience with service meshes, proxies, transport-aware systems, traffic steering, or network observability for cost-attribution.\nHave worked on cloud cost attribution, workload-level infrastructure observability, eBPF, VPC Flow Logs, or similar telemetry systems\nHave MLOps experience with KServe, Kubeflow, MLflow, model-serving infrastructure, GPU workloads, or other AI/ML platform systems.\nHave contributed to brownfield infrastructure migration, Terraform/OpenTofu import workflows, drift detection, policy-as-code, or infrastructure governance.\nHave experience replacing brittle YAML/Helm-heavy abstractions with typed platform tooling, GitOps generators, CDK-style infrastructure definitions, or internal developer platforms.\nIf you think this role is right for you, apply today! We’re excited to share more details, learn about your experience, and discover together if we’re the perfect fit for each other.\nCommitment to Diversity\nAs we continue to grow with customers and team members worldwide, we are committed to cultivating an environment where everyone’s unique perspectives are heard and valued. The diversity of our team will enable us to build the most inclusive product and workplace possible. We encourage applications from all backgrounds, identities, abilities, and life experiences.\nAdditional Information\nHeadcount: This is a vacancy at Bettermode.\nCompensation Range:CA$160K-$180K/annually for Canada-based candidates\n\nAI Use: Large language models (LLM) might be used in the hiring process for this position to screen, assess or select job applicants.","description_format":"text","description_chars":9016,"description_truncated":false,"requirements":{"experience_years_min":null,"management_years_min":null,"team_size_min":null,"manages_managers":false,"education":{"level":"phd","optional":false},"security_clearance":false,"languages":[]},"benefits":["Parental leave"],"hiring_locations":[{"name":"Canada","iso":"CA","kind":"country"}],"hiring_excludes":[],"relocation_offered":false,"industries":["Artificial Intelligence"],"lifecycle":[{"event":"open","at":"2026-09-28T10:08:33Z"}],"liveness":{"score":5,"band":"cold","label":"Long shot","p_open":1,"p_active":0.196,"p_room":0.28,"age_days":153,"expected_fill_days":30,"reasons":["conf:13","win:tail","crowd:"],"computed_at":"2026-10-01T05:45:00Z"},"pay":{"stated_usd_annual":126485,"is_top_pay":false},"html_url":"https://alion.io/job/bettermode-senior-platform-systems-engineer","json_url":"https://alion.io/job/bettermode-senior-platform-systems-engineer.json","meta":{"generated_at":"2026-10-01T21:48:46Z","cache_seconds":300,"methodology":"https://alion.io/methodology","terms":"https://alion.io/terms","contact":"https://alion.io/contact","api":"https://alion.io/developers","usage":{"tier":"crawler","counted_by":"address","units_charged":1,"used_today":4389,"day_limit":5000,"remaining_today":611,"minute_limit":60,"resets_at":"2026-10-02T00:00:00Z"}}}