743,556open jobs
44,602companies
107,147added this week
Browse all
Salary
$161k – $194k per year
Location
Remote (United States)
Seniority
Senior · 5+ years exp

Confirmed on the employer's own hiring board on Sep 24, 2026. First seen by Alion on Sep 24, 2026. Smartsheet scores B on the Alion truth index.

Overview
Company
Impact
Profile match
Smartsheet is an American work management company founded in 2005 that built a collaborative platform around the spreadsheet interface most business teams already knew rather than asking them to learn project management software. Its customers use it for project and portfolio management, resource planning, approvals and reporting across marketing, construction, professional services and operations teams. Headquartered in Bellevue, Washington, it was listed on the New York Stock Exchange until being taken private by Blackstone and Vista Equity Partners in early 2025.

For over 20 years, Smartsheet has empowered teams to manage work seamlessly and scale solutions smarter. Now, in our most ambitious chapter yet, we are uniting human teams with AI agents. By orchestrating the work agents do best, automating manual tasks and uncovering insights at scale, we create the space for people to focus on what truly matters: judgment, creativity, and big thinking. That is magic at work, and it’s what we show up for every day.

Smartsheet’s Observability Engineering team owns how the company sees itself: the collection, modeling, storage, and analysis of metrics, logs, distributed traces, and events across a global, multi-region infrastructure. As a Senior Software Engineer I on this team, you will be hands-on-keyboard building that platform (instrumentation libraries, telemetry pipelines, SLO and alerting systems, dashboards as code), and you will be the engineer who connects it to everything around it. Observability only pays off when signals flow into the systems where engineers actually work: CI/CD, the service catalog, incident management, ticketing, chat, and the automated remediation that closes the loop without waking anyone up.

We run observability as an internal product, and the engineers of Smartsheet are its users. That means published interfaces, versioned client libraries, honest deprecation paths, and SLOs on the platform itself. It also means we measure ourselves on adoption rather than on components shipped: a beautifully engineered tracing pipeline that three teams use is a failure. You will own a slice of that product end to end, including the golden path that makes correct instrumentation the easy choice, the documentation and hands-on sessions that get teams there, and the numbers that tell you which part to fix next.

This is a build role, not a configure role, and it is engineering first. You will write the services, the integrations, and the automation that turn a collection of separate tools into one coherent platform, and you will own the architecture decisions that make it hold up as it grows. Where tooling and teaching compete for the same problem, we prefer the tool: a CI check that rejects an unlabeled metric works for every team forever, while a workshop works for the people in the room. You will report to the Team Lead, Observability Engineering, and partner closely with the Principal Engineer setting telemetry data platform direction.

You Will

  • Architect and build the observability platform:  Design and ship end-to-end capability for metrics, logs, distributed traces, and events across multiple regions and environments, owning components from collection through storage, query, and presentation
  • Run observability as an internal product:  Build for the engineers who depend on you, with published interfaces, versioned client libraries, clear deprecation paths, and SLOs on the platform itself, so teams rely on it the way they rely on any production service
  • Drive instrumentation with OpenTelemetry:  Build and maintain shared instrumentation libraries, collector deployments, semantic conventions, context propagation, and sampling strategies so service teams get correlated signals by default rather than by effort
  • Engineer telemetry pipelines at scale:  Build high-volume collection, enrichment, redaction, and routing pipelines with the reliability, backpressure handling, and tiered retention that multi-region, high-cardinality traffic requires
  • Connect the platform end to end:  Integrate the observability toolchain with CI/CD, service catalog, incident management, ticketing, chat, and feature-flag systems through REST APIs, webhooks, and event-driven services, so that a deploy, an alert, an incident, and a ticket form one continuous thread instead of four disconnected ones
  • Build automated and self-healing remediation:  Design event-driven and agent-assisted workflows that detect, diagnose, and resolve known failure modes automatically, with human-in-the-loop approval gates as the safety mechanism for anything consequential
  • Make reliability measurable:  Implement SLOs, error budgets, and golden-signal alerting as code, and drive down alert noise so that a page means something is genuinely wrong
  • Build the golden path:  Own the paved road for instrumenting a new service, including scaffolding and templates, versioned Terraform modules, and CI checks that catch missing or malformed telemetry before merge. Make the correct path the easy path, so coverage comes from good defaults rather than from chasing teams
  • Instrument the platform itself:  Track coverage, onboarding time, time to first useful dashboard, query performance, and cost per service, and build the guardrails that keep cardinality, sampling, and retention proportional to the value of the signal. Let those numbers decide what you build next instead of guessing
  • Build the enablement layer:  Write documentation as code, reference architectures, and worked examples that scale past the conversations you can personally have, build the onboarding path that takes a team from zero to instrumented without a meeting, and run the workshops, office hours, and game days that exercise dashboards and alerts under realistic failure
  • Raise the technical bar:  Lead code reviews and architecture discussions, author the instrumentation standards other teams build against, mentor engineers on signal design and cost-aware instrumentation, and grow a group of instrumentation champions who carry the practice inside their own teams
  • Apply AI where it earns its place:  Use AI tooling to improve your own and the team’s efficiency across coding, testing, design, and troubleshooting, and help instrument Smartsheet’s AI and agentic systems so their behavior is as observable as any other service
  • Turn incidents into durable improvements:  Join the team’s on-call rotation, drive root-cause analysis, and close every incident with an instrumentation change, an automation, or a documented lesson that reaches the teams who need it

You Have

  • 5+ years building and operating highly scalable, highly available distributed systems, platform services, or observability infrastructure
  • 5+ years programming in Go, Python, Java/Kotlin, or TypeScript/Node.js, with the ability to move fluently between backend services and the operational tooling around them
  • Experience building internal platforms, developer tooling, or an internal developer platform (e.g., Backstage), with a product mindset toward internal users
  • Hands-on depth across all three primary signals (metrics, logs, and distributed tracing) on a major commercial or open-source observability platform, including practical understanding of cardinality, sampling, and cost mechanics, and experience defining SLOs and alerting for production systems
  • Practical OpenTelemetry experience: collectors, instrumentation (auto and manual), semantic conventions, and context propagation across service boundaries
  • Experience building high-volume data or telemetry pipelines: log and metric shippers, streaming transport, transformation and enrichment, and search or time-series backends
  • Strong REST API and integration engineering skills, including OAuth and other authentication patterns, webhooks, and event-driven architectures connecting third-party SaaS platforms
  • Advanced AWS and Kubernetes expertise (EKS, ECS, EC2, Lambda, and the managed messaging and eventing services that tie them together), delivered through Terraform and GitOps-based workflows
  • Demonstrated technical teaching: workshops, onboarding curricula, internal courses, conference or meetup talks, or a track record of raising a team’s capability. You write things down, and what you write gets used
  • Experience measuring and driving adoption of a platform or tool, and comfort treating low adoption as a product problem rather than a user problem
  • 6+ months of professional experience leveraging AI to enhance engineering productivity, with a view on where it helps and where it does not
  • A track record of leading large projects autonomously: decomposing ambiguous problems, planning the work, and carrying it to production
  • Computer Science degree, Engineering degree, or equivalent practical experience
  • Legal eligibility to work in the U.S. on an ongoing basis

Nice to Have

  • Workflow automation or orchestration experience (managed state machines, workflow engines, or RPA platforms) with human-in-the-loop approval patterns
  • Experience instrumenting LLM or agentic systems, including token, latency, quality, and cost telemetry
  • Experience migrating or consolidating overlapping observability tooling without a coverage gap in between, especially where the hard part was the people rather than the technology
  • Docs-as-code toolchains and internal developer portals, including the discipline of keeping generated reference material accurate as the platform changes
  • Community building or public speaking: conference talks, meetup organizing, or an internal guild you started and kept alive
  • Multi-region and data residency experience, including telemetry redaction and handling of sensitive data across jurisdictions
  • Regulated or government environment experience (FedRAMP, GovCloud) and the instrumentation constraints that come with it
  • Prometheus, Grafana, and Alertmanager, or comparable open-source observability tooling
  • Frontend experience building operational UIs, dashboards, or internal tools

Current US Perks & Benefits:

  • Employer subsidized medical/vision and dental coverage for full-time employees
  • 401k Match to help you save for your future (50% of your contribution up to the first 6% of your eligible pay)
  • Monthly stipend to support your work and productivity
  • Flexible Time Away Program, plus Sick Time Off
  • US employees are automatically covered under Smartsheet-sponsored life insurance, short-term, and long-term disability plans
  • US employees receive 12 paid holidays per year
  • Up to 24 weeks of Parental Leave
  • Personal paid Volunteer Day to support our community
  • Opportunities for professional growth and development including access to Udemy online courses
  • Company Funded Perks, including a counseling membership, local retail discounts, and your own personal Smartsheet account
  • Teleworking options from any registered location in the U.S. (role specific)

Smartsheet provides a competitive base salary range for roles that may be hired in different geographic areas we are licensed to operate our business from. Actual compensation is determined by several factors including, but not limited to, level of professional, educational experience, skills, and specific candidate location. In addition, this role will be eligible for a market competitive incentive opportunity.

US Base Salary Pay Range

$161,250—$193,750 USD

Get to Know Us:

At Smartsheet, your ideas are heard, your potential is supported, and your contributions have real impact. You’ll have the freedom to explore, push boundaries, and grow beyond your role. We welcome diverse perspectives and nontraditional paths-because we know that impact comes from individuals who care deeply and challenge thoughtfully. When you’re doing work that stretches you, excites you, and connects you to something bigger, that’s magic at work. Let’s build what’s next, together.

Equal Opportunity Employer:

Smartsheet is an Equal Opportunity (EEO) employer committed to fostering an inclusive environment with the best employees. It is our policy to provide equal employment opportunities to all qualified applicants in accordance with applicable laws in the US, UK, Australia, Germany, Costa Rica, Japan, Bulgaria, India, and Singapore. All qualified applicants will receive consideration without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, age, protected veteran or disabled status, or genetic information. 

If there are preparations we can make to help ensure you have a comfortable and positive interview experience, please let us know.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
743,556 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Backend
Similar stack
Same company
In your city
$206k – $286k per year • In office • 5+ years exp • Bachelor's Degree • Seattle
Python
Java
Kotlin
SQL
Scala
Kotlin
Mockito
AI/ML
Spark
Mobile
JUnit
Management
Stripe
QA
TestNG
Apply
$166k per year • Hybrid • Full-Time • Park
Java
Java
Maven
Spring Boot
Databases
Apache Kafka
DevOps
GitLab CI
CI/CD
Jenkins
AWS
Trunk-Based Development
Bitbucket
GitLab
Cybersecurity
SonarQube
Veracode
Management
Confluence
Jira
Apply
$170k – $200k per year • Equity • Remote (United States) • Full-Time • 7+ years exp • Seattle
DevOps
GCP
Azure
Apply
$87k – $157k per year • In office • Full-Time • 4+ years exp • Bachelor's Degree • Tewksbury
SQL
C++
C++
OpenGL
Qt
AI/ML
CUDA Toolkit
CUDA
DevOps
Red Hat
CI/CD
Linux
Management
Agile
Scrum
Apply
$108k – $195k per year • Hybrid • Public Trust • Full-Time • 8+ years exp • Bachelor's Degree • Gaithersburg
C++
Ada
AI/ML
Claude
AI Agents
OpenAI Codex
DevOps
Red Hat
CI/CD
Linux
TCP/IP
Management
Agile
Apply
≈ $139k – $230k per year (Estimated) • Remote (United States) • Full-Time • 5+ years exp • Bachelor's Degree • United States
Java
Databases
PostgreSQL
DevOps
Terraform
CI/CD
AWS
Kubernetes
Amazon EKS
Linux
Apply
Treasury Analyst 1 hour ago
≈ $44k – $119k per year (Estimated) • In office • Full-Time • London
Python
SQL
Apply
$166k per year • Hybrid • Full-Time • Park
Java
Java
Maven
Spring Boot
Databases
Apache Kafka
DevOps
GitLab CI
CI/CD
Jenkins
AWS
Trunk-Based Development
Bitbucket
GitLab
Cybersecurity
SonarQube
Veracode
Management
Confluence
Jira
Apply
$120k – $160k per year • Remote (United States) • Full-Time • High School Diploma • United States
Python
PowerShell
AI/ML
AI Agents
DevOps
CI/CD
Platform Engineering
IAM
Cybersecurity
Zero Trust
PKI
Management
ITSM
Apply
≈ $29k – $63k per year (Estimated) • Hybrid • Full-Time • 7+ years exp • Bachelor's Degree • Noida • Gurgaon
Python
SQL
Python
Flask
Django
Databases
MySQL
PostgreSQL
DevOps
Rest API
Terraform
CI/CD
Git
AWS
Docker
Kubernetes
AWS Lambda
Amazon EC2
Amazon S3
API Gateway
Management
Agile
Apply
$112k – $210k per year • Remote (United States) • 8+ years exp • Bachelor's Degree
AI/ML
AI Agents
DevOps
CI/CD
Management
Smartsheet
Apply
$175k – $245k per year • Remote (United States) • 7+ years exp • Bachelor's Degree
Java
Kotlin
TypeScript
Databases
MySQL
Snowflake
Databricks
DynamoDB
Apache Kafka
AI/ML
AI Agents
Flink
DevOps
Rest API
Azure
AWS
Kubernetes
Amazon ECS
Amazon Kinesis
Management
Smartsheet
Apply
$175k – $245k per year • Remote (United States) • 8+ years exp • Bachelor's Degree
Python
Databases
Databricks
AI/ML
MLFlow
Prompt Engineering
AI Agents
LLM
RAG
Hallucination
LLMOps
Context Engineering
Multi-Agent Systems
DevOps
GCP
Azure
AWS
Management
Smartsheet
Apply
$223k – $258k per year • Remote (United States) • Full-Time • 10+ years exp • Bachelor's Degree • Bellevue
Python
Go
Java
SQL
Scala
Databases
Snowflake
Databricks
Apache Iceberg
Delta Lake
Apache Kafka
AI/ML
Spark
Model Context Protocol
MLFlow
AI Agents
LLM
DevOps
Terraform
OpenTelemetry
Datadog
Fluent Bit
Fluentd
Prometheus
GitOps
ArgoCD
AWS
Kubernetes
Grafana
Amazon EKS
AWS Fargate
AWS Lambda
Amazon EC2
Alertmanager
FinOps
Amazon ECS
Amazon CloudWatch
Amazon Kinesis
Cybersecurity
FedRAMP
Management
Smartsheet
Apply
Remote (Bulgaria) • 3+ years exp
Java
Kotlin
Swift
Objective-C
Java
RxJava
AI/ML
AI Agents
Mobile
JUnit
Clean Architecture
Management
Smartsheet
Agile
Apply
See all jobs
This is one of many
743,556 more open roles from verified company boards, updated every day.