1,164,481open jobs
65,945companies
208,354added this week
Browse all
Salary
≈ $117k – $229k per year (Estimated)
Location
In office (United States)
Seniority
Senior · 5+ years exp

Confirmed on the employer's own hiring board on Oct 3, 2026. First seen by Alion on Oct 1, 2026. Sand Technologies scores A on the Alion truth index.

Overview
Company
Impact
Profile match
Sand Technologies (Sand Tech Holdings) is a global artificial intelligence and digital transformation company that partners with governments, cities, and enterprises to modernize critical infrastructure across healthcare, telecommunications, water, energy, and public systems. The company develops enterprise-grade data platforms, digital twin solutions, and specialized software - such as its Health Operating System (HOS) - to enable real-time operational decision-making in complex environments. Headquartered with operations spanning Africa, Europe, the UK, and the US, it also collaborates with initiative networks like ALX to train digital tech talent and deploy forward-deployed engineering teams worldwide.

Site Reliability Engineer

About Sand

Sand Technologies is a global Physical AI company using data and AI to make critical industries work better. We partner with governments, cities and enterprises to improve how essential systems operate across healthcare, water, energy, telecommunications and infrastructure.

Our work delivers proven real-world impact. We have built AI systems that help manage London’s water supply, supported telecom network planning across hundreds of cities, and developed digital healthcare platforms serving tens of millions of people across Africa. From intelligent command centers to AI-powered infrastructure platforms, we help organizations sense, analyze and act in complex environments.

Our people are ambitious, curious and relentlessly practical. Our teams work alongside clients in the field, solving hard problems and deploying solutions that last. With colleagues across Africa, Europe, the UK and the US, we operate across the full stack - from research and engineering to deployment and capability building.

Our mission is simple: to harness AI to solve humanity’s most pressing challenges.

Our platform

Sand builds the intelligent platforms behind critical infrastructure: water, healthcare, telecommunications, energy, for enterprises and governments worldwide. Our AI and data intelligence systems help manage London's water supply and power digital health platforms serving tens of millions of people. We have delivered successfully into over fifty client environments globally and are rapidly growing our footprint in the US. The successful applicant for this role will have an opportunity to grow with these expansion efforts, taking on significant technical and operational leadership responsibilities as we scale.

We are investing in a platform, Symmetri, which underpins and accelerates our client impact. It integrates heterogeneous operational data, models the entities and relationships of an entire domain, and lets operators of critical infrastructure understand the state of their world, reason about it, and act. It runs in multi-tenant cloud, in single-tenant customer environments, on premises, and is already deployed and in use in multiple fully air-gapped installations at ministry and utility level.

Our US customers are city governments, water utilities and public agencies who will run their operations on these systems and audit how we deploy and manage them.

About the role

Sand operates in partnership with senior executives at our customers. Your primary client within these relationships will be the customer's enterprise IT organization. You will spend as much time in front of a municipal CIO's security team, an architecture review board and a cloud governance committee as you will in a terminal. You will be asked where the data sits, who has access to it, how it is encrypted, and what happens during an incident, and you will need to answer those questions accurately and calmly without escalating every one of them.

Symmetri is designed to be highly modular and extensible, supporting rapid development and the mechanisms to support reusable intelligence capabilities across clients without sharing their data. As it continues to become the standard way we deliver, the line between deploying the platform and building it converges. You will be expected to see that coming and drive the strategy and execution that helps us get there.

Your role will be to lead and deliver the deployment and the operation of cloud infrastructure, including Symmetri environments, for US customers, end to end.

What you’ll do

  • Provisioning: Standing up infrastructure for a new environment, where the shape of the build depends on the environment type and variant, from managed cloud through single-tenant to on premises based on established patterns and practices.
  • Configuration: Tuning environments to a customer's needs as we land and expand on the use cases they are consuming: performance, scale, resource ceilings, network boundaries, identity, applying their security model rather than ours.
  • Lifecycle: Running the upgrade and maintenance cycles, and coordinating releases across live customer production estates without breaking the operations that depend on them.
  • Operation: Monitoring, alerting, and being first responder for critical production outages at any time of the day or night - which we see as an exceptional occurrence to be solved for, not the norm.
  • Collaboration & Growth: You will be the first US member of our Delivery Infra & Enablement function, but you will not need to work in isolation. You will collaborate directly with an established and experienced remote enablement team and regional operations counterparts in other geographies. This will assist you to establish a full US-based operations team and set the standard of how US deployment and operations actually works in practice.

Your first 6 months

  • You have taken a US customer environment from provisioning to production and you are the named owner of it.
  • Runbooks for that environment exist, are accurate, and someone other than you (or their agent) could follow them safely.
  • Monitoring and alerting are in place and tuned, so that a page means something and silence means something too.
  • You have been through at least one customer security review as our technical voice in the room.
  • You can say, with evidence, which parts of our global deployment approach need to work differently in the United States.

Who you are

  • 5 to 8 years in platform engineering, cloud infrastructure, DevOps or site reliability, operating production systems that mattered to someone.
  • Cloud infrastructure, hands on. Azure and AWS. You have provisioned and run production data infrastructure rather than only managing it, with a strong understanding of networking, IAM and secure data handling, and how to design effectively within a client’s existing landscape and governance constraints.
  • Networking you can reason about under pressure. VPCs and VNets, subnets, security groups, DNS, load balancing, private connectivity, and the patterns for getting services to talk to each other across a customer's internal and external boundaries.
  • Kubernetes and containers. Pods, deployments, services, namespaces. Comfortable in kubectl or any kube-api interface for inspection and troubleshooting. You do not need to have written an operator.
  • Infrastructure as code. You are comfortable with using IaC by default, from the start, with more than one toolchain. Demonstrable experience across one or more of Pulumi, Terragrunt, Terraform, OpenTofu, CDK or equivalents is required. Pulumi is used for Symmetri and experience with it is a strong plus. 
  • Python for automation. As an organisation with strong data science roots we are strongly biased to python for many coding use cases, so familiarity is important and experience is a plus.
  • Change management and branching discipline. Standard git workflows, and the willingness to follow and improve a team's existing strategy rather than your own.
  • Resource planning. Sizing compute, memory and storage against what a customer actually needs and what their environment can actually give you, including when that environment is on premises and finite.
  • Incident response. You have been on call. You know the difference between fixing an outage and fixing the cause of one, you are able to run effective root cause analysis processes that result in long term improvements.
  • Client-facing competence with enterprise IT. You can hold a technical conversation with a customer's infrastructure and security people, be trusted by them, and represent a commitment without over-promising.

You will not be equally strong across all of these aspects. Strong in most, familiar with the rest, unafraid of learning fast while being aware of and drawing in support for your current limitations.

Advantageous experience

  • Public sector, utility or other regulated operational environments, and the security review processes that come with them
  • Air-gapped, sovereign or disconnected deployments of data-intensive systems
  • Observability and monitoring stacks, and building the alerting rather than only responding to it
  • Serving ML models and LLM systems in production, including offline

Location

We are looking to hire on the East Coast in the United States, as this is where the majority of our clients are based.

How we work

This role has an on-call component and accountability. Real production systems, real operational consequences.

It also has less scaffolding than you might expect. We optimize first to fall in love with our clients' challenges and to deliver impactful solutions. This means you will not always receive neatly scoped work, and part of the job is creating that clarity yourself and sharing it. We are not looking for someone to invent everything from scratch or replicate exactly what they have done in the past either. 

We believe in strong opinions, loosely held. We have established patterns and approaches that work, built by teams that have run these environments for years, and the person who succeeds here is one who learns them properly first and then improves them while participating in raising our global standard.

Due to the highly collaborative and internationally distributed nature of our work, successful candidates must be comfortable operating in small teams while contributing to larger, globally coordinated efforts. A strong sense of ownership, self-motivation and discipline in maintaining clear and consistent communication through virtual collaboration tools and video conferencing is essential.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
1,164,481 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

DevOps
Similar stack
Same company
United States
≈ $77k – $157k per year (Estimated) • In office • 2+ years exp • Tampa
Databases
Apache Kafka
DevOps
Rest API
CI/CD
Management
Agile
Apply
≈ $101k – $196k per year (Estimated) • In office • 5+ years exp • Columbus
Python
Java
Java
Maven
Spring Boot
Spring MVC
Spring Cloud
Gradle
Databases
Apache Kafka
AI/ML
Flink
Anomaly Detection
Human-in-the-Loop
Machine Learning
DevOps
CI/CD
Git
AWS
Kubernetes
Bitbucket
Amazon EKS
AWS Fargate
Amazon S3
Amazon ECS
Amazon Kinesis
Management
Agile
Apply
≈ $84k – $171k per year (Estimated) • In office • 3+ years exp • Wilmington
AI/ML
AI Agents
DevOps
AWS
FinOps
Analytics
ETL/ELT
Apply
≈ $82k – $167k per year (Estimated) • In office • 4+ years exp • Jersey City
Python
SQL
Databases
PostgreSQL
AI/ML
Copilot
LLM
LLM Guardrails
DevOps
Terraform
GCP
CI/CD
AWS
Bitbucket
Management
Confluence
Jira
Agile
Apply
≈ $108k – $210k per year (Estimated) • In office • 3+ years exp • Bachelor's Degree • Milford
Design
SolidWorks
Management
Microsoft Office
Apply
≈ $30k – $59k per year (Estimated) • Remote (EAEU) • Moscow
Python
JavaScript
Java
Node JS
Bash
Groovy
Databases
Redis
MinIO
Apache Kafka
DevOps
Terraform
Ansible
Zabbix
Helm
Prometheus
VictoriaMetrics
CI/CD
GitOps
ArgoCD
Jenkins
Git
Docker
Kubernetes
Grafana
Bitbucket
Linux
Unix
Apply
≈ $93k – $184k per year (Estimated) • In office • TS/SCI • 4+ years exp • Bachelor's Degree • Palm Bay
Python
Verilog
SystemVerilog
VHDL
DevOps
Linux
Chips/EDA
Xilinx Vivado
Management
Agile
Apply
$127k – $236k per year • In office • Secret • 16+ years exp • Bachelor's Degree • Rochester
Python
Rust
C++
Apply
$36k per year • Hybrid • 5+ years exp • Saint Petersburg
Python
Python
Flask
FastAPI
Databases
PostgreSQL
Redis
PostGIS
RabbitMQ
DevOps
GitLab
Astra Linux
Apply
≈ $17k – $45k per year (Estimated) • Hybrid • Saint Petersburg
Python
Bash
DevOps
Zabbix
Grafana
Linux
Windows
DHCP
VPN
VLAN
Cybersecurity
Wireshark
Tcpdump
Apply
Platform Engineer 2 months ago
≈ $75k – $193k per year (Estimated) • In office • 5+ years exp • Bachelor's Degree • Rwanda
AI/ML
Physical AI
DevOps
Helm
OpenTelemetry
Prometheus
CI/CD
GitOps
ArgoCD
AWS
Kubernetes
Grafana
Platform Engineering
Amazon EKS
Apply
Platform Engineer 3 months ago
≈ $75k – $193k per year (Estimated) • In office • 5+ years exp • Bachelor's Degree • Rwanda
AI/ML
Physical AI
DevOps
Helm
OpenTelemetry
Prometheus
CI/CD
GitOps
ArgoCD
AWS
Kubernetes
Grafana
Platform Engineering
Amazon EKS
Apply
≈ $108k – $226k per year (Estimated) • In office • 3+ years exp • United States
Python
TypeScript
SQL
AI/ML
AI Agents
LLM
Physical AI
Machine Learning
DevOps
Platform Engineering
Apply
≈ $93k – $220k per year (Estimated) • In office • 5+ years exp • Bachelor's Degree • Nigeria
Databases
Databricks
AI/ML
Replit
Physical AI
DevOps
GCP
Azure
AWS
Robotics
Digital Twin
Analytics
ETL/ELT
Management
Miro
Agile
Apply
≈ $186k – $349k per year (Estimated) • In office • 10+ years exp • Bachelor's Degree • United States
AI/ML
Physical AI
Apply
≈ $96k – $217k per year (Estimated) • In office • Full-Time • 3+ years exp • Bachelor's Degree • United States
AI/ML
AI Agents
Apply
In office • Internship • United States
Analytics
Microsoft Excel
Apply
$59k – $81k per year • Hybrid • Full-Time • 2+ years exp • Bachelor's Degree • United States
Cybersecurity
HIPAA
Management
Outlook
Apply
$40k – $52k per year • Remote (United States, EST hours) • Full-Time • 1+ year exp • PhD • United States
Cybersecurity
HIPAA
Management
Outlook
Apply
Site Manager Solar 5 hours ago
≈ $94k – $179k per year (Estimated) • In office • 5+ years exp • United States
Apply
See all jobs
This is one of many
1,164,481 more open roles from verified company boards, updated every day.