1,169,581open jobs
66,097companies
208,043added this week
Browse all
Salary
≈ $112k – $205k per year (Estimated)
Location
Remote (United States)
Seniority
Senior · 8+ years exp
Employment
Full-Time

Confirmed on the employer's own hiring board on Oct 4, 2026. First seen by Alion on Sep 24, 2026.

Overview
Company
Impact
Profile match
HCSS is a Texas company founded in 1986 that builds software for heavy civil construction. Its products cover estimating, project management, fleet maintenance and safety for contractors. Thousands of infrastructure builders use its systems.

We are HCSS. For the last 40 years, we have been developing software to help construction companies streamline their operations.Based in Sugar Land, TX, our mission is helping customers achieve excellence through our proven customer-centric, end-to-end solutions and exceptionally helpful service, while providing a great life for our employees. With this mission at the core of everything we do, HCSS is a pioneer and leader in the construction software space and a consistently recognized employer. We have earned Best Companies to Work for in Texashonors for 18consecutive years and have been named a USA Today Top Workplace. HCSS has also been recognized by Built Inas a Best Place to Work in Greater Houstonand by Construction Executivefor our technology innovation, reflecting our strong culture, industry leadership, and commitment to excellence.

WHO WE NEED:

As a Senior DevOps Engineer with a focus on Site Reliability Engineering (SRE), you will play a key role in driving infrastructure resilience, availability, and operational excellence across our cloud environments. A core responsibility of this role is managing and optimizing Azure SQL Elastic Pools, ensuring performance, cost-efficiency, observability, and automation are aligned with business objectives. You will lead initiatives around high availability, disaster recovery, incident response, and reliability automation. Your expertise in Azure or AWS, observability tools such as Grafana, and scalable infrastructure will be essential to ensuring our systems are robust, performant, and recoverable.

Qualifications:

  • 8+ years of experience in DevOps or SRE roles with a strong focus on cloud infrastructure and systems reliability
  • 3+ years of hands-on experience with managing Azure SQL Elastic Pools, including performance tuning, scaling, and automation
  • 5+ years of expertise in Azure cloud services including networking, compute, databases, and identity
  • 3+ years of experience applying SRE principles including SLIs, SLOs, and incident management best practices
  • Extensive experience with Infrastructure as Code tools such as Terraform, Bicep, or ARM templates
  • Extensive experience building and managing CI/CD pipelines with tools like Azure DevOps or GitHub Actions
  • Strong scripting skills using Azure CLI and PowerShell for automation and operational tasks
  • Experience with monitoring and observability platforms, ideally Grafana, or a strong foundation in similar tools

Soft Skills:

  • Strong troubleshooting and problem-solving abilities.
  • Excellent communication skills and a collaborative mindset to work with cross-functional teams.
  • Ability to work independently, manage multiple tasks, and prioritize efficiently.
  • A proactive attitude toward continuous improvement and learning.

Preferred Qualifications:

  • Managed 10+ Elastic pools and 100+ databases in Azure
  • Advanced level certifications on cloud infrastructure like Az-400 or equivalent

Role Responsibilities:

Azure Elastic Pool Management:

  • Take ownership of the design, scaling, and optimization of Azure SQL Elastic Pools
  • Monitor and tune pool performance to ensure efficiency and SLA compliance
  • Establish observability and alerting for SQL resource consumption, errors, and performance anomalies
  • Automate provisioning, scaling, and failover using infrastructure and scripting tools
  • Collaborate with database and application teams to align on resource usage strategies

High Availability and Disaster Recovery:

  • Design and implement highly available and fault tolerant systems
  • Develop and maintain disaster recovery strategies across critical services
  • Perform regular failover testing, documentation, and validation of recovery procedures
  • Work closely with infrastructure and development teams to ensure business continuity objectives are met

Monitoring, Observability, and Automation:

  • Implement and manage observability stacks with logs, metrics, traces, and alerting
  • Create dashboards and alerts in Grafana or similar platforms to track key system indicators
  • Develop automated solutions for provisioning, monitoring, and maintenance tasks
  • Continuously improve system visibility and reduce time to detect and resolve issues
  • Assist in the automation of performance testing to proactively identify bottlenecks, validate scalability and ensure reliable system behavior

Incident Management and Operational Excellence:

  • Establish and refine incident response processes, escalation workflows, and resolution protocols
  • Lead root cause analysis, post-incident reviews, and continuous improvement efforts, including following up to ensure identified improvements to the application or process are implemented.
  • Develop and maintain runbooks, diagnostic tools, and automated remediation solutions
  • Champion a blameless culture of reliability and operational readiness across engineering teams

Cloud Infrastructure Management:

  • Architect and manage scalable and secure cloud infrastructure in Azure or AWS
  • Provision and manage services including compute, networking, storage, and containerized workloads
  • Continuously monitor performance, latency, and uptime to ensure system health
  • Apply cost optimization practices while aligning infrastructure with business goals

Infrastructure as Code (IaC):

  • Define and implement infrastructure using tools such as Terraform
  • Maintain modular, version-controlled infrastructure code that supports environment consistency
  • Apply automation and policy enforcement to reduce drift and improve auditability
  • Ensure IaC best practices are embedded in the development lifecycle

Collaboration and Leadership:

  • Mentor and support junior DevOps engineers through code reviews, knowledge sharing, and technical guidance
  • Partner with cross-functional teams including development, security, and database operations to drive initiatives
  • Act as a subject matter expert in SRE practices and reliability-driven engineering
  • Lead continuous improvement efforts across infrastructure and operations processes

Travel Requirements:

  • Remote Requirements
    • Employees will be expected to come into the office on a periodic basis.
    • Baseline expectations for roles are as follows but may fluctuate based on manager’s discretion:
      • Individual Contributor - Up to 2x per year
  • Employees will be expected to attend HCSS sponsored events per manager discretion (ex. UGM)

BENEFITS & PERKS:

Part of our mission is to provide a great life for our employees. We believe that when our people are happy, they do their best work. Some of the benefits and perks we offer include:

  • Flexibility to work Remotely
  • Medical, dental, and vision coverage with company-paid and employee-paid options
  • Paid holidays, sick days, and personal time off
  • Employee Resource Groups (ERGs) that foster connection and inclusion
  • On-site amenities including a covered basketball court, soccer field, track, pickleball/tennis courts, gym, etc.
  • Dog-friendly campus and WiFi-accessible courtyards
  • 401(k) with a 5% company match
  • Coverage for employee professional development and wellness
  • And more!
Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
1,169,581 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

DevOps
Similar stack
Same company
In your city
≈ $120k – $219k per year (Estimated) • Remote (United States) • Part-Time • United States
Python
DevOps
Terraform
Helm
OpenTelemetry
Datadog
Prometheus
GitOps
AWS
Kubernetes
Grafana
Amazon EKS
Error Budget
SLI/SLO/SLA
IAM
Amazon CloudWatch
Cybersecurity
SOC 2
Web3
Staking
Apply
System Engineer 4 years ago
≈ $56k – $156k per year (Estimated) • Remote (United States) • 3+ years exp • Bachelor's Degree
Python
PowerShell
DevOps
Puppet
Ansible
Zabbix
Kibana
Chef
Grafana
Configuration Management
Nagios
Windows
VLAN
BGP
OSPF
Cybersecurity
Active Directory
Management
Jira
ServiceNow
ITSM
Apply
Remote (Portugal) • Full-Time • 3+ years exp • Lisbon
DevOps
Windows Server
Windows
DNS
DHCP
Cybersecurity
Active Directory
Apply
≈ $98k – $191k per year (Estimated) • In office • Full-Time • 5+ years exp • Bachelor's Degree • Malvern
Python
PowerShell
Databases
Microsoft Fabric
AI/ML
Claude
Model Context Protocol
Embeddings
RAG
OpenAI
Anthropic
Human-in-the-Loop
LLM Guardrails
Copilot Studio
DevOps
Rest API
Terraform
Azure DevOps
GitHub Actions
Azure
CI/CD
Kubernetes
Bicep
Azure AKS
Cybersecurity
Zero Trust
Microsoft Entra ID
Analytics
Azure Data Factory
Management
ServiceNow
Power Automate
SharePoint
Apply
$117k – $130k per year • In office • Full-Time • 7+ years exp • Bachelor's Degree • Sacramento
Python
Java
SQL
DevOps
AWS
Platform Engineering
Analytics
ETL/ELT
Apply
$160k – $180k per year • In office • Bachelor's Degree • Phoenix
Python
JavaScript
TypeScript
SQL
Databases
PostgreSQL
Mobile
JUnit
DevOps
Rest API
Azure DevOps
GitHub Actions
Azure
CI/CD
Jenkins
Git
SOAP
Management
Jira
ServiceNow
Agile
Scrum
QA
TestNG
Selenium
Cypress
Playwright
Postman
Rest-Assured
Pytest
Apply
$160k – $240k per year • In office • 5+ years exp • Bachelor's Degree • Phoenix
SQL
DevOps
Azure DevOps
Azure
Analytics
Tableau
Power BI
Management
Confluence
Microsoft Project
Jira
SharePoint
Agile
Waterfall
Apply
AI Native Engineer 7 hours ago
$140k – $200k per year • In office • United States
Python
JavaScript
TypeScript
Node JS
Databases
DynamoDB
AI/ML
Claude Code
Model Context Protocol
AI Agents
AWS Bedrock
Frontend
React.js
DevOps
Terraform
Azure DevOps
Azure
AWS
AWS Lambda
Amazon S3
AWS Step Functions
API Gateway
QA
Playwright
Apply
≈ $47k – $105k per year (Estimated) • In office • 3+ years exp • Bachelor's Degree • 6th of October City
SQL
Analytics
SAP BusinessObjects
Apply
≈ $81k – $180k per year (Estimated) • In office • 6+ years exp • Bachelor's Degree • 6th of October City
Python
Java
SQL
C#
Databases
HBase
AI/ML
Hadoop
Analytics
ETL/ELT
Erwin
Apply
≈ $36k – $63k per year (Estimated) • Hybrid • Part-Time • Sugar Land
DevOps
Wi-Fi
Apply
≈ $51k – $98k per year (Estimated) • Hybrid • Full-Time • Bachelor's Degree • Sugar Land
DevOps
Wi-Fi
Analytics
Power BI
Management
Trello
Confluence
Jira
Google Docs
Agile
Apply
≈ $73k – $159k per year (Estimated) • Remote (United States) • Full-Time • 4+ years exp
DevOps
Wi-Fi
Marketing
Salesforce
Apply
≈ $104k – $214k per year (Estimated) • In office • Full-Time • 3+ years exp • Bachelor's Degree • United States
DevOps
Wi-Fi
Management
Agile
Apply
≈ $93k – $210k per year (Estimated) • Hybrid • Full-Time • 7+ years exp • Sugar Land
DevOps
Wi-Fi
Apply
See all jobs
This is one of many
1,169,581 more open roles from verified company boards, updated every day.