{"id":1667269,"url":"https://alion.io/job/ercot-senior-systems-reliability-engineer","title":"Senior Systems Reliability Engineer","company":{"id":6006,"name":"Ercot","domain":"ercot.com","url":"https://alion.io/company/ercot","size_band":"1001-5000","is_staffing_agency":false,"employer_type":"direct","is_intermediary":false,"listed_via":null,"ats_vendor":"Workday","truth_index":{"grade":"B","score":75,"open_postings":4,"ghost_share":0,"stale_share":1,"repost_share":0,"time_to_fill_p50_days":null,"computed_at":"2026-10-09T06:01:00Z"}},"role":"Backend","role_family":"Backend","seniority":"senior","employment_type":"full_time","work_mode":"hybrid","remote_scope":null,"remote_scope_basis":null,"remote_working_hours":null,"hiring_geo_confidence":"structured","locations":["Taylor, United States"],"countries":["US"],"hiring_countries":[],"hiring_countries_total":0,"salary":{"min":109000,"max":150000,"currency":"USD","period":"year","gross":null,"usd_annual":150000},"salary_estimate":null,"experience_years_min":5,"visa_sponsorship":false,"relocation_package":false,"has_equity":false,"technologies":[{"name":"ActiveMQ","optional":false},{"name":"Ansible","optional":false},{"name":"Apache Kafka","optional":false},{"name":"Azure","optional":false},{"name":"Bash","optional":false},{"name":"Chaos Engineering","optional":false},{"name":"CI/CD","optional":false},{"name":"Datadog","optional":false},{"name":"Dynatrace","optional":false},{"name":"Error Budget","optional":false},{"name":"Grafana","optional":false},{"name":"ITIL","optional":false},{"name":"Java","optional":false},{"name":"Kubernetes","optional":false},{"name":"Linux","optional":false},{"name":"Loki","optional":false},{"name":"Micrometer","optional":false},{"name":"Mimir","optional":false},{"name":"OpenShift","optional":false},{"name":"OpenTelemetry","optional":false},{"name":"Oracle","optional":false},{"name":"Platform Engineering","optional":false},{"name":"PostgreSQL","optional":false},{"name":"Progressive Delivery","optional":false},{"name":"Python","optional":false},{"name":"Rest API","optional":false},{"name":"Self-Healing","optional":false},{"name":"SLI/SLO/SLA","optional":false},{"name":"SOAP","optional":false},{"name":"Splunk","optional":false},{"name":"Spring Boot","optional":false},{"name":"Terraform","optional":false}],"status":"live","first_seen_at":"2026-08-06T00:00:00Z","employer_posted_date":"2026-08-06","last_verified_at":"2026-10-09T22:47:32Z","board_verified":true,"closed_at":null,"days_open":65,"trust":{"level":"ok","repost_count":null,"flags":[],"days_open":65},"description":"At ERCOT, our diverse and dynamic work environment provides a platform on which employees can work together to build the future of the Texas power grid and wholesale market utilizing the latest technologies and resources. We encourage you to join our talented, dedicated workforce to develop world-class solutions for today and tomorrow’s energy challenges while learning new skills and growing your career.\nERCOT is committed to fostering inclusion at alllevels of our company. It is the cornerstone of our corporate values of accountability, leadership, innovation, trust, and expertise. We know that individuals with a wide variety of talents, ideas, and experiences propel the innovation that drives our success. An inclusive and diverseworkforce strengthens us and allows for a collaborative environment to solve the challenges that face our industry today and in the future.\nJOB SUMMARY\nThe Senior Systems Reliability Engineer applies software engineering discipline to reliability problems - designing, building, and operating the systems that make production software measurable, scalable, and self-healing. This role treats operational challenges as engineering problems: when a process is manual, it gets automated; when a failure mode is unknown, it gets instrumented; when a system degrades, the degradation is understood before it recurs.\nAt this level, the specialist owns SLO and error budget frameworks for assigned systems, architects the observability stack that the team relies on, leads engineering-driven incident response, and holds NERC/CIP compliance responsibility for assigned systems. This role partners directly with Software Engineers as a technical peer - participating in design reviews, influencing architecture decisions for reliability, and building the production readiness standards that govern how software ships. Advancement to Lead is based on demonstrated ability to define reliability engineering standards at the platform level, influencing practice across multiple teams and portfolios.\nJOB DUTIES\nPerforms complex reliability engineering work autonomously; recognized subject matter expert within the team and adjacent teams.\nDesigns and builds production software systems, reliability tooling, and automation frameworks; treats operational problems as engineering problems to be solved through code.\nOwns SLO governance, error budget management, and observability architecture for assigned systems; leads engineering-driven incident response including failover scenarios.\nHolds NERC/CIP compliance responsibility for assigned systems; formally mentors less experienced specialists; may coordinate team delivery and on-call activities.\nADDITIONAL JOB DUTIES\nCore Expectations\nThe following expectations apply at all Systems Reliability Specialist levels. Scope and independence expand with each level.\nEngineer reliability solutions: when a process is manual and repeatable, automate it; when a failure mode is opaque, instrument it; when a system is fragile, redesign the failure boundary.\nDefine and own SLIs and SLOs for assigned systems; treat error budgets as a shared engineering contract with development teams, not an operations metric.\nRespond to production incidents as an engineer: form a hypothesis, isolate the failure, resolve it, and close the loop with a post-mortem that addresses root cause.\nInstrument systems so that on-call responders have sufficient telemetry to diagnose and act without tribal knowledge.\nParticipate in 24/7 on-call rotation; treat every alert as signal - either actionable or worth eliminating.\nWrite production-quality code: reliability tooling, automation frameworks, and operational software are held to the same engineering standards as application code.\nPartner with development teams as a peer in design reviews; reliability is designed in, not bolted on after deployment.\nReliability Engineering\nSenior specialists design and build the engineering systems that make production software reliable. This is software engineering applied to operational problems - the output is code, frameworks, and automated systems, not tickets and runbooks alone.\nDesign, build, and maintain reliability tooling: automated remediation systems, self-healing infrastructure components, and operational software that reduces human intervention in production.\nOwn SLO and error budget definitions for assigned systems; review error budget consumption with development teams and drive engineering decisions based on budget status.\nArchitect and implement chaos engineering programs: define failure injection scenarios, automate resilience tests, and validate recovery behavior against defined SLOs.\nBuild and maintain CI/CD reliability gates: automated canary analysis, progressive delivery validation, and rollback triggers based on SLI thresholds.\nDesign capacity planning models for assigned systems; build tooling to project resource needs and surface capacity risks before they affect availability.\nContribute to production readiness reviews: define and enforce the engineering criteria that a system must meet before it ships to production.\nReduce operational toil through engineering: measure toil, track reduction targets, and build the automation that eliminates it.\nObservability & Instrumentation\nObservability is an engineering discipline. Senior specialists design and build the telemetry systems that make production behavior understandable - not just monitored.\nArchitect MLTP (Metrics, Logs, Traces, Profiling) observability solutions using the Grafana LGTM stack (Loki, Grafana, Tempo, Mimir), Dynatrace APM, Splunk, and Datadog.\nDefine and enforce instrumentation standards: structured logging schemas, metric naming conventions, trace context propagation, and continuous profiling configuration for assigned systems.\nBuild distributed tracing coverage across service boundaries; identify and close observability gaps that produce blind spots during incidents.\nDesign SLI instrumentation: translate user-facing reliability requirements into specific, measurable signals that accurately represent system health from the user's perspective.\nBuild and maintain alerting frameworks: alerts must be actionable, calibrated to SLO burn rate, and free of noise; own alert quality as an engineering output.\nCorrelate application performance data - JVM heap behavior, GC pressure, thread contention - with infrastructure events to enable root cause analysis across layers.\nIncident Response & Problem Management\nIncident response at this level is an engineering activity. Senior specialists lead the technical response to high-severity events, own the post-mortem process, and drive the engineering work that prevents recurrence.\nLead high-severity incident response for assigned systems, including dual-datacenter failover execution; own the technical resolution from detection through remediation.\nApply structured root cause analysis: distinguish symptoms from causes, identify contributing factors across system layers, and drive remediation that addresses root cause rather than surface behavior.\nAuthor post-mortems that produce actionable engineering work items - not process improvements alone; track remediation to completion and validate effectiveness.\nDiagnose complex cross-layer failures: Java/JVM application failures, distributed system race conditions, database connection pool exhaustion, messaging system backpressure, and cross-datacenter synchronization issues.\nBuild and maintain incident response runbooks as engineering artifacts: automated where feasible, version-controlled, and validated during chaos engineering exercises.\nParticipate in blameless post-mortem facilitation; model the engineering culture that treats incidents as system failures, not human failures.\nJava Application Reliability\nThe primary application platform is Java/Spring Boot. Senior specialists are expected to operate at the intersection of application engineering and reliability - understanding the runtime deeply enough to diagnose, tune, and improve production behavior.\nDiagnose and resolve Java application performance problems in production: heap memory pressure, garbage collection tuning, thread pool exhaustion, connection leak detection, and class loading anomalies.\nPerform JVM performance analysis using heap dumps, thread dumps, and continuous profiling; translate findings into engineering recommendations for development teams.\nInstrument Spring Boot applications with production-grade observability: Micrometer metrics, structured logging with correlation IDs, and distributed trace integration.\nDiagnose failures across the Java application stack: Spring Boot service behavior, PostgreSQL and Oracle query performance, Kafka and ActiveMQ messaging reliability, and REST/SOAP API integration failures.\nContribute to Java application design reviews with a reliability lens: identify failure modes, single points of failure, and observability gaps before code ships to production.\nPlatform & Infrastructure Engineering\nSenior specialists build and maintain the platform engineering components that reliability depends on - container orchestration, infrastructure automation, deployment tooling, and environment governance.\nDesign and operate Kubernetes and OpenShift workloads for reliability: resource quotas, pod disruption budgets, horizontal pod autoscaling, and liveness and readiness probe engineering.\nBuild infrastructure-as-code for reliability infrastructure: Terraform modules, Ansible/AAP playbooks, and Azure Resource Manager templates that are tested, version-controlled, and peer-reviewed.\nOwn dual-datacenter reliability architecture for assigned systems: synchronization validation, automated failover triggering, traffic management, and recovery time objective verification.\nDesign and automate environment promotion pipelines: ensure that configuration, secrets, and infrastructure state are consistent and validated across development, test, staging, and production.\nBuild and maintain automated patch compliance workflows; integrate CVE remediation into CI/CD pipelines rather than treating it as a manual operational process.\nNERC/CIP Compliance\nNERC/CIP compliance for assigned systems is an engineering responsibility at this level - not a documentation exercise. Senior specialists implement controls through code and automation wherever possible.\nOwn NERC/CIP compliance for assigned systems: interpret applicable reliability standards, implement required controls, maintain evidence documentation, and prepare for regulatory audit.\nEngineer compliance controls into the platform where possible: automated hardening scripts, configuration drift detection, access control validation, and audit log integrity verification.\nMaintain currency on applicable NERC/CIP standards and ERCOT-specific regulatory requirements; escalate emerging compliance risks to the Lead or Manager.\nParticipate in regulatory audit preparation: produce control evidence, respond to auditor inquiries, and coordinate with compliance stakeholders on findings remediation.\nTechnical Leadership & Mentoring\nHold formal mentoring responsibility for Systems Reliability Specialist I and II team members: structured coaching on SRE practices, code review for reliability tooling, and career development conversations.\nServe as the recognized technical authority on reliability engineering and Java application operations for the team; adjacent teams and development engineers seek out this specialist for guidance.\nLead design reviews for systems within the team's scope; identify reliability risks and observability gaps before systems reach production.\nSet engineering standards for the team: post-mortem quality, observability instrumentation, chaos engineering practices, and on-call readiness.\nContribute to the broader engineering organization: internal technical talks, SRE practice documentation, and shared tooling that other teams can adopt.\nEXPERIENCE\nMinimum 5 years of progressive experience in syste...","description_format":"text","description_chars":14758,"description_truncated":true,"requirements":{"experience_years_min":5,"management_years_min":null,"team_size_min":null,"manages_managers":false,"education":{"level":"bachelor","optional":false},"security_clearance":false,"languages":[]},"benefits":[],"hiring_locations":[{"name":"United States","iso":"US","kind":"country"}],"hiring_excludes":[],"relocation_offered":false,"industries":["Energy & Utilities","Government","Power Grid"],"lifecycle":[{"event":"open","at":"2026-10-02T06:13:41Z"}],"visa":[],"liveness":{"score":15,"band":"cold","label":"Long shot","p_open":1,"p_active":0.529,"p_room":0.28,"age_days":64,"expected_fill_days":29,"reasons":["conf:2","velocity","win:tail","crowd:"],"computed_at":"2026-10-09T06:01:00Z"},"pay":{"stated_usd_annual":150000,"is_top_pay":false},"html_url":"https://alion.io/job/ercot-senior-systems-reliability-engineer","json_url":"https://alion.io/job/ercot-senior-systems-reliability-engineer.json","meta":{"generated_at":"2026-10-10T00:37:37Z","cache_seconds":300,"methodology":"https://alion.io/methodology","terms":"https://alion.io/terms","contact":"https://alion.io/contact","api":"https://alion.io/developers","about":"Alion is a live layer of people, companies and AI agents: who they are, whether they are real and active right now, what they do and how to work with them, readable by people and by agents and paid per call.","catalog":"https://alion.io/catalog.json","usage":{"tier":"crawler","counted_by":"address","units_charged":1,"used_today":1014,"day_limit":5000,"remaining_today":3986,"minute_limit":60,"resets_at":"2026-10-11T00:00:00Z"}},"offers":[{"id":"company.slices","title":"One company in depth, by slice","status":"live","price":{"credits":0.02,"usd":0.002,"plus_per_slice":{"credits":0.05,"usd":0.005}},"unit":"per company, plus each slice with data","note":"the employer in depth","call":{"mcp_tool":"get_company","arguments":{"id":6006},"rest":"https://alion.io/mcp/rest/get_company?id=6006"},"human":"https://alion.io/catalog?offer=company.slices&for=job%2Fercot-senior-systems-reliability-engineer"},{"id":"market.stats","title":"A market slice: pay, demand and time to fill","status":"live","price":{"credits":1,"usd":0.1},"unit":"per slice","note":"pay, demand and time to fill for this role and place","call":{"mcp_tool":"market_stats"},"human":"https://alion.io/catalog?offer=market.stats&for=job%2Fercot-senior-systems-reliability-engineer"},{"id":"job.search","title":"Open jobs by role, technology, place, pay and visa","status":"live","price":{"credits":0.02,"usd":0.002},"unit":"per posting in a list","note":"similar open postings","call":{"mcp_tool":"search_jobs"},"human":"https://alion.io/catalog?offer=job.search&for=job%2Fercot-senior-systems-reliability-engineer"},{"id":"company.verify","title":"Is this company real and active right now","status":"pilot","price":null,"unit":"per company","request":{"url":"https://alion.io/catalog/request","method":"POST","body":"{\"offer\": \"company.verify\", \"for\": \"job/ercot-senior-systems-reliability-engineer\", \"note\": \"what you need it for\"}"},"human":"https://alion.io/catalog?offer=company.verify&for=job%2Fercot-senior-systems-reliability-engineer"}]}