{"id":2039489,"url":"https://alion.io/job/glint-tech-solutions-lead-site-reliability-engineer","title":"Lead Site Reliability Engineer","company":{"id":3854730,"name":"Glint Tech Solutions","domain":"glinttechsolutions.com","url":"https://alion.io/company/glint-tech-solutions","size_band":null,"is_staffing_agency":true,"employer_type":"agency","is_intermediary":false,"listed_via":null,"ats_vendor":"Manatal","truth_index":null},"role":"DevOps","role_family":"DevOps","seniority":"lead","employment_type":null,"work_mode":"on_site","remote_scope":null,"remote_scope_basis":null,"remote_working_hours":null,"hiring_geo_confidence":"structured","locations":["Buffalo, United States"],"countries":["US"],"hiring_countries":[],"hiring_countries_total":0,"salary":null,"salary_estimate":{"min_usd":118000,"max_usd":242000,"period":"year","method":"role_seniority_country_remote_cell","sample_n":501},"experience_years_min":null,"visa_sponsorship":false,"relocation_package":false,"has_equity":false,"technologies":[{"name":"Azure","optional":false},{"name":"CI/CD","optional":false},{"name":"Dynatrace","optional":false},{"name":"Incident Management","optional":false},{"name":"OpenTelemetry","optional":false},{"name":"Platform Engineering","optional":false},{"name":"Self-Healing","optional":false},{"name":"Terraform","optional":false},{"name":"Agile","optional":true},{"name":"PowerShell","optional":true},{"name":"Python","optional":true}],"status":"live","first_seen_at":"2026-10-07T17:22:34Z","employer_posted_date":"2026-10-07","last_verified_at":"2026-10-11T17:20:30Z","board_verified":true,"closed_at":null,"days_open":4,"trust":{"level":"ok","repost_count":null,"flags":[],"days_open":4},"description":"Job Title: Lead Site Reliability Engineer\nLocation: Remote within the USA, or onsite in Buffalo, NY / Wilmington, DE (client preference for candidates near these areas). New hires are required to work onsite at the client's office for the first 2-3 weeks (treated as a business trip; travel expenses covered by the company).\nCompany Overview\nGlint Tech Solutions is a women-owned, global IT staffing and recruiting firm serving enterprise clients across the USA and Canada.\nProject Description\nA leading financial services client is seeking a Lead Site Reliability Engineer responsible at the expert level for ensuring the reliability, scalability, performance, and operational excellence of critical banking platforms and applications. This senior individual contributor will design, implement, and improve SRE practices across the software development lifecycle, working closely with application development, infrastructure, platform engineering, and business teams to enhance system resiliency through automation, observability, testing, and proactive operational management, while coaching and influencing others.\nKey Responsibilities\nDesign, implement, and support highly available, scalable, and resilient applications and cloud infrastructure following enterprise SRE best practices\nDefine, implement, and monitor SLOs, SLIs, and error budgets for critical business services\nDevelop observability strategies using Dynatrace, OpenTelemetry (OTel), distributed tracing, metrics, logging, dashboards, and alerting\nAnalyze production telemetry to proactively identify performance bottlenecks, reliability risks, and capacity constraints\nLead incident response for high-severity production events and facilitate Root Cause Analysis (RCA)\nDrive operational excellence through automation of deployments, recovery procedures, and reliability controls\nDesign and execute automated regression testing strategies to validate stability and performance\nCreate and maintain Infrastructure as Code (IaC) solutions using Terraform\nSupport and optimize Microsoft Azure environments, including App Services, scaling, and deployment automation\nUtilize Azure Monitor, Application Insights, and Log Analytics to improve platform visibility\nDrive performance testing, resiliency testing, and disaster recovery preparedness\nLead capacity planning, performance tuning, and workload optimization\nDevelop operational runbooks, incident playbooks, and standard operating procedures\nMentor engineers on observability, cloud engineering, automation, and SRE principles\nAdhere to Company risk and regulatory standards, policies, and controls\nMandatory Skills\nStrong hands-on experience with Dynatrace, OpenTelemetry (OTel), distributed tracing, metrics collection, and centralized logging\nProven experience designing and executing automated regression testing frameworks\nStrong proficiency in Infrastructure as Code (IaC) using Terraform\nExperience with CI/CD pipelines, deployment automation, and operational tooling\nExpert knowledge of production systems monitoring, incident management, and operational troubleshooting\nStrong understanding of application performance management, distributed systems, and cloud-native architectures\nStrong experience with Microsoft Azure (App Services, Resource Groups, networking, scaling, deployment/release management)\nExperience with Azure Monitor, Application Insights, Log Analytics, and Azure dashboards/alerting\nExperience supporting cloud-native and hybrid infrastructure environments\nDemonstrated experience implementing SRE practices - SLOs, SLIs, error budgets, incident/problem management, RCA, reliability automation\nAbility to improve system reliability through performance tuning, capacity planning, and observability-driven insights\nExperience developing automated recovery mechanisms and self-healing solutions\nKnowledge of resiliency engineering patterns, disaster recovery planning, and high-availability architectures\nNice-to-Have Skills\nExperience supporting large-scale enterprise applications in regulated environments\nExperience working in Agile and DevOps operating models\nAbility to work autonomously and lead complex reliability initiatives\nExperience partnering with architecture, infrastructure, cybersecurity, and application development teams\nScripting/automation experience with PowerShell, Python, or Bash\nIndustry certifications in Azure, Terraform, Cloud Engineering, or Site Reliability Engineering\nProven experience leading major incident response and post-incident improvement efforts","description_format":"text","description_chars":4535,"description_truncated":false,"requirements":{"experience_years_min":null,"management_years_min":null,"team_size_min":null,"manages_managers":false,"education":null,"security_clearance":false,"languages":[]},"benefits":[],"hiring_locations":[{"name":"United States","iso":"US","kind":"country"}],"hiring_excludes":[],"relocation_offered":false,"industries":["Professional Services","Recruiting & Staffing"],"lifecycle":[{"event":"open","at":"2026-10-07T17:22:34Z"}],"visa":[],"liveness":{"score":54,"band":"ok","label":"Likely open","p_open":1,"p_active":0.542,"p_room":1,"age_days":2,"expected_fill_days":23,"reasons":["conf:3","agency","velocity","win:early","comp:brand"],"computed_at":"2026-10-10T05:45:15Z"},"pay":null,"html_url":"https://alion.io/job/glint-tech-solutions-lead-site-reliability-engineer","json_url":"https://alion.io/job/glint-tech-solutions-lead-site-reliability-engineer.json","meta":{"generated_at":"2026-10-11T19:39:19Z","cache_seconds":300,"methodology":"https://alion.io/methodology","terms":"https://alion.io/terms","contact":"https://alion.io/contact","api":"https://alion.io/developers","about":"Alion is a live layer of people, companies and AI agents: who they are, whether they are real and active right now, what they do and how to work with them, readable by people and by agents and paid per call.","catalog":"https://alion.io/catalog.json","usage":{"tier":"crawler_verified","counted_by":"address","units_charged":1,"used_today":6757,"day_limit":null,"remaining_today":null,"minute_limit":300,"resets_at":"2026-10-12T00:00:00Z"}},"offers":[{"id":"company.slices","title":"One company in depth, by slice","status":"live","price":{"credits":0.02,"usd":0.002,"plus_per_slice":{"credits":0.05,"usd":0.005}},"unit":"per company, plus each slice with data","note":"the employer in depth","call":{"mcp_tool":"get_company","arguments":{"id":3854730},"rest":"https://alion.io/mcp/rest/get_company?id=3854730"},"human":"https://alion.io/catalog?offer=company.slices&for=job%2Fglint-tech-solutions-lead-site-reliability-engineer"},{"id":"market.stats","title":"A market slice: pay, demand and time to fill","status":"live","price":{"credits":1,"usd":0.1},"unit":"per slice","note":"pay, demand and time to fill for this role and place","call":{"mcp_tool":"market_stats"},"human":"https://alion.io/catalog?offer=market.stats&for=job%2Fglint-tech-solutions-lead-site-reliability-engineer"},{"id":"job.search","title":"Open jobs by role, technology, place, pay and visa","status":"live","price":{"credits":0.02,"usd":0.002},"unit":"per posting in a list","note":"similar open postings","call":{"mcp_tool":"search_jobs"},"human":"https://alion.io/catalog?offer=job.search&for=job%2Fglint-tech-solutions-lead-site-reliability-engineer"},{"id":"company.verify","title":"Is this company real and active right now","status":"pilot","price":null,"unit":"per company","request":{"url":"https://alion.io/catalog/request","method":"POST","body":"{\"offer\": \"company.verify\", \"for\": \"job/glint-tech-solutions-lead-site-reliability-engineer\", \"note\": \"what you need it for\"}"},"human":"https://alion.io/catalog?offer=company.verify&for=job%2Fglint-tech-solutions-lead-site-reliability-engineer"}]}