{"id":1951140,"url":"https://alion.io/job/trintech-lead-site-reliability-engineer","title":"Lead Site Reliability Engineer","company":{"id":2212182,"name":"Trintech","domain":"trintech.com","url":"https://alion.io/company/trintech","size_band":"201-500","is_staffing_agency":false,"employer_type":"direct","is_intermediary":false,"listed_via":null,"ats_vendor":"Workday","truth_index":{"grade":"B","score":79,"open_postings":7,"ghost_share":0,"stale_share":0.857,"repost_share":0,"time_to_fill_p50_days":null,"computed_at":"2026-10-10T05:45:15Z"}},"role":"DevOps","role_family":"DevOps","seniority":"lead","employment_type":"full_time","work_mode":"hybrid","remote_scope":null,"remote_scope_basis":null,"remote_working_hours":null,"hiring_geo_confidence":"structured","locations":["Bengaluru, India"],"countries":["IN"],"hiring_countries":[],"hiring_countries_total":0,"salary":null,"salary_estimate":{"min_usd":29000,"max_usd":73000,"period":"year","method":"role_seniority_country_remote_cell","sample_n":12},"experience_years_min":null,"visa_sponsorship":false,"relocation_package":false,"has_equity":false,"technologies":[{"name":"Azure","optional":false},{"name":"Datadog","optional":false},{"name":"Grafana","optional":false},{"name":"Kubernetes","optional":false},{"name":"Linux","optional":false},{"name":"OpenShift","optional":false},{"name":"PagerDuty","optional":false},{"name":"Platform Engineering","optional":false},{"name":"Prometheus","optional":false},{"name":"Sumo Logic","optional":false},{"name":"VMWare","optional":false},{"name":"Windows","optional":false},{"name":"ArgoCD","optional":true},{"name":"Azure DevOps","optional":true},{"name":"SaltStack","optional":true},{"name":"Terraform","optional":true}],"status":"closed","first_seen_at":"2026-10-06T10:05:45Z","employer_posted_date":"2026-10-06","last_verified_at":"2026-10-10T16:29:22Z","board_verified":false,"closed_at":"2026-10-10T16:29:22Z","days_open":4,"trust":{"level":"not_scored","repost_count":null,"flags":[],"days_open":4},"description":"Description\nAbout the role\nAt Trintech, the Lead Site Reliability Engineer leads the technical delivery of reliability improvements across a broad service portfolio. You will turn production evidence into a prioritised engineering plan, coordinate work across teams and ensure incident learning results in measurable improvement. This is a technical individual-contributor role with no direct reports.\nYou will remain hands-on across Kubernetes and OKD/OpenShift, Azure, VMware-based infrastructure, and Linux and Windows workloads. Grafana, Prometheus, Datadog, Sumo Logic and PagerDuty support investigation, service visibility and incident response.\nYour impact\nOwn the integrated reliability engineering plan across services, regions and workstreams, with clear technical priorities and accountable follow-through.\nStrengthen the SRE team's effectiveness through practical direction, better operational tools, sustainable response practices and verified engineering outcomes.\nWhat you will do\nBuild a prioritised reliability roadmap from customer impact, incident recurrence, service risk and operational effort, agreeing delivery capacity with the SRE manager.\nLead adoption of meaningful service indicators, objectives and error-budget decisions with product and service owners, establishing ownership and review practices.\nCoordinate alert improvements across Grafana, Prometheus, Datadog, Sumo Logic and PagerDuty, addressing duplication, missing coverage, routing and diagnostic context.\nLead technical response to complex incidents and strengthen regional handovers, escalation readiness and blameless reviews; verify completion of consequential follow-ups.\nDirect and contribute to automation that removes repeated operational work, with code review, testing, bounded permissions and safe recovery behaviour.\nCoordinate readiness, capacity and recovery work across application, data, infrastructure and Platform Engineering teams, resolving dependencies and technical blockers.\nReview high-risk changes and remain hands-on through diagnosis, prototypes and targeted implementation where direct involvement improves the outcome.\nReport customer impact, reliability trends, repeat failures, alert quality and engineering progress; raise gaps in capacity or ownership with clear recommendations.\nWhat you will bring\nA record of leading reliability improvements across multiple production services and teams, with evidence of sustained outcomes after delivery.\nStrong practical judgement across Kubernetes, Linux, application behaviour, networks, storage and hybrid infrastructure, with the ability to involve specialists effectively.\nHands-on depth in monitoring, PromQL, log analysis and alert design, including the ability to distinguish telemetry faults from service failures.\nExperience leading significant incidents, improving on-call practices and driving corrective actions through to verified resolution.\nCredible software and automation skills, including reviewing and implementing maintainable, tested operational tooling and controlled changes.\nAbility to coordinate dependencies, mentor technical leaders and negotiate delivery priorities and reliability trade-offs without relying on line-management authority.\nNice to have\nExperience with OKD/OpenShift, Azure, VMware, Windows, Datadog, Sumo Logic, PagerDuty, Terraform, SaltStack, Azure DevOps or Argo CD.\nExperience supporting distributed teams and SaaS services with database, integration, batch-processing or customer-facing financial-workflow dependencies.\nHow you will work and lead\nSet technical priorities and guide Senior and Principal contributors, agreeing resource commitments with their managers and retaining clear service ownership.\nPartner with the SRE manager on on-call readiness and protected engineering capacity; provide coaching and technical input into development and recruitment.\nMake delivery and recovery decisions within agreed authority, escalating business-risk acceptance, funding and organisation-wide architecture choices to accountable owners.\nWhat success looks like\nThe reliability roadmap delivers measurable reductions in customer disruption, recurring incidents and avoidable operational effort.\nCritical services have agreed owners, meaningful health measures, actionable alerts and tested recovery procedures.\nRegional handovers, escalation and incident follow-ups work consistently, with response load and capacity constraints visible to management.\nTeams deliver coordinated improvements with verified outcomes while engineers develop stronger independent technical judgement.\nAt our core, Trintechers stand committed to fostering a culture rooted in our core values - Humble, Empowered, Reliable, and Open. Together, these values guide our actions, define our identity, and inspire us to continuously strive for excellence in everything we do.\nAll qualified applicants will receive consideration for employment without regard to race, color, religion, sex, sexual orientation, gender identity, national origin or disability.","description_format":"text","description_chars":5040,"description_truncated":false,"requirements":{"experience_years_min":null,"management_years_min":null,"team_size_min":null,"manages_managers":false,"education":null,"security_clearance":false,"languages":[]},"benefits":[],"hiring_locations":[{"name":"India","iso":"IN","kind":"country"}],"hiring_excludes":[],"relocation_offered":false,"industries":["Financial Services","Accounting & Tax Software"],"lifecycle":[{"event":"open","at":"2026-10-06T10:05:45Z"},{"event":"close","at":"2026-10-10T16:29:22Z"}],"visa":[],"liveness":null,"pay":null,"html_url":"https://alion.io/job/trintech-lead-site-reliability-engineer","json_url":"https://alion.io/job/trintech-lead-site-reliability-engineer.json","meta":{"generated_at":"2026-10-11T19:32:31Z","cache_seconds":300,"methodology":"https://alion.io/methodology","terms":"https://alion.io/terms","contact":"https://alion.io/contact","api":"https://alion.io/developers","about":"Alion is a live layer of people, companies and AI agents: who they are, whether they are real and active right now, what they do and how to work with them, readable by people and by agents and paid per call.","catalog":"https://alion.io/catalog.json","usage":{"tier":"crawler_verified","counted_by":"address","units_charged":1,"used_today":6460,"day_limit":null,"remaining_today":null,"minute_limit":300,"resets_at":"2026-10-12T00:00:00Z"}},"offers":[{"id":"company.slices","title":"One company in depth, by slice","status":"live","price":{"credits":0.02,"usd":0.002,"plus_per_slice":{"credits":0.05,"usd":0.005}},"unit":"per company, plus each slice with data","note":"the employer in depth","call":{"mcp_tool":"get_company","arguments":{"id":2212182},"rest":"https://alion.io/mcp/rest/get_company?id=2212182"},"human":"https://alion.io/catalog?offer=company.slices&for=job%2Ftrintech-lead-site-reliability-engineer"},{"id":"market.stats","title":"A market slice: pay, demand and time to fill","status":"live","price":{"credits":1,"usd":0.1},"unit":"per slice","note":"pay, demand and time to fill for this role and place","call":{"mcp_tool":"market_stats"},"human":"https://alion.io/catalog?offer=market.stats&for=job%2Ftrintech-lead-site-reliability-engineer"},{"id":"job.search","title":"Open jobs by role, technology, place, pay and visa","status":"live","price":{"credits":0.02,"usd":0.002},"unit":"per posting in a list","note":"similar open postings","call":{"mcp_tool":"search_jobs"},"human":"https://alion.io/catalog?offer=job.search&for=job%2Ftrintech-lead-site-reliability-engineer"},{"id":"company.verify","title":"Is this company real and active right now","status":"pilot","price":null,"unit":"per company","request":{"url":"https://alion.io/catalog/request","method":"POST","body":"{\"offer\": \"company.verify\", \"for\": \"job/trintech-lead-site-reliability-engineer\", \"note\": \"what you need it for\"}"},"human":"https://alion.io/catalog?offer=company.verify&for=job%2Ftrintech-lead-site-reliability-engineer"}]}