{"id":2145083,"url":"https://alion.io/job/drivetrain-site-reliability-engineer-sre","title":"Site Reliability Engineer - SRE","company":{"id":2915,"name":"Drivetrain","domain":"drivetrain.io","url":"https://alion.io/company/drivetrain","size_band":null,"is_staffing_agency":false,"employer_type":"direct","is_intermediary":false,"listed_via":null,"ats_vendor":"Lever","truth_index":{"grade":"D","score":54,"open_postings":31,"ghost_share":0.774,"stale_share":0,"repost_share":0,"time_to_fill_p50_days":null,"computed_at":"2026-10-10T05:45:15Z"}},"role":"DevOps","role_family":"DevOps","seniority":"senior","employment_type":"full_time","work_mode":"remote","remote_scope":"stated_countries","remote_scope_basis":"board_field","remote_working_hours":null,"hiring_geo_confidence":"structured","locations":["India"],"countries":["IN"],"hiring_countries":["IN"],"hiring_countries_total":1,"salary":null,"salary_estimate":{"min_usd":18000,"max_usd":42000,"period":"year","method":"role_seniority_country_cell","sample_n":46},"experience_years_min":5,"visa_sponsorship":false,"relocation_package":false,"has_equity":false,"technologies":[{"name":"Amazon CloudWatch","optional":false},{"name":"Amazon EC2","optional":false},{"name":"Amazon EKS","optional":false},{"name":"Amazon S3","optional":false},{"name":"AWS","optional":false},{"name":"CI/CD","optional":false},{"name":"Configuration Management","optional":false},{"name":"DNS","optional":false},{"name":"Docker","optional":false},{"name":"GCP","optional":false},{"name":"Git","optional":false},{"name":"Google GKE","optional":false},{"name":"Grafana","optional":false},{"name":"IAM","optional":false},{"name":"Incident Management","optional":false},{"name":"Istio","optional":false},{"name":"Jenkins","optional":false},{"name":"Kubernetes","optional":false},{"name":"Kustomize","optional":false},{"name":"Linkerd","optional":false},{"name":"Prometheus","optional":false},{"name":"Python","optional":false},{"name":"Service Mesh","optional":false},{"name":"SQL","optional":false},{"name":"Terraform","optional":false},{"name":"Zero Trust","optional":false}],"status":"live","first_seen_at":"2025-09-01T14:13:54Z","employer_posted_date":"2025-09-01","last_verified_at":"2026-10-11T20:33:19Z","board_verified":true,"closed_at":null,"days_open":405,"trust":{"level":"ghost","repost_count":0,"flags":["stale","company_stale"],"days_open":404},"description":"As a Senior Site Reliability Engineer at Drivetrain, you will be a cornerstone of our engineering organization, ensuring our fast-growing SaaS platform remains highly available, performant, and secure. At this stage of our growth, scaling infrastructure efficiently while maintaining the rigorous security and reliability standards required for financial data is paramount. You will take ownership of our multi-cloud infrastructure, drive automation, champion observability, and collaborate closely with development teams to build a culture of reliability from code commit to production.\nKey Responsibilities\nCloud Infrastructure & Orchestration\nMulti-Cloud Management: Architect, manage, and continuously optimize highly available cloud infrastructure across both AWS and GCP. Balance workload demands to ensure maximum cost-efficiency, scalability, and strict security compliance across both platforms.\n\nAdvanced Kubernetes Orchestration: Lead the design, deployment, and management of scalable Kubernetes clusters. Utilize configuration management tools like Kustomize to enforce standardized, repeatable, and automated deployment configurations across all environments.\n\nService Mesh & Security Integration: Implement and maintain service mesh technologies (e.g., Istio, Linkerd) to secure, control, and observe service-to-service communication. Drive container security best practices, including image scanning, runtime protection, and strict RBAC enforcement.\n\nCI/CD & Automation\nPipeline Engineering: Architect, maintain, and optimize robust CI/CD pipelines using Git and Jenkins. Focus on reducing deployment friction, accelerating release velocity, and enforcing automated testing and security gates.\n\nInfrastructure as Code (IaC): Treat infrastructure as software. Write, review, and maintain Terraform modules to provision and manage cloud resources predictably and safely.\n\nOperational Automation: Aggressively reduce operational toil. Develop robust Python scripts and tooling to automate routine maintenance, data backups, scaling operations, and system recovery processes.\n\nObservability & Reliability\nComprehensive Monitoring: Design and enhance our observability stack to provide deep, real-time insights into system health. Manage and scale tools including Prometheus, Grafana, ELK/EFK stack, AWS CloudWatch, and GCP Operations Suite.\n\nReliability Engineering: Spearhead reliability initiatives critical to a scaling SaaS platform. Drive rigorous capacity planning exercises to stay ahead of growth.\n\nIncident Management & SLOs: Own the incident response lifecycle. Facilitate blameless postmortems to extract actionable learnings. Define, track, and enforce SLIs, SLOs, and SLAs, ensuring the platform consistently meets its reliability guarantees.\n\nCollaboration & Leadership\nDevOps Culture: Act as an embedded reliability advocate. Collaborate closely with software engineers early in the development lifecycle to ensure applications are designed for deployability, scalability, and resilience.\n\nContinuous Improvement: Proactively identify system bottlenecks and architectural weaknesses. Contribute to process improvements, build internal developer tooling, and maintain comprehensive documentation to elevate team productivity and system understanding.\n\nRequired Proficiency & Qualifications\nExperience: 5+ years of hands-on experience in Site Reliability Engineering, DevOps, or Cloud Infrastructure roles, preferably within a fast-paced SaaS environment.\n\nCloud Platforms: Deep, proven proficiency in AWS (EC2, EKS, RDS, VPC, IAM, S3) AND GCP (GKE, Compute Engine, Cloud SQL, IAM, Cloud Storage). Ability to navigate and optimize multi-cloud architectures.\n\nContainerization: Expert-level knowledge of Docker and Kubernetes, including advanced deployment strategies and lifecycle management.\n\nAutomation/IaC: Strong programming skills in Python and extensive experience with Terraform.\n\nObservability: Hands-on expertise building dashboards and alerting systems using Prometheus, Grafana, and log aggregation stacks (ELK/EFK).\n\nNetworking & Security: Solid understanding of cloud networking (VPC peering, load balancing, DNS) and zero-trust security principles in a containerized environment.","description_format":"text","description_chars":4218,"description_truncated":false,"requirements":{"experience_years_min":5,"management_years_min":null,"team_size_min":null,"manages_managers":false,"education":null,"security_clearance":false,"languages":[]},"benefits":[],"hiring_locations":[{"name":"India","iso":"IN","kind":"country"}],"hiring_excludes":[],"relocation_offered":false,"industries":["Mobile App Development Services","Web Development Services","Custom Software Development"],"lifecycle":[{"event":"open","at":"2026-10-09T04:13:13Z"}],"visa":[],"liveness":{"score":3,"band":"cold","label":"Long shot","p_open":1,"p_active":0.109,"p_room":0.28,"age_days":403,"expected_fill_days":29,"reasons":["conf:2","stale_co","ghost","win:tail","crowd:"],"computed_at":"2026-10-10T05:45:15Z"},"pay":null,"html_url":"https://alion.io/job/drivetrain-site-reliability-engineer-sre","json_url":"https://alion.io/job/drivetrain-site-reliability-engineer-sre.json","meta":{"generated_at":"2026-10-11T20:39:20Z","cache_seconds":300,"methodology":"https://alion.io/methodology","terms":"https://alion.io/terms","contact":"https://alion.io/contact","api":"https://alion.io/developers","about":"Alion is a live layer of people, companies and AI agents: who they are, whether they are real and active right now, what they do and how to work with them, readable by people and by agents and paid per call.","catalog":"https://alion.io/catalog.json","usage":{"tier":"crawler","counted_by":"address","units_charged":1,"used_today":21,"day_limit":5000,"remaining_today":4979,"minute_limit":60,"resets_at":"2026-10-12T00:00:00Z"}},"offers":[{"id":"company.slices","title":"One company in depth, by slice","status":"live","price":{"credits":0.02,"usd":0.002,"plus_per_slice":{"credits":0.05,"usd":0.005}},"unit":"per company, plus each slice with data","note":"the employer in depth","call":{"mcp_tool":"get_company","arguments":{"id":2915},"rest":"https://alion.io/mcp/rest/get_company?id=2915"},"human":"https://alion.io/catalog?offer=company.slices&for=job%2Fdrivetrain-site-reliability-engineer-sre"},{"id":"market.stats","title":"A market slice: pay, demand and time to fill","status":"live","price":{"credits":1,"usd":0.1},"unit":"per slice","note":"pay, demand and time to fill for this role and place","call":{"mcp_tool":"market_stats"},"human":"https://alion.io/catalog?offer=market.stats&for=job%2Fdrivetrain-site-reliability-engineer-sre"},{"id":"job.search","title":"Open jobs by role, technology, place, pay and visa","status":"live","price":{"credits":0.02,"usd":0.002},"unit":"per posting in a list","note":"similar open postings","call":{"mcp_tool":"search_jobs"},"human":"https://alion.io/catalog?offer=job.search&for=job%2Fdrivetrain-site-reliability-engineer-sre"},{"id":"company.verify","title":"Is this company real and active right now","status":"pilot","price":null,"unit":"per company","request":{"url":"https://alion.io/catalog/request","method":"POST","body":"{\"offer\": \"company.verify\", \"for\": \"job/drivetrain-site-reliability-engineer-sre\", \"note\": \"what you need it for\"}"},"human":"https://alion.io/catalog?offer=company.verify&for=job%2Fdrivetrain-site-reliability-engineer-sre"}]}