{"id":1272338,"url":"https://alion.io/job/instalily-mid-level-site-reliability-engineer","title":"Mid-Level Site Reliability Engineer","company":{"id":2283862,"name":"InstaLILY","domain":"instalily.ai","url":"https://alion.io/company/instalily","size_band":"51-200","is_staffing_agency":false,"employer_type":"direct","is_intermediary":false,"listed_via":null,"ats_vendor":"Greenhouse","truth_index":{"grade":"A","score":86,"open_postings":18,"ghost_share":0,"stale_share":0.556,"repost_share":0,"time_to_fill_p50_days":null,"computed_at":"2026-10-01T05:45:00Z"}},"role":"DevOps","role_family":"DevOps","seniority":"middle","employment_type":null,"work_mode":"on_site","remote_scope":null,"remote_scope_basis":null,"remote_working_hours":null,"hiring_geo_confidence":"structured","locations":["New York, United States"],"countries":["US"],"hiring_countries":[],"hiring_countries_total":0,"salary":null,"salary_estimate":{"min_usd":98000,"max_usd":200000,"period":"year","method":"role_seniority_country_remote_cell","sample_n":397},"experience_years_min":3,"visa_sponsorship":false,"relocation_package":false,"has_equity":true,"technologies":[{"name":"AI Agents","optional":false},{"name":"ArgoCD","optional":false},{"name":"AWS","optional":false},{"name":"Azure","optional":false},{"name":"CI/CD","optional":false},{"name":"Datadog","optional":false},{"name":"DNS","optional":false},{"name":"Edge AI","optional":false},{"name":"GCP","optional":false},{"name":"GitHub Actions","optional":false},{"name":"GitLab CI","optional":false},{"name":"GitOps","optional":false},{"name":"HIPAA","optional":false},{"name":"IAM","optional":false},{"name":"ISO 27001","optional":false},{"name":"Jenkins","optional":false},{"name":"Kubernetes","optional":false},{"name":"Machine Learning","optional":false},{"name":"OpenTelemetry","optional":false},{"name":"OpenTofu","optional":false},{"name":"Prometheus","optional":false},{"name":"SOC 2","optional":false},{"name":"Terraform","optional":false}],"status":"live","first_seen_at":"2026-06-13T21:19:17Z","employer_posted_date":"2026-07-24","last_verified_at":"2026-10-01T07:49:26Z","board_verified":true,"closed_at":null,"days_open":109,"trust":{"level":"stale","repost_count":0,"flags":["stale"],"days_open":109},"description":"About InstaLILY\nInstaLILY is an AI products and infrastructure company that puts execution at the frontier of enterprise AI. That work begins with Lily™, the world's first AI Forward Deployed Engineer, which learns how a business works, builds the software it needs, and goes live in days. It does not leave when the work ships; it stays and keeps the software working as the business changes. Lily runs wherever the work happens, in the cloud, on-premise, or at the edge, through InstaLILY's Small Data Center, built with NVIDIA technology. Founded in 2023 by Amit Shah and Sumantro Das, InstaLILY has raised nearly $100 million from Energize Capital, Insight Partners, and Home Depot Ventures. Headquartered in New York, with offices in San Francisco, London, and Toronto, InstaLILY serves leading companies across construction, industrial distribution, logistics, healthcare, and other operationally intensive industries. Learn more at https://instalily.ai/.\nThe Traction\nRevenue grew 5x over the past year, and Lily has driven over $200M in new annual sales for a single customer. We serve some of the largest operators in our industries, including SRS Distribution (part of The Home Depot family), United Rentals, and Henry Schein, and we work closely with the Google DeepMind and NVIDIA ecosystems.\nHow We Work\nWe work in small teams with real ownership: clear problems, direct access to the customers whose work you're changing, and room to ship. Your code runs in live production systems inside billion-dollar operations, so you see your impact directly. People who do well here want that proximity to the work. We're growing fast, and the people who join now shape what this company becomes. Everything runs on three principles: Customers, Culture, and Code.\nRole Overview\nInstalily, a cutting-edge AI startup, is seeking a curious and highly skilled Site Reliability Engineer to help build the Internal Developer Platform (IDP) that powers our AI agent platform. We are redefining how organizations leverage AI using vertical agents, and we’re looking for engineers who think of developer experience as a product. You will build the paved roads, golden paths, and self-service tooling that allow every team at Instalily to ship AI products with speed and confidence.\nAs a Mid-Level Site Reliability Engineer, you will help build and own meaningful pieces of our IDP and the multi-cloud, Kubernetes-based infrastructure beneath it. You’ll partner closely with AI and Software Engineers to turn rough edges into self-service abstractions. You will benefit from world-class mentorship from highly-regarded executive leaders at Internet Retailer 100 brands.\nResponsibilities\nBuild out Instalily’s Internal Developer Platform - including the developer portal, golden paths, and one-click developer workflows.\nHelp design and stand up our Kubernetes platform; operate clusters and workloads as services migrate, focusing on networking, autoscaling, RBAC, and reliability.\nDesign, implement, and maintain cloud infrastructure across AWS, GCP, and/or Azure to support the AI agent platform.\nDevelop and maintain Infrastructure as Code (IaC) using OpenTofu, focusing on modularity and safe rollouts.\nBuild and improve CI/CD pipelines and GitOps workflows (e.g., ArgoCD, Flux) for seamless deployment.\nImplement platform security best practices, including IAM, network segmentation, and policy-as-code.\nImplement and tune logging, monitoring, and alerting using tools such as Datadog, Prometheus, and OpenTelemetry.\nOptimize platform environments for cost, performance, and reliability; participate in on-call rotation.\nTreat developers as customers - gather feedback and measure platform adoption to iterate on the developer experience.\nMentor more junior engineers and contribute to architectural discussions and technical reviews.\nRequirements\nBachelor’s or Master’s degree in Computer Science, Engineering, or a related field, or equivalent practical experience.\n3 to 5 years of experience as a platform, cloud, infrastructure, or DevOps engineer.\nProduction experience operating Kubernetes (managing clusters, upgrades, and reliability), with greenfield build-out experience being a strong plus.\nStrong hands-on experience with at least one major cloud platform (AWS, GCP, or Azure); multi-cloud experience is preferred.\nExperience contributing to or building Internal Developer Platforms (golden paths, paved roads) is a strong plus.\nProficiency with Infrastructure as Code (OpenTofu or Terraform) and GitOps workflows (e.g., ArgoCD, Flux).\nExperience with CI/CD tools and practices such as GitHub Actions, Jenkins, or GitLab CI.\nSolid understanding of cloud networking concepts (VPCs, load balancers, DNS, CDNs, service meshes).\nWorking knowledge of cloud security principles, IAM, and compliance frameworks (SOC 2, HIPAA, ISO 27001).\nA product-minded approach to internal tooling, focusing on adoption and feedback loops rather than just uptime.\nStrong problem-solving skills and ability to work in fast-paced, collaborative environments.\nExcellent communication skills to engage effectively with technical and non-technical teams.\nInterest in AI and machine learning infrastructure (GPU workloads, model serving, vector databases) is a plus.\nWhat You'll Get\nProven product: Customers are live; this isn't a bet on an unproven thesis\nAI-native: In how we build, how we work, and what we sell\nStage: Early enough to shape how the company scales\nGlobal: Based in New York with offices in SF and London\nCulture: Sharp, low-ego team that keeps raising the bar\nGrowth: The learning curve is steep\nCompensation and Benefits\nSalary Range: $150,000-$190,000 per year, commensurate with experience\nEquity: Stock options awards, and refreshers for top performers\nBenefits: Medical, Dental, Vision, 401K, in-office Lunch reimbursement, Wellbeing Stipend, Generous Parental leave, PTO and 10 US Federal Holidays, and more!\nQuality Over Quantity\nTo ensure a focused, high-quality hiring experience, we kindly ask candidates to limit their applications to 3 open requisitions at any given time. Applying strategically to roles that best align with your skills and career goals gives you the highest chance of standing out. Have you interviewed with us in the past 12 months? We encourage you to reach out directly to your previous interviewer rather than submitting a new application.\nInstaLILY is committed to providing an inclusive and barrier-free recruitment process. If you require an accommodation, please let us know, and we will work with you to meet your needs.","description_format":"text","description_chars":6564,"description_truncated":false,"requirements":{"experience_years_min":3,"management_years_min":null,"team_size_min":null,"manages_managers":false,"education":{"level":"bachelor","optional":false},"security_clearance":false,"languages":[]},"benefits":["401k plan","Equity","Parental leave","Stock options"],"hiring_locations":[],"hiring_excludes":[],"relocation_offered":false,"industries":["Artificial Intelligence","AI Agents"],"lifecycle":[{"event":"open","at":"2026-09-26T00:38:17Z"}],"liveness":{"score":7,"band":"cold","label":"Long shot","p_open":1,"p_active":0.264,"p_room":0.28,"age_days":109,"expected_fill_days":35,"reasons":["conf:10","stale_co","win:tail","crowd:"],"computed_at":"2026-10-01T05:45:00Z"},"pay":null,"html_url":"https://alion.io/job/instalily-mid-level-site-reliability-engineer","json_url":"https://alion.io/job/instalily-mid-level-site-reliability-engineer.json","meta":{"generated_at":"2026-10-01T17:02:19Z","cache_seconds":300,"methodology":"https://alion.io/methodology","terms":"https://alion.io/terms","contact":"https://alion.io/contact","api":"https://alion.io/developers","usage":{"tier":"crawler","counted_by":"address","units_charged":1,"used_today":7,"day_limit":5000,"remaining_today":4993,"minute_limit":60,"resets_at":"2026-10-02T00:00:00Z"}}}