{"id":2071190,"url":"https://alion.io/job/atlan-senior-software-engineer-reliability","title":"Senior Software Engineer - Reliability","company":{"id":63265,"name":"Atlan","domain":"atlan.com","url":"https://alion.io/company/atlan","size_band":"201-500","is_staffing_agency":false,"employer_type":"direct","is_intermediary":false,"listed_via":null,"ats_vendor":"Ashby","truth_index":null},"role":"Backend","role_family":"Backend","seniority":"senior","employment_type":"full_time","work_mode":"remote","remote_scope":"stated_countries","remote_scope_basis":"board_field","remote_working_hours":null,"hiring_geo_confidence":"structured","locations":[],"countries":[],"hiring_countries":["IN"],"hiring_countries_total":1,"salary":null,"salary_estimate":{"min_usd":21000,"max_usd":47000,"period":"year","method":"role_seniority_country_remote_cell","sample_n":18},"experience_years_min":null,"visa_sponsorship":false,"relocation_package":false,"has_equity":false,"technologies":[{"name":"AI Agents","optional":false},{"name":"LLM Evaluation","optional":false},{"name":"LLM Guardrails","optional":false}],"status":"live","first_seen_at":"2026-10-08T06:08:22Z","employer_posted_date":"2026-10-08","last_verified_at":"2026-10-11T03:10:41Z","board_verified":true,"closed_at":null,"days_open":2,"trust":{"level":"ok","repost_count":null,"flags":[],"days_open":2},"description":"Who We Are\nAtlan is building the context layer for enterprise AI. Enterprises are pouring money into AI and most of it dies in production because the AI does not understand the business context around the data. That is the problem Atlan solves.\nGartner has named context the defining issue in enterprise AI, and calls context graphs the essential infrastructure for AI agents, naming Atlan one of three vendors already building it. We are also the only vendor named a Leader across all four major Gartner and Forrester evaluations for data catalogs, data governance and metadata management, the foundation the context layer runs on.\nCome build the infrastructure that AI runs on.\nThe Team\nYou'll join the Reliability team, which runs a multi-agent AI SRE platform that is already live in production. Today, agents investigate incidents, remediate them, and resolve networking tickets across every tenant we run. A small core team built it and shipped it fast. The scope has now outgrown them.\nBe clear on what this role is not. It is not an incident-command SRE seat. Our view is simple: if a human is fixing production by hand, the system has failed. It is also not a greenfield charter. You'll join a mature, opinionated codebase with documented architecture decisions and eval-gated pull requests, and you'll make it better.\nWhy now: the platform works, and the hardest problems are wide open. Making agent investigations accurate enough that engineers trust them without re-checking is a genuine frontier problem in agentic reliability. Other Atlan teams are lining up to put their own reliability metrics on the platform. What you build in the next year decides how far autonomous operations can go at Atlan.\nWhat You Will Do\nMake investigation agents right, not just fast. Push root-cause accuracy toward the point where engineers trust the answer without re-checking it. Treat every wrong diagnosis as a class of failure to remove, not a one-off bug to patch.\n\nGrow auto-remediation coverage, safely. Take remediation from a handful of playbooks to dozens. Each one earns autonomy in stages: dry run before execute, human-approved before autonomous.\n\nOwn the evals that decide when an agent can be trusted. Build and run fault-injection benchmarks and eval harnesses. Set pass marks before the run, report results with honest denominators, and validate the harness itself, not just the agent.\n\nDesign the gates that let an agent write to production. Fail-closed checks, kill switches, approval flows and blast-radius limits. You start from the question \"what happens when the agent is wrong?\"\n\nBuild new agents where the toil says they belong. Every agent maps to a category of toil it removes, and each one you ship should make the next one cheaper to build.\n\nBuild the platform for other teams. Clean interfaces, guardrails and safe defaults, so teams beyond reliability can trust the platform on their own production. Your hiring manager owns cross-team alignment. You earn adoption through what you build.\n\nShip the unglamorous fix when it matters. When something is on fire, you stop the bleeding first, then come back and remove the class.\n\nWhat Makes You a Match\nYou've lived operational toil. You've carried a pager, owned incidents end to end, or worked a support or escalation queue long enough to know which pain is worth removing and why. The toil was yours, not something you read about.\n\nYou remove classes, not tasks. You can point to a recurring operational problem you made disappear, explain the category behind it, and tell us why you picked that one. Not a faster runbook. Gone.\n\nYou start from the problem and measure by adoption. When you describe something you built, you reach on your own for who used it and what changed, with real numbers and real denominators. You won't call one success a rate. You can also name something you built that didn't get adopted, and what you learned.\n\nAI has changed how you work, structurally. You've rebuilt a core part of your own work end to end with AI, so it's different, not just faster. You've shipped an AI-native workflow that other people depend on. If you've architected agents that take action on their own, you describe the guardrails before the capability.\n\nYou have failure-mode humility. You know what an agent being wrong on production costs, because you've been on the other end of it. You think in false-positive rates, rollback paths and blast radius.\n\nYou build for someone else's production. Clean interfaces, documented failure modes, safe defaults. Experience getting other teams to adopt a platform you built helps, but the instinct matters more than the track record.\n\nYou raise the bar in an existing codebase. You can join a high-rigor system, contribute at its standard quickly, and improve it rather than rewrite it.\n\nYou write code when code is the right tool. We run deterministic code for routing, filtering and safety before any generative step. You think in systems and use agents as leverage, not as a replacement for judgment.\n\nReliability work genuinely energises you. You see operational efficiency as mission-critical for a company whose product is trust in data and AI, not as a stepping stone to something else.\n\nThis probably isn't for you if you want to design the perfect autonomous platform before shipping anything, your instinct is to make toil faster rather than make it disappear, or you're fluent in agents but have never owned the consequences of one being wrong.\nWorking hours: We work best with overlap between roughly 11am and 8pm IST.\nWhy Atlan?\nJoining Atlan means being part of a global movement to help data teams do their life’s best work. Here’s what you can expect:\nCompetitive Compensation: We benchmark at the top of the market and keep compensation simple: strong base salary, performance-based variable pay, and impact-driven equity (for most roles), so your total rewards grow in step with the value you create over time.\n\nAI Native Culture: Atlan is where AI-native builders come to build the systems the future of work will run on. AI isn’t an add-on, it’s woven into how we build, think, and work every day, empowering every Atlanian to move faster and create a bigger impact.\n\nHealth & Wellness: From Day-1 health, dental, vision, and mental health to flexible health stipends, we design benefits offerings that lead in each country we're in.\n\nFlexible Time Off & Leave Policies: We trust you to own your energy: flexible time off and modern leave so you can unplug properly, support yourself and your loved ones, and come back ready to drive an impact.\n\nAccelerated Growth & Learning: Develop at an uncommon velocity through cutting-edge tech, complex implementations, and an experienced team that values mastery.\n\nGlobal, Remote-First, High-Trust: Work from anywhere with a diverse team across 15+ countries, in a trust-first, async environment that gives you true flexibility and ownership over how you work.\n\nMore About Us\nAtlan is building the shared context layer that enterprises need so AI can operate on trusted, governed context. The conversation has moved from data leaders asking: “Can we trust the data in our stack?” to businesses asking: “Can we trust AI inside the business?”\nWe are the missing infrastructure for businesses becoming AI-forward - the connective tissue between their data stack, operational systems, and AI agents.\nTo learn more, visit www.atlan.com and follow us on LinkedIn.\nEqual Opportunity Employer\nAtlan is committed to building an inclusive, diverse, and authentic workplace. We do not discriminate based on race, color, religion, national origin, age, disability, sex, gender identity or expression, sexual orientation, marital status, military or veteran status, or any other legally protected characteristic.\nRecruitment Fraud Alert\nYour safety is important to us. Be cautious of recruiting emails from domains other than @atlan.com. If you're ever unsure about a communication, don't click any links - visit atlan.com/careers directly for confirmed position openings.\nAtlan will never ask candidates to make a payment to apply for or obtain a position. We also will never request financial information, passwords, or other sensitive information from candidates. If you receive a communication that you believe may not be legitimate, do not send money or sensitive information. Please report the same on .","description_format":"text","description_chars":8400,"description_truncated":false,"requirements":{"experience_years_min":null,"management_years_min":null,"team_size_min":null,"manages_managers":false,"education":null,"security_clearance":false,"languages":[]},"benefits":["Equity","Flexible time off"],"hiring_locations":[{"name":"India","iso":"IN","kind":"country"}],"hiring_excludes":[],"relocation_offered":false,"industries":["Data Governance & Catalogs"],"lifecycle":[{"event":"open","at":"2026-10-08T07:06:36Z"}],"visa":[],"liveness":{"score":86,"band":"hot","label":"Hiring now","p_open":1,"p_active":0.86,"p_room":1,"age_days":1,"expected_fill_days":42,"reasons":["conf:0","win:early"],"computed_at":"2026-10-10T05:45:15Z"},"pay":null,"html_url":"https://alion.io/job/atlan-senior-software-engineer-reliability","json_url":"https://alion.io/job/atlan-senior-software-engineer-reliability.json","meta":{"generated_at":"2026-10-11T04:53:41Z","cache_seconds":300,"methodology":"https://alion.io/methodology","terms":"https://alion.io/terms","contact":"https://alion.io/contact","api":"https://alion.io/developers","about":"Alion is a live layer of people, companies and AI agents: who they are, whether they are real and active right now, what they do and how to work with them, readable by people and by agents and paid per call.","catalog":"https://alion.io/catalog.json","usage":{"tier":"crawler","counted_by":"address","units_charged":1,"used_today":4054,"day_limit":5000,"remaining_today":946,"minute_limit":60,"resets_at":"2026-10-12T00:00:00Z"}},"offers":[{"id":"company.slices","title":"One company in depth, by slice","status":"live","price":{"credits":0.02,"usd":0.002,"plus_per_slice":{"credits":0.05,"usd":0.005}},"unit":"per company, plus each slice with data","note":"the employer in depth","call":{"mcp_tool":"get_company","arguments":{"id":63265},"rest":"https://alion.io/mcp/rest/get_company?id=63265"},"human":"https://alion.io/catalog?offer=company.slices&for=job%2Fatlan-senior-software-engineer-reliability"},{"id":"market.stats","title":"A market slice: pay, demand and time to fill","status":"live","price":{"credits":1,"usd":0.1},"unit":"per slice","note":"pay, demand and time to fill for this role and place","call":{"mcp_tool":"market_stats"},"human":"https://alion.io/catalog?offer=market.stats&for=job%2Fatlan-senior-software-engineer-reliability"},{"id":"job.search","title":"Open jobs by role, technology, place, pay and visa","status":"live","price":{"credits":0.02,"usd":0.002},"unit":"per posting in a list","note":"similar open postings","call":{"mcp_tool":"search_jobs"},"human":"https://alion.io/catalog?offer=job.search&for=job%2Fatlan-senior-software-engineer-reliability"},{"id":"company.verify","title":"Is this company real and active right now","status":"pilot","price":null,"unit":"per company","request":{"url":"https://alion.io/catalog/request","method":"POST","body":"{\"offer\": \"company.verify\", \"for\": \"job/atlan-senior-software-engineer-reliability\", \"note\": \"what you need it for\"}"},"human":"https://alion.io/catalog?offer=company.verify&for=job%2Fatlan-senior-software-engineer-reliability"}]}