{"id":2071656,"url":"https://alion.io/job/qualys-senior-database-reliability-engineer","title":"Senior Database Reliability Engineer","company":{"id":18966,"name":"Qualys","domain":"qualys.com","url":"https://alion.io/company/qualys","size_band":"1001-5000","is_staffing_agency":false,"employer_type":"direct","is_intermediary":false,"listed_via":null,"ats_vendor":"Workday","truth_index":null},"role":"Data Science","role_family":"Data Science","seniority":"senior","employment_type":"full_time","work_mode":"on_site","remote_scope":null,"remote_scope_basis":null,"remote_working_hours":null,"hiring_geo_confidence":"structured","locations":["Pune, India"],"countries":["IN"],"hiring_countries":[],"hiring_countries_total":0,"salary":null,"salary_estimate":{"min_usd":21000,"max_usd":42000,"period":"year","method":"role_seniority_country_remote_cell","sample_n":54},"experience_years_min":5,"visa_sponsorship":false,"relocation_package":false,"has_equity":false,"technologies":[{"name":"Ansible","optional":false},{"name":"Apache Kafka","optional":false},{"name":"ElasticSearch","optional":false},{"name":"GitOps","optional":false},{"name":"Grafana","optional":false},{"name":"Kubernetes","optional":false},{"name":"Linux","optional":false},{"name":"OpenSearch","optional":false},{"name":"OpenTelemetry","optional":false},{"name":"Prometheus","optional":false},{"name":"Python","optional":false},{"name":"Redis","optional":false},{"name":"Terraform","optional":false},{"name":"AWS","optional":true},{"name":"Azure","optional":true},{"name":"ClickHouse","optional":true},{"name":"Flink","optional":true},{"name":"GCP","optional":true},{"name":"Spark","optional":true}],"status":"live","first_seen_at":"2026-10-08T07:19:22Z","employer_posted_date":"2026-10-08","last_verified_at":"2026-10-09T01:13:10Z","board_verified":true,"closed_at":null,"days_open":0,"trust":{"level":"ok","repost_count":null,"flags":[],"days_open":0},"description":"Come work at a place where innovation and teamwork come together to support the most exciting missions in the world!\nSenior Database Reliability Engineer (DBRE) - Data Platform\nOpenSearch Must-Have | Kafka | Redis | Python Automation | Production Reliability\nLevel\nSenior DBRE (Senior Individual Contributor)\nDomain\nDistributed Data Infrastructure\nPrimary Focus\nHands-on reliability and operational ownership for assigned OpenSearch/Elasticsearch, Kafka, and Redis services; Python automation and disciplined production execution\nWork Expectations\nU.S. Eastern Time business hours; on-call participation as required\n1. Role Summary\nWe are seeking an experienced Senior DBRE to improve the reliability, performance, scalability, and automation of mission-critical distributed data platforms. OpenSearch/Elasticsearch is the core must-have technology, complemented by production experience with Kafka, Redis, Linux systems, observability, and Python automation.\nThis is a hands-on role. The Senior DBRE independently operates assigned clusters and services, troubleshoots production issues, performs upgrades and lifecycle work, builds practical automation, and contributes to platform design and production-readiness standards.\n2. Level Scope and Leadership Expectations\nThe Senior DBRE is a high-impact hands-on engineer focused on independent execution within an assigned platform scope. Senior engineers solve common and moderately complex production problems, improve automation and runbooks, and contribute to designs and standards while escalating broader architectural decisions appropriately.\nExpectation at This Level\nScope of Ownership\nOwns assigned clusters, services, maintenance activities, incidents, upgrades, backup/restore, capacity reviews, and operational improvements with limited supervision.\nTechnical Depth\nDiagnoses common and moderately complex issues involving shards, indexing, search, Kafka lag, Redis memory behavior, JVM, Linux, storage, and networking.\nAutomation\nBuilds and maintains tested Python tools, operational CLIs, API integrations, dashboards, alerts, and runbooks that reduce recurring toil.\nExecution\nExecutes migrations, rolling upgrades, scaling, and DR exercises with change-management discipline and appropriate technical review.\nCollaboration\nPartners with Lead and Staff engineers, SRE, application, infrastructure, storage, network, and cloud teams; documents findings and communicates risks early.\nArchitecture Influence\nContributes evidence and operational experience to design reviews; does not independently set enterprise-wide platform architecture.\n3. Key Responsibilities\nA. OpenSearch and Elasticsearch Operations\nOperate, scale, and improve assigned OpenSearch/Elasticsearch clusters and node topologies.\nTroubleshoot shard allocation, indexing throughput, slow search and aggregation workloads, cluster state, JVM heap/GC, disk, memory, and recovery issues.\nImplement and maintain index templates, shard strategies, ISM/ILM policies, rollover, retention, snapshot, restore, and rolling-upgrade procedures.\nB. Kafka and Redis Operations\nOperate Kafka topics, partitions, replication, retention, consumer groups, and broker health; diagnose lag, rebalance storms, and capacity bottlenecks.\nOperate Redis Cluster and Sentinel environments; address memory fragmentation, eviction behavior, persistence, replication lag, connection limits, and failover issues.\nC. Automation and Infrastructure as Code\nDevelop clean, modular Python automation for health checks, maintenance, validation, reporting, and safe remediation.\nUse Terraform, Ansible, GitOps workflows, or Kubernetes operators where applicable to make platform operations repeatable.\nD. Reliability and Incident Response\nInstrument actionable metrics and alerts using Prometheus, Grafana, OpenTelemetry, or comparable platforms.\nParticipate in on-call, stabilize incidents, document root cause, and complete preventive actions.\nContribute to capacity planning, production-readiness reviews, runbooks, and design reviews.\n4. Required Qualifications and Experience\n5+ years of relevant DBRE, SRE, database, or distributed-systems operations experience in production environments.\nStrong hands-on OpenSearch and/or Elasticsearch experience, including shards, indexing, search, lifecycle management, JVM behavior, and snapshot/restore.\nProduction experience with at least one of Kafka or Redis; experience with both is strongly preferred.\nStrong Linux and systems troubleshooting skills across CPU, memory, storage I/O, networking, and process behavior.\nAbility to write maintainable Python automation beyond simple one-off shell scripts.\nExperience with monitoring, incident response, controlled production changes, upgrades, and operational documentation.\nClear communication, sound judgment, and willingness to escalate risk early.\n5. Preferred Qualifications\nHands-on experience with both Kafka and Redis.\nTerraform, Ansible, GitOps, Kubernetes, ECK/OpenSearch, Strimzi, or Redis operators.\nCloud-managed data services in AWS, Azure, or GCP.\nFamiliarity with ClickHouse, Pinot, vector databases, Flink, Spark Streaming, or Kafka Connect.\n6. Expected Outcomes in the First 6-12 Months\nEstablish measurable baselines and deliver material improvements in reliability, alert quality, MTTR, and recurring operational toil for assigned services.\nAutomate high-frequency maintenance and validation workflows using maintainable Python tooling.\nImprove runbook coverage, upgrade readiness, backup/restore confidence, and capacity visibility.\nContribute a well-supported evaluation or operational-readiness recommendation for one relevant emerging technology when business needs require it.\n7. What This Role Does Not Include\nThis is not a people-management role.\nThis is not passive monitoring, ticket routing, or a generic operations queue; hands-on troubleshooting and improvement are expected.\nThis role does not independently own enterprise-wide architecture or technical strategy.\nThis role does not own application code or product feature delivery, although close partnership with application teams is required.\nThe role does not require equal mastery of every listed platform; OpenSearch/Elasticsearch is the anchor expertise, with complementary Kafka and/or Redis depth.\n8. Work Expectations and Location\nThis role requires working during U.S. Eastern Time (ET) business hours and participating in on-call support as required.","description_format":"text","description_chars":6438,"description_truncated":false,"requirements":{"experience_years_min":5,"management_years_min":null,"team_size_min":null,"manages_managers":false,"education":null,"security_clearance":false,"languages":[]},"benefits":[],"hiring_locations":[],"hiring_excludes":[],"relocation_offered":false,"industries":["Cybersecurity","Endpoint Security","Information Security"],"lifecycle":[{"event":"open","at":"2026-10-08T07:19:22Z"}],"visa":[],"liveness":{"score":86,"band":"hot","label":"Hiring now","p_open":1,"p_active":0.86,"p_room":1,"age_days":0,"expected_fill_days":54,"reasons":["conf:2","win:early","comp:brand"],"computed_at":"2026-10-09T04:04:29Z"},"pay":null,"html_url":"https://alion.io/job/qualys-senior-database-reliability-engineer","json_url":"https://alion.io/job/qualys-senior-database-reliability-engineer.json","meta":{"generated_at":"2026-10-09T04:04:29Z","cache_seconds":300,"methodology":"https://alion.io/methodology","terms":"https://alion.io/terms","contact":"https://alion.io/contact","api":"https://alion.io/developers","about":"Alion is a live layer of people, companies and AI agents: who they are, whether they are real and active right now, what they do and how to work with them, readable by people and by agents and paid per call.","catalog":"https://alion.io/catalog.json","usage":{"tier":"crawler","counted_by":"address","units_charged":1,"used_today":3510,"day_limit":5000,"remaining_today":1490,"minute_limit":60,"resets_at":"2026-10-10T00:00:00Z"}},"offers":[{"id":"company.slices","title":"One company in depth, by slice","status":"live","price":{"credits":0.02,"usd":0.002,"plus_per_slice":{"credits":0.05,"usd":0.005}},"unit":"per company, plus each slice with data","note":"the employer in depth","call":{"mcp_tool":"get_company","arguments":{"id":18966},"rest":"https://alion.io/mcp/rest/get_company?id=18966"},"human":"https://alion.io/catalog?offer=company.slices&for=job%2Fqualys-senior-database-reliability-engineer"},{"id":"market.stats","title":"A market slice: pay, demand and time to fill","status":"live","price":{"credits":1,"usd":0.1},"unit":"per slice","note":"pay, demand and time to fill for this role and place","call":{"mcp_tool":"market_stats"},"human":"https://alion.io/catalog?offer=market.stats&for=job%2Fqualys-senior-database-reliability-engineer"},{"id":"job.search","title":"Open jobs by role, technology, place, pay and visa","status":"live","price":{"credits":0.02,"usd":0.002},"unit":"per posting in a list","note":"similar open postings","call":{"mcp_tool":"search_jobs"},"human":"https://alion.io/catalog?offer=job.search&for=job%2Fqualys-senior-database-reliability-engineer"},{"id":"company.verify","title":"Is this company real and active right now","status":"pilot","price":null,"unit":"per company","request":{"url":"https://alion.io/catalog/request","method":"POST","body":"{\"offer\": \"company.verify\", \"for\": \"job/qualys-senior-database-reliability-engineer\", \"note\": \"what you need it for\"}"},"human":"https://alion.io/catalog?offer=company.verify&for=job%2Fqualys-senior-database-reliability-engineer"}]}