{"id":1304449,"url":"https://alion.io/job/tonghuaholding-site-reliability-engineer-system-admin","title":"Site Reliability Engineer / System Admin","company":{"id":3813977,"name":"tonghuaholding","domain":"tonghuaholding.com","url":"https://alion.io/company/tonghuaholding","size_band":null,"is_staffing_agency":false,"employer_type":"direct","is_intermediary":false,"listed_via":null,"ats_vendor":null,"truth_index":null},"role":"DevOps","role_family":"DevOps","seniority":"middle","employment_type":"full_time","work_mode":"on_site","remote_scope":null,"remote_scope_basis":null,"remote_working_hours":null,"hiring_geo_confidence":"structured","locations":["Bangkok, Thailand"],"countries":["TH"],"hiring_countries":[],"hiring_countries_total":0,"salary":null,"salary_estimate":null,"experience_years_min":4,"visa_sponsorship":false,"relocation_package":false,"has_equity":false,"technologies":[{"name":"Anomaly Detection","optional":false},{"name":"Ansible","optional":false},{"name":"Bash","optional":false},{"name":"CI/CD","optional":false},{"name":"Configuration Management","optional":false},{"name":"DHCP","optional":false},{"name":"DNS","optional":false},{"name":"Docker","optional":false},{"name":"GitHub Actions","optional":false},{"name":"Go","optional":false},{"name":"Grafana","optional":false},{"name":"HAProxy","optional":false},{"name":"Kubernetes","optional":false},{"name":"Linux","optional":false},{"name":"Loki","optional":false},{"name":"MinIO","optional":false},{"name":"MySQL","optional":false},{"name":"Nginx","optional":false},{"name":"OpenShift","optional":false},{"name":"PostgreSQL","optional":false},{"name":"Prometheus","optional":false},{"name":"Proxmox VE","optional":false},{"name":"Python","optional":false},{"name":"Qdrant","optional":false},{"name":"Redis","optional":false},{"name":"TCP/IP","optional":false},{"name":"Terraform","optional":false},{"name":"VMWare","optional":false},{"name":"VPN","optional":false},{"name":"YugabyteDB","optional":false}],"status":"live","first_seen_at":"2026-09-25T03:38:32Z","employer_posted_date":null,"last_verified_at":"2026-09-25T03:38:32Z","board_verified":false,"closed_at":null,"days_open":6,"trust":{"level":"not_scored","repost_count":null,"flags":[],"days_open":6},"description":"As a Site Reliability Engineer / System Administrator at THCloud.AI, you will be responsible for maintaining the reliability, scalability, and efficiency of our AI and blockchain infrastructure across on-premise and multi-cloud environments. You will drive automation and operational excellence by designing, implementing, and managing CI/CD pipelines, monitoring system health, and proactively addressing potential issues before they impact performance. Key Responsibilities:1. Maintain, monitor, and troubleshoot the company's cloud, blockchain, AI and associated business systems across on-premise and multi-cloud environments.2. Deploy and manage applications on Linux platforms and virtualized infrastructure (Proxmox, VMware, OpenShift), handling system installations, configurations, and ongoing maintenance tasks.3. Develop, implement, and manage CI/CD pipelines using tools such as GitHub Actions, Ansible, and Kubernetes to ensure seamless and efficient deployment workflows.4. Design high-availability systems with load balancing (HAProxy, Nginx), caching (Redis), and failover configurations.5. Conduct daily monitoring, data backup, and recovery using open-source monitoring tools (Prometheus, Grafana, Loki) for performance reporting, issue tracking, and proactive health checks.6. Perform anomaly detection, root cause analysis, and automated alerting to address and prevent system failures and performance bottlenecks.7. Automate operational tasks and improve system resilience through scripting (Bash, Python, or Golang) and configuration management tools.8. Maintain and optimize infrastructure components such as Docker, Kubernetes, databases (PostgreSQL, MySQL), and distributed storage (Ceph, MinIO).9. Setup VPN, VPC, and secure networking for client environments with proper isolation and security.10. Collaborate with cross-functional teams to support infrastructure improvements, incident response, and operational resilience.\nQualifications\nBachelor’s degree in Computer Science, Information Technology, or a related field, with 4+ years of relevant experience in DevOps, SRE, or similar roles\n\nDemonstrated experience with production-grade infrastructure in high-availability (e.g. load balancing) and high-performance environments (e.g. cache optimization)\n\nProficiency in Linux administration and containerization (Docker, Kubernetes)\n\nStrong knowledge of CI/CD processes and automation tools (Ansible, Terraform) and experience scripting (Python, Shell) for operational automation\n\nSolid understanding of networking protocols (TCP/IP, DNS, DHCP) and networking expertise (VPN, VPC, firewalls)\n\n* Hands-on experience with on-premise virtualization (VMware , ProxMox, OpenShift or similar) and cloud platforms\n\nProficient in monitoring and logging solutions (Prometheus, Grafana, Loki) for proactive system management\n\nFamiliarity with database management and distributed storage solutions, particularly PostgreSQL,YugabyteDB, Qdrant and MinIO\n\nMulti-cloud and hybrid environment experience\n\nAbility to communicate in English at a conversational level\nBenefits\n- เวลาในการทำงานที่ยืดหยุ่น- ประกันสังคม- กองทุนสำรองเลี้ยงชีพ- ลาพักร้อน- ลากิจ 6 วัน- ที่จอดรถยนต์และจักรยานยนต์ฟรี- งบประมาณสำหรับการอบรมและพัฒนาทักษะความรู้- เลี้ยงสังสรรค์ประจำปี- Mini Snack Bar- Boardgame Night / Table Tennis Area","description_format":"text","description_chars":3325,"description_truncated":false,"requirements":{"experience_years_min":4,"management_years_min":null,"team_size_min":null,"manages_managers":false,"education":null,"security_clearance":false,"languages":[]},"benefits":[],"hiring_locations":[],"hiring_excludes":[],"relocation_offered":false,"industries":["Blockchain & Crypto"],"lifecycle":[{"event":"open","at":"2026-09-26T12:33:12Z"}],"liveness":{"score":85,"band":"hot","label":"Hiring now","p_open":1,"p_active":0.852,"p_room":1,"age_days":6,"expected_fill_days":18,"reasons":["seen:6","velocity","win:early"],"computed_at":"2026-10-01T05:45:00Z"},"pay":null,"html_url":"https://alion.io/job/tonghuaholding-site-reliability-engineer-system-admin","json_url":"https://alion.io/job/tonghuaholding-site-reliability-engineer-system-admin.json","meta":{"generated_at":"2026-10-02T03:27:04Z","cache_seconds":300,"methodology":"https://alion.io/methodology","terms":"https://alion.io/terms","contact":"https://alion.io/contact","api":"https://alion.io/developers","usage":{"tier":"crawler","counted_by":"address","units_charged":1,"used_today":4790,"day_limit":5000,"remaining_today":210,"minute_limit":60,"resets_at":"2026-10-03T00:00:00Z"}}}