747,188open jobs
44,848companies
107,947added this week
Browse all
Location
Hybrid (Seoul, South Korea)
Seniority
Junior · 2+ years exp
Employment
Full-Time

Confirmed on the employer's own hiring board on Sep 25, 2026. First seen by Alion on Aug 26, 2026.

Overview
Company
Impact
Profile match
AI Cloud for All Models. Train, fine-tune, and serve across clouds: self-serve A100 and H100, B200 to GB300 as reserved capacity, multi-node clusters, persistent workspaces, and transparent pricing.

About the Role

VESSL AI의 Senior Site Reliability Engineer 는 VESSL GPU 클라우드 플랫폼의 가용성과 성능을 책임집니다. 메트릭·로그·트레이스 기반의 Observability 체계를 구축해 이상 징후를 조기에 탐지하고, 장애 발생 시 On-call 대응과 Root Cause Analysis를 통해 신속하게 복구하며, 반복적인 운영 작업은 자동화로 대체해 시스템 전반의 안정성과 운영 효율을 함께 끌어올립니다.

GPU 클라우드 플랫폼은 대규모 GPU 클러스터, InfiniBand·RoCE 기반 고성능 네트워크, Kubernetes·Slurm 기반 스케줄링 시스템 등 여러 레이어로 구성됩니다. GPU 워크로드는 몇 시간에서 몇 주까지 이어지는 경우가 많고, 노드 하나의 장애나 네트워크 지연만으로도 진행 중이던 학습·추론 작업 전체가 중단될 수 있어 일반적인 웹 서비스보다 훨씬 높은 수준의 안정성이 요구됩니다. 이를 위해서는 GPU, NIC, 드라이버, 커널에서부터 스케줄러, 네트워크 패브릭에 이르기까지 시스템 전 레이어의 상태를 지속적으로 모니터링하고, 장애가 발생했을 때 원인을 빠르게 좁혀 복구하며, 반복되는 운영 작업을 자동화로 줄여나가는 기능이 필요합니다.

What you will do

  • Reliability & Observability: 메트릭·로그·트레이스 등 Observability 스택을 구축하고, SLI/SLO를 정의해 플랫폼의 안정성을 정량적으로 추적

  • Incident Response: On-call 로테이션에 참여해 장애 상황에 대응하고, Root Cause Analysis와 Postmortem을 통해 재발을 방지

  • Automation: 반복되는 운영 작업(Toil)을 자동화 스크립트·파이프라인·IaC로 대체해 운영 효율을 개선

  • Infra 설계 및 운영: Kubernetes, GPU 클러스터, 네트워크 등 인프라 레이어의 설계·운영에 참여하고 지속적으로 개선

  • Capacity 검증 지원: 신규 인프라 배포 시 East-West/North-South 네트워크 구성, 스토리지 처리량, BIOS·펌웨어 설정 등 운영 관점의 요구사항을 정의하고 검증에 참여

Qualifications

  • 컴퓨터공학, 전자공학, AI 등 관련 전공 학사 이상 학위 혹은 이에 준하는 실무 경험

  • 2년 이상의 Software Engineering 경험

  • Linux 시스템에 대한 이해와 트러블슈팅 경험

  • Python, Go 등의 언어로 자동화 스크립트나 도구를 작성해 본 경험

  • Kubernetes, Docker 등 컨테이너 오케스트레이션에 대한 기본적인 이해

  • 온프레미스 또는 퍼블릭 클라우드 인프라 운영 경험, 혹은 이에 준하는 학습·실습 경험

  • 장애 상황에서 데이터를 기반으로 원인을 좁혀나가는 문제 해결 능력

  • On-call 로테이션에 참여 가능하신 분

Helpful experience (not required)

  • GPU/TPU/NPU 등 가속기 기반으로 상용 학습/인퍼런스 시스템을 구축/운영해 본 경험

  • Slurm, Kubernetes 기반 대규모 GPU 스케줄링 환경 운영 경험

  • NVIDIA GPU Operator를 활용해 Kubernetes 클러스터의 GPU 드라이버·패브릭 버전을 관리해 본 경험

  • Prometheus, Grafana, OpenTelemetry 등 Observability 스택을 구축해 본 경험

  • InfiniBand/RoCEv2 등 고성능 네트워크 기반 멀티노드 클러스터 환경 운영 경험

  • Terraform, Ansible 등 IaC(Infrastructure as Code) 툴 활용 경험

Life & Benefit

함께 변화를 만들어갈 수 있도록, 도전과 성장을 지원

  • 연간 최대 120만원 한도 내 온/오프라인 자기계발 지원

  • 연 1회 미국 시장 경험 기회 제공 (Global Exposure pass)

  • 성장에 필요한 도서 실물 구매 지원 또는 전자도서관 이용

  • 구성원 간의 1on1 음료 지원

업무 생산성을 높여 몰입할 수 있는 환경

  • 오전 8시~11시 사이 선택하는 시차출퇴근제 운영

  • 이니셔티브 중심의 조직 목표와 Align되어 몰입하는 협업 방식

  • 월 1회 Allhands + Team Gathering 통한 업무 공유

  • 늦은 시간까지 근무 시, 야근식대/택시비 지원

  • 구성원 간의 1on1 음료 지원

몰입한 만큼 휴식과 생활 편의 지원

  • 개인 간식비 지원 (월 한도)

  • 장기근속자 리프레시 휴가 제공

  • 종합건강검진비 및 휴가 지원 (연 1회)

  • 입사 N주년 축하 선물 제공

  • 생일 반차 휴가 제공

  • 명절 선물, 각종 휴가 및 경조금 지원

  • 본인 및 배우자 출산휴가비 지원

Joinning Process

서류전형 → Coding Test → Technical Interview → Resume/Culture Interview → CEO Interveiw 순으로 진행됩니다.

  • 위 내용은 베슬에이아이코리아 경력 채용 기본 프로세스이며, 경우에 따라 절차가 가감될 수 있습니다.

    • 지원서 (경력 세부 기술) 및 포트폴리오 (또는 Git 링크)를 필수로 제출해주세요. (양식 자유)

    • Technical Interview 는 개발 실무자가 참여하며, 구조 설계 및 구현 방식 등 기술 중심의 대화로 문제 해결 역량을 파악하는 시간으로 최대 2시간 정도 소요됩니다.

    • Resume/Culture Interview 는 유관 경험 중심의 기술 역량 및 문화적 핏을 알아보는 시간으로 소속 매니저와 팀 멤버가 참여하며 각 1시간 정도 소요됩니다.

    • 경력직의 경우, 인터뷰 마지막 단계 이후 Reference Check 를 진행하고 있습니다.

    • 이력서 및 제출서류에 허위 사실이 발견될 경우, 합격 발표 후라도 입사가 취소될 수 있습니다.

  • 근무 형태

    • 정규직 (수습 3개월)

    • 3개월의 수습 피드백 기간 후, 업무 성과 평가 결과에 따라 최종 합류 여부가 결정됩니다.

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
747,188 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

DevOps
Similar stack
Same company
Seoul
Platform Engineer 27 days ago
≈ $37k – $111k per year (Estimated) • Hybrid • 3+ years exp • Bachelor's Degree • Seoul
Python
SQL
Databases
SAP HANA
AI/ML
LangChain
Model Context Protocol
NLP
PyTorch
Hugging Face
Context Engineering
arXiv
Agentic Workflows
Multi-Agent Systems
Tool Use
Machine Learning
DevOps
Docker
Apply
$343k – $686k per year • In office • Tokyo
DevOps
Terraform
CloudFormation
AWS
Kubernetes
Apply
$72k – $100k per year • In office • Full-Time • 3+ years exp • Bachelor's Degree • Upton
Python
C++
AI/ML
Ray
Computer Use
DevOps
Podman
CI/CD
Docker
Kubernetes
Linux
Unix
Apply
≈ $63k – $151k per year (Estimated) • Remote (United States) • Full-Time • 2+ years exp • Bachelor's Degree
Python
Bash
DevOps
Terraform
CI/CD
AWS
Kubernetes
GitLab
Amazon ECS
Linux
DNS
Management
Agile
Apply
≈ $35k – $94k per year (Estimated) • In office • Full-Time • Tokyo
Python
JavaScript
Java
Kotlin
Ruby
Java
Spring Boot
Ruby
Ruby on Rails
Databases
PostgreSQL
Redis
ElasticSearch
Google BigQuery
BigQuery
AI/ML
Copilot
Cursor
Claude
ChatGPT
Claude Code
Dify
Model Context Protocol
Vertex AI
Cline
Gemini
OpenAI
Devin
Frontend
Vue.js
DevOps
Terraform
GitLab CI
Azure
CI/CD
AWS
Docker
GitHub
Windows
Design
Figma
Management
Slack
Confluence
Google Workspace
Apply
≈ $34k – $93k per year (Estimated) • In office • Full-Time • Tokyo
Python
JavaScript
Java
PHP
Kotlin
Ruby
Java
Spring Framework
Ruby
Ruby on Rails
Databases
PostgreSQL
Redis
ElasticSearch
Frontend
Vue.js
JQuery
DevOps
GitLab CI
Azure
CI/CD
AWS
Docker
AWS Lambda
Amazon EC2
GitHub
Amazon S3
Amazon ECS
Windows
Management
Slack
Confluence
Apply
≈ $18k – $52k per year (Estimated) • Hybrid • Full-Time • Wrocław
Python
JavaScript
SQL
DevOps
Rest API
VMWare
CI/CD
AWS
Management
Jira
ServiceNow
Scrum
Kanban
QA
Selenium
Playwright
Apply
≈ $90k – $170k per year (Estimated) • In office • Full-Time • Hamburg • Frankfurt am Main • Munich • Berlin
Python
SQL
Apex
Apex
Salesforce Data Cloud
Databases
Snowflake
Databricks
Google BigQuery
BigQuery
AI/ML
AI Agents
Apply
$85k – $108k per year • In office • Full-Time • Lisbon
Python
Java
DevOps
GCP
AWS
Docker
Cybersecurity
Burp Suite
ISO 27001
OWASP Top 10
PCI DSS
Threat Modeling
Apply
Hybrid • Full-Time • 1+ year exp • Seoul
Python
PowerShell
DevOps
Git
Kubernetes
Windows
TCP/IP
DNS
DHCP
VPN
Wi-Fi
Cybersecurity
Okta
Microsoft Entra ID
Management
Linear
Slack
Jira
Google Workspace
Apply
≈ $55k – $137k per year (Estimated) • Hybrid • Full-Time • 5+ years exp • Seoul
Python
AI/ML
TPU
InfiniBand
NVLink
DevOps
SLURM
Git
Docker
Kubernetes
Chaos Engineering
eBPF
SLI/SLO/SLA
Linux
Apply
≈ $26k – $60k per year (Estimated) • Hybrid • Full-Time • Seoul
Python
AI/ML
vLLM
AI Agents
SGLang
TensorRT
TensorRT-LLM
InfiniBand
DevOps
Git
Linux
Chips/EDA
PoC Library
Apply
≈ $53k – $117k per year (Estimated) • Hybrid • Full-Time • 5+ years exp • Seoul
Python
AI/ML
vLLM
AI Agents
SGLang
TensorRT
TensorRT-LLM
InfiniBand
DevOps
Git
HPC
Linux
Chips/EDA
PoC Library
Apply
In office • Full-Time • Seoul
SQL
DevOps
Git
Analytics
Tableau
Looker
Marketing
Salesforce
HubSpot
Apply
In office • Seoul
SQL
Databases
MySQL
Oracle
Amazon Aurora
DevOps
AWS
Apply
≈ $51k – $114k per year (Estimated) • In office • 4+ years exp • Seoul
Management
Google Sheets
Apply
≈ $35k – $85k per year (Estimated) • In office • 3+ years exp • Bachelor's Degree • Seoul
Apply
≈ $49k – $108k per year (Estimated) • In office • 3+ years exp • Bachelor's Degree • Seoul
SQL
Analytics
Power BI
Looker
Microsoft Excel
Apply
≈ $35k – $85k per year (Estimated) • In office • 3+ years exp • Bachelor's Degree • Seoul
Apply
See all jobs
This is one of many
747,188 more open roles from verified company boards, updated every day.