368,530open jobs
9,432companies
50,439added this week
Browse all
Location
Remote/Hybrid (Seoul, South Korea)
Seniority
Senior · 5+ years exp
Employment
Full-Time
Overview
Company
Impact
Profile match
Gauss Labs is an industrial artificial intelligence company founded in 2020 by SK hynix to apply machine learning to semiconductor manufacturing. Its products improve metrology, defect detection and yield analysis inside fabrication plants. The company operates research teams in Korea and the United States.

Gauss Labs is an industrial AI company on a mission to revolutionize manufacturing with AI, starting with the semiconductor sector. Panoptes is an AI-based virtual metrology solution deployed in high-volume manufacturing fabs, helping customers improve yield, reduce costs, and accelerate production. Our software runs in our customers' own managed environments, and we're seeking a Site Reliability Engineer to own the reliability of the infrastructure and platform that Panoptes runs on. You will keep the platform available, performant, and scalable; own monitoring, alerting, incident first-response, and the on-call rotation; and build the automation and observability that let engineering teams operate their services safely.

Responsibilities

  • Platform reliability and operations: Own platform-layer reliability across both environments. In our internal cloud environment: full ownership - cluster health, resource management (CPU/memory/OOM), scheduling, autoscaling, Kubernetes/EKS lifecycle. In the customer environment: operate directly at the application-namespace level and for the customer-controlled cluster/node layer, diagnose and clearly communicate what's needed, and operate the platform within their setup, decisions, and constraints.
  • Monitoring and Alerting: Build and maintain robust monitoring and alerting for the infrastructure and platform layer to proactively identify and resolve issues before they impact the platform.
  • Incident Response: Own incident first-response for the platform layer and participate in the on-call rotation to minimize downtime and restore service quickly.
  • Automation: Develop automation tools and scripts to streamline operations, reduce manual effort, and enable engineering teams to operate their own services safely.
  • Capacity Planning: Forecast resource needs, optimize resource utilization, and ensure the platform infrastructure can handle increasing workloads.
  • Deployment infrastructure: Build and maintain CI/CD pipelines and deployment infrastructure for the platform.
  • Continuous Improvement: Drive a culture of continuous improvement by identifying opportunities to enhance platform reliability, performance, and efficiency.

Basic Qualifications

  • Bachelor's degree in computer science, engineering, or a related discipline
  • 5+ years of industry experience as a Site Reliability Engineer or in platform/infrastructure engineering
  • Hands-on experience operating Kubernetes in production (EKS preferred): cluster lifecycle, scheduling, autoscaling, resource management
  • Experience with cloud platforms (AWS preferred) and containerization technologies (Docker, Kubernetes)
  • Experience with observability and alerting tools (Prometheus, Grafana, ElasticSearch, Jaeger)
  • Experience with scripting languages (Python, Bash)
  • Working knowledge of GitHub, GitHub Actions, and CI/CD concepts
  • Strong problem-solving and troubleshooting skills
  • Working proficiency in English for internal documentation and technical coordination

Preferred Qualifications

  • Knowledge of AI/ML infrastructure and workloads.
  • Knowledge of database technologies (MongoDB, PostgreSQL)
  • Experience operating software in customer-managed (on-prem or customer-cloud) environments
  • Exposure to manufacturing, semiconductor, or enterprise B2B customer environments
Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
368,530 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
Seoul
$34k – $72k per year (Estimated) • In office • Contractor • Toulouse
DevOps
Amazon EC2
Amazon EKS
AWS
AWS Lambda
Azure
CI/CD
CloudFormation
GitLab CI
Terraform
Kubernetes
Amazon ECS
Amazon S3
GitLab
IAM
Apply
$28k – $65k per year (Estimated) • In office • Full-Time • 12+ years exp • Bengaluru
Bash
PowerShell
Python
Node JS
JavaScript
Node JS
Commander.js
AI/ML
AI Agents
DevOps
Amazon EC2
Amazon EKS
AWS
Azure
Kubernetes
Amazon ECS
IAM
Cybersecurity
Crowdstrike
Zero Trust
Apply
AI Engineer 1 day ago
$25k – $103k per year (Estimated) • In office • Full-Time • Bachelor's Degree • Gurgaon
Python
SQL
Databases
Databricks
Microsoft Fabric
AI/ML
AI Agents
Embeddings
Gemini
Hallucination
LangChain
LangGraph
LLM
Multimodal AI
Prompt Engineering
PyTorch
RAG
Semantic Search
Spark
TensorFlow
Hugging Face
LLM Guardrails
LLMOps
OpenAI
Semantic Search
DevOps
AWS
Azure
CI/CD
Apply
$71k – $170k per year (Estimated) • In office • Full-Time • Netanya
Python
TypeScript
AI/ML
Accelerate
Fine-tuning
LangChain
LLM
NLP
Prompt Engineering
PyTorch
RAG
Edge AI
Hugging Face
AI Agents
DevOps
AWS
Azure
Docker
GCP
Kubernetes
Apply
$18k – $64k per year (Estimated) • In office • Full-Time • 2+ years exp • Toulouse
C#
SQL
C#
.NET
Entity Framework Core
Databases
Azure SQL Database
DevOps
Azure
Azure DevOps
CI/CD
Docker
Apply
Remote/Hybrid • Full-Time • 10+ years exp • Seoul
Python
Databases
OpenSearch
DevOps
Grafana
Kubernetes
Prometheus
Management
Jira
Apply
Remote/Hybrid • Full-Time • 3+ years exp • Bachelor's Degree • Seoul
Python
SQL
AI/ML
Edge AI
DevOps
AWS
Azure
GCP
Apply
$144k – $278k per year (Estimated) • Remote/Hybrid • Full-Time • 8+ years exp • Bachelor's Degree • Palo Alto
Python
AI/ML
Claude
Claude Code
Copilot
Fine-tuning
NumPy
Pandas
PyTorch
Scikit-learn
TensorFlow
AI Agents
DevOps
AWS
Azure
CI/CD
Docker
GCP
Git
Kubernetes
Apply
Remote/Hybrid • Full-Time • 3+ years exp • PhD • Seoul
Python
AI/ML
JAX
PyTorch
TensorFlow
Transformers
Time Series Forecasting
Apply
Product Manager 1 day ago
In office • Full-Time • 5+ years exp • Bachelor's Degree • Seoul
Apply
In office • Full-Time • 5+ years exp • Bachelor's Degree • Seoul
Python
DevOps
AWS
CI/CD
GCP
Marketing
Salesforce
Apply
In office • Full-Time • Seoul
C++
SystemVerilog
Verilog
Apply
In office • Contractor • Bachelor's Degree • Seoul
AI/ML
Mamba
Multimodal AI
TPU
Apply
See all jobs
This is one of many
368,530 more open roles from verified company boards, updated every day.