The Role
We’re looking for Forward Deployed Site Reliability Engineers who can help us build, operate, and maintain high-performance, scalable, and reliable services for our production infrastructure, primarily across on-prem environments for the US Government. Forward Deployed Site Reliability Engineers combine engineering experience and an innate drive to improve existing systems and processes, with the creativity to develop novel solutions to evolving challenges. Our team strives to automate processes wherever possible, using whichever tools are best for the job. You’ll travel to various locations where you will be the expert for Palantir’s infrastructure, helping partner teams build & configure their hardware and network for software to operate reliably within.
We strongly believe in engineering teams being responsible for the operations of their services in production. In this role, you’ll work closely with engineers to advocate and participate in sensible, scalable, systems design and share responsibility with them in diagnosing, resolving, and preventing production issues.
Core Responsibilities
- Maintaining availability of physical Linux servers that power the Palantir platform in air-gapped production environments
- Design, deploy, and operate infrastructure to support customer & product requirements via modern orchestration & monitoring platforms
- Collaborate closely with product teams on requirements & SLOs for deploying software into air-gapped environments
- Identifying, troubleshooting, and solving network & systems issues
- Scripting to automate away routine operational tasks
- Provide technical troubleshooting support for production issues, ensuring timely resolution and minimal impact on operations. Participate in a support on-call schedule
What We Value
- Confidence in troubleshooting complex systems issues independently using stack traces and observability & systems tools
- Comfort with configuration management, load balancing, monitoring & alerting infrastructure, and container orchestration on small hardware form factors.
- Demonstrated ability to continuously learn and work independently, making decisions with minimal supervision while working in secure facilities
- Experience with containers (Docker/Podman) and orchestration (OpenShift/Kubernetes) at scale is a plus
- Preferred Certifications: DOD 8570 IAT Level II or greater (CISSP, Sec+), Unix/Linux Computing Environment (e.g Linux+, RHCE)
What We Require
- Available for 50% travel (domestic and international)
- 4+ years of experience with Linux system administration (RHEL or equivalent preferred)
- Experience with hardware environments, including setup, configuration, and management of physical servers and networking equipment
- Familiarity with monitoring systems using tools like Prometheus and writing health checks
- Proficiency with at least one programming or scripting language, such as Java, Go, Python, JavaScript, Bash, or similar languages.
- Strong engineering background, preferred in fields such as Computer Science, Mathematics, Software Engineering, Physics, and Data Science.
- Active US Security Clearance at or above the Top Secret level

