{"id":1293259,"url":"https://alion.io/job/princeton-university-cloud-engineer","title":"Cloud Engineer","company":{"id":15250,"name":"Princeton University","domain":"princeton.edu","url":"https://alion.io/company/princeton","size_band":"201-500","is_staffing_agency":false,"employer_type":"direct","is_intermediary":false,"listed_via":null,"ats_vendor":"iCIMS","truth_index":null},"role":"DevOps","role_family":"DevOps","seniority":"senior","employment_type":null,"work_mode":"remote","remote_scope":"stated_countries","remote_scope_basis":"board_field","remote_working_hours":null,"hiring_geo_confidence":"structured","locations":["Princeton, United States"],"countries":["US"],"hiring_countries":["US"],"hiring_countries_total":1,"salary":{"min":71723,"max":77240,"currency":"USD","period":"year","gross":null,"usd_annual":77240},"salary_estimate":null,"experience_years_min":5,"visa_sponsorship":false,"relocation_package":false,"has_equity":false,"technologies":[{"name":"Anomaly Detection","optional":false},{"name":"Azure","optional":false},{"name":"Azure AKS","optional":false},{"name":"Azure DevOps","optional":false},{"name":"Bicep","optional":false},{"name":"CI/CD","optional":false},{"name":"Databricks","optional":false},{"name":"Docker","optional":false},{"name":"FinOps","optional":false},{"name":"GitHub Actions","optional":false},{"name":"Go","optional":false},{"name":"Grafana","optional":false},{"name":"HIPAA","optional":false},{"name":"Hugging Face","optional":false},{"name":"IAM","optional":false},{"name":"ISO 27001","optional":false},{"name":"Kubernetes","optional":false},{"name":"Loki","optional":false},{"name":"Machine Learning","optional":false},{"name":"Microsoft Entra ID","optional":false},{"name":"MLFlow","optional":false},{"name":"Node JS","optional":false},{"name":"Prometheus","optional":false},{"name":"Python","optional":false},{"name":"React.js","optional":false},{"name":"Rest API","optional":false},{"name":"SOC 2","optional":false},{"name":"Terraform","optional":false},{"name":"TypeScript","optional":false},{"name":"JavaScript","optional":true}],"status":"live","first_seen_at":"2026-09-26T08:34:04Z","employer_posted_date":"2026-09-26","last_verified_at":"2026-09-26T13:36:04Z","board_verified":true,"closed_at":null,"days_open":1,"trust":{"level":"ok","repost_count":null,"flags":[],"days_open":1},"description":"Overview\nThe Accelerator seeks a part-time Cloud Engineer to design, build, and operate the secure cloud infrastructure that powers large-scale academic research on the information environment. Working as part of a small, high-trust cross-functional team, this individual will contribute across the full stack - from infrastructure and DevOps to backend services and data pipelines - and will have meaningful ownership over the technical systems that enable researchers at Princeton and across a global consortium to do their work.\nThis is a role for a senior, self-directing engineer who is equally comfortable designing architecture and writing code, and who takes satisfaction in building systems that are reliable, secure, and well-understood by the people who depend on them. The right candidate brings deep cloud expertise alongside strong software engineering fundamentals - someone who can own infrastructure end to end and contribute meaningfully to application development.\nThis position is classified at the Senior Engineer level, corresponding to 5-8 years of relevant experience. The individual will:\nPlan and execute work independently, applying sound judgment in the evaluation, selection, and adaptation of technical approaches across infrastructure, security, and software development.\nDesign and implement solutions with broad ownership - from initial architecture through deployment and ongoing operations - with supervisory input primarily at the level of objectives and critical decisions rather than day-to-day methods.\nDevise new approaches to novel problems, drawing on extensive knowledge across cloud infrastructure, DevSecOps, and software engineering disciplines.\nServe as a technical resource for the broader team, contributing to architectural decisions and engineering standards.\nThis is a part-time, benefits-eligible, 12-month term position. The prorated base salary range is approximately $71,723/year - $77,240/year.\nA remote work arrangement within the United States may be considered for candidates with the appropriate background and experience.\nResponsibilities\nCloud Infrastructure\nDesign, deploy, and maintain cloud infrastructure on Azure, with responsibility for performance, cost-effectiveness, and reliability across research and production environments.\nArchitect and manage Databricks workspaces, including compute cluster configuration, access controls, and cost optimization for large-scale data processing workflows.\nManage Azure networking, storage, identity (Azure AD / Entra ID), and resource governance across multiple environments.\nImplement infrastructure-as-code using Terraform and/or Bicep; maintain version-controlled, reproducible infrastructure definitions including modules, remote state management, and PR-based workflow.\nDeploy, operate, and maintain AKS clusters running containerized workloads - including containerized data crawlers - managing deploys, scaling, health monitoring, patching, and upgrades.\nAdminister Azure Blob Storage, including lifecycle policies, redundancy configuration, and access tier management.\nManage Azure networking and security, including Private Link, network rules, RBAC, and secrets hygiene across environments.\nOwn Azure cost management: budget alerts, cost/cluster policies, anomaly detection and response, and FinOps practices to keep infrastructure spend predictable and efficient.\nSoftware Development & DevOps\nDesign, build, and maintain backend services, APIs, and data pipelines using Python and/or TypeScript/Node.js.\nDevelop and maintain CI/CD pipelines using GitHub Actions, ensuring reliable and automated delivery of infrastructure and application changes.\nBuild and maintain internal tooling that improves the experience and efficiency of the research and operations teams.\nContribute to frontend integrations where needed; comfortable working across the stack on a small team.\nData Engineering & ML Infrastructure\nDevelop and support data pipelines for ingesting, transforming, and serving large-scale behavioral and social media datasets to researchers.\nImplement and maintain infrastructure for machine learning workflows, including model serving, experiment tracking, and compute resource management.\nSupport integration with ML frameworks and tools (e.g., MLflow, Hugging Face, or equivalent) within the managed environment.\nSecurity & Compliance\nImplement and maintain security controls across all systems, including encryption at rest and in transit, identity and access management, network segmentation, and secrets management.\nDesign and operate environments meeting IRB, data governance, and institutional compliance requirements; ensure adherence to standards equivalent to SOC 2, HIPAA, or ISO 27001 as applicable.\nConduct regular security reviews, vulnerability assessments, and penetration test coordination; manage remediation tracking.\nImplement audit logging, access controls, and data handling procedures for sensitive research data in compliance with IRB protocols and data use agreements.\nObservability & Operations\nOperate, patch, and upgrade the self-hosted observability stack - Grafana (dashboards), Loki (log aggregation), and Prometheus (metrics) - including security patching and version upgrades; implement and maintain alerting, distributed tracing, and platform-wide monitoring.\nOwn incident response, root cause analysis, and operational reliability for production systems.\nDevelop and maintain runbooks, architecture documentation, and operational procedures.\nQualifications\nSkills and Experience\nRequired\n5-8 years of experience in cloud engineering, DevOps, or a software engineering role with significant infrastructure ownership.\nBachelor's degree in Computer Science, Engineering, or a related field or equivalent work experience.\nStrong proficiency in Python; experience with at least one additional language (TypeScript/Node.js, Go, or equivalent).\nDeep hands-on experience with Azure cloud services, including compute, networking, storage, identity, and managed services; familiarity with Azure CAF landing zones, subscription governance, and resource management at scale.\nProficiency with Terraform, including module development, remote state management, and PR-based workflow; Bicep familiarity a plus.\nExperience designing and implementing CI/CD pipelines, preferably using GitHub Actions.\nProduction experience with Kubernetes / AKS - deploys, scaling, health management, upgrades, and cluster operations.\nSolid Docker and container image management skills; experience building and maintaining containerized services in production.\nAzure networking and security fundamentals, including Private Link, network rules, NSGs, and RBAC; comfort managing secrets hygiene across environments.\nAzure cost management and FinOps awareness: budget alerts, cost/cluster policies, anomaly detection and response.\nComfort operating in and improving existing codebases with limited live handoff - able to orient independently, read unfamiliar infrastructure, and contribute quickly without extensive documentation.\nExperience with Databricks or equivalent large-scale data processing platforms.\nSolid understanding of data security principles, IAM patterns, and compliance frameworks (SOC 2, HIPAA, ISO 27001, or equivalent).\nExperience operating a self-hosted observability stack - specifically Grafana, Loki, and Prometheus - including patching, upgrades, and dashboard maintenance; equivalent stack experience considered.\nAbility to work independently on complex, ambiguous problems and communicate technical decisions clearly to non-technical stakeholders.\nStrong written communication skills; comfortable producing architecture documentation, runbooks, and technical specifications.\nPreferred\nExperience supporting research computing or academic data infrastructure environments.\nFamiliarity with ML infrastructure tooling (MLflow, Hugging Face Hub, model serving frameworks).\nExperience with IRB-compliant research data environments or sensitive data handling at scale.\nFrontend development experience (React or equivalent) - useful on a small cross-functional team.\nRelevant certifications: Azure Administrator (AZ-104), Azure Solutions Architect (AZ-305), Azure DevOps Engineer (AZ-400), or equivalent.\nRequirements\nA combination of relevant work experience and education equivalent to 5-8 years of hands-on cloud engineering or software engineering experience, with a demonstrable record of owning and delivering complex infrastructure and software projects. A bachelor's degree in Computer Science, Engineering, or a related field is preferred but not required - equivalent professional experience will be considered.\nAbout the Accelerator\nThe Accelerator at Princeton's School of Public and International Affairs (SPIA) builds shared infrastructure for large-scale, multi-institutional research on the information environment. Our platform - the Accelerator Cloud Environment (ACE) - runs on Azure with AKS-hosted data crawlers, Blob Storage, and a Databricks pipeline, with infrastructure managed as code in Terraform and CI/CD in GitHub Actions. We serve a global consortium of academic researchers and currently support 11 active research projects across 7 institutions. We are a small team that moves quickly, values technical ownership, and cares deeply about the quality and security of the systems we build.\nPrinceton University is an Equal Opportunity Employer and all qualified applicants will receive consideration for employment without regard to age, race, color, religion, sex, sexual orientation, gender identity or expression, national origin, disability status, protected veteran status, or any other characteristic protected by law.\nThe University considers factors such as (but not limited to) scope and responsibilities of the position, candidate's qualifications, work experience, education/training, key skills, market, collective bargaining agreements as applicable, and organizational considerations when extending an offer. The posted salary range represents the University's good faith and reasonable estimate for a full-time position; salaries for part-time positions are pro-rated accordingly.\nIf the salary range on the posted position shows an hourly rate, this is the baseline; the actual hourly rate may be higher, depending on the position and factors listed above.\nThe University also offers a comprehensive benefit program to eligible employees. Please see this link for more information.\nStandard Weekly Hours\n20.00Eligible for Overtime\nNoBenefits Eligible\nYesProbationary Period\n180 daysEssential Services Personnel (see policy for detail)\nNoEstimated Appointment End Date\n9/30/2027Physical Capacity Exam Required\nNoValid Driver’s License Required\nNoExperience Level\nMid-Senior LevelSalary Range\n$130,000 to $140,000","description_format":"text","description_chars":10805,"description_truncated":false,"requirements":{"experience_years_min":5,"management_years_min":null,"team_size_min":null,"manages_managers":false,"education":{"level":"bachelor","optional":true},"security_clearance":false,"languages":[]},"benefits":[],"hiring_locations":[{"name":"United States","iso":"US","kind":"country"}],"hiring_excludes":[],"relocation_offered":false,"industries":["Education","Higher Education"],"lifecycle":[{"event":"open","at":"2026-09-26T08:34:04Z"}],"liveness":{"score":90,"band":"hot","label":"Hiring now","p_open":1,"p_active":0.903,"p_room":1,"age_days":1,"expected_fill_days":13,"reasons":["conf:40","velocity","win:early"],"computed_at":"2026-09-28T05:45:00Z"},"pay":{"stated_usd_annual":77240,"is_top_pay":false},"html_url":"https://alion.io/job/princeton-university-cloud-engineer","json_url":"https://alion.io/job/princeton-university-cloud-engineer.json","meta":{"generated_at":"2026-09-28T06:39:50Z","cache_seconds":300,"methodology":"https://alion.io/methodology","terms":"https://alion.io/terms","contact":"https://alion.io/contact","api":"https://alion.io/developers","usage":{"tier":"crawler","counted_by":"address","units_charged":1,"used_today":4751,"day_limit":5000,"remaining_today":249,"minute_limit":60,"resets_at":"2026-09-29T00:00:00Z"}}}