{"id":1200889,"url":"https://alion.io/job/openteams-platform-engineer-aiml-infrastructure","title":"Platform Engineer, AI/ML Infrastructure","company":{"id":678978,"name":"OpenTeams","domain":"openteams.com","url":"https://alion.io/company/openteams","size_band":"51-200","is_staffing_agency":false,"employer_type":"direct","is_intermediary":false,"listed_via":null,"ats_vendor":"Greenhouse","truth_index":null},"role":"AI/ML","role_family":"AI/ML","seniority":"middle","employment_type":null,"work_mode":"hybrid","remote_scope":null,"remote_scope_basis":null,"remote_working_hours":null,"hiring_geo_confidence":"structured","locations":["Austin, United States"],"countries":["US"],"hiring_countries":[],"hiring_countries_total":0,"salary":{"min":120000,"max":250000,"currency":"USD","period":"year","gross":null,"usd_annual":250000},"salary_estimate":null,"experience_years_min":4,"visa_sponsorship":false,"relocation_package":false,"has_equity":false,"technologies":[{"name":"AWS","optional":false},{"name":"Azure","optional":false},{"name":"CI/CD","optional":false},{"name":"GCP","optional":false},{"name":"GitOps","optional":false},{"name":"Helm","optional":false},{"name":"Kubernetes","optional":false},{"name":"OpenTofu","optional":false},{"name":"Prometheus","optional":false},{"name":"Pulumi","optional":false},{"name":"Python","optional":false},{"name":"Terraform","optional":false},{"name":"Agentic Workflows","optional":true},{"name":"KServe","optional":true},{"name":"LLM","optional":true},{"name":"NIST 800-53","optional":true},{"name":"NumPy","optional":true},{"name":"PyTorch","optional":true},{"name":"SBOM","optional":true},{"name":"SciPy","optional":true},{"name":"Travis CI","optional":true},{"name":"vLLM","optional":true}],"status":"live","first_seen_at":"2026-09-24T18:47:16Z","employer_posted_date":"2026-09-25","last_verified_at":"2026-09-29T19:43:58Z","board_verified":true,"closed_at":null,"days_open":5,"trust":{"level":"ok","repost_count":null,"flags":[],"days_open":5},"description":"Who We Are\nEvery organization runs on intelligence: years of accumulated knowledge, decisions, and context. As AI takes on more of that work, companies face a choice: rent that intelligence from vendors who keep the data, the context, and the results, or own it.\nOpenTeams exists to make ownership possible.\nFounded by Travis Oliphant, creator of NumPy and SciPy, and built by people with deep roots across the open-source ecosystem, including NumPy, SciPy, PyTorch, and Jupyter, we help enterprises and governments build AI they control, govern, and evolve themselves.\nIf that sounds like your kind of work, we'd like to meet you.\nPlatform Engineer, AI/ML Infrastructure\nLocation: U.S - Remote OR Hybrid - Washington, DC, Denver, CO or Colorado Springs, CO. \nWork Authorization: U.S. citizenship required\nClearance: U.S.-Remote Opening: An active clearance is not required. Candidates must be eligible and willing to obtain and maintain a U.S. security clearance. Hybrid Opening: An active TS/SCI clearance is preferred. Candidates may also be considered if they previously held a TS/SCI with CI polygraph or currently hold an active TS or Secret clearance.\nSalary Range: $120,000-$250,000 USD, dependent on experience level and location\nOpenings: Two positions are available:\nOne hybrid position: Candidates must be located in or willing to work hybrid from Washington, DC; Denver, CO; or Colorado Springs, CO. An active TS/SCI clearance is required. Up to 15% travel is required.\nOne U.S.-remote position: Candidates may work remotely from anywhere in the United States. An active clearance is not required, but candidates must be willing and able to undergo the process required to obtain and maintain a U.S. security clearance.\nCandidates will be considered for the opening that best aligns with their location, clearance status, experience, and work preferences. Candidates who meet the requirements for multiple openings may be considered for more than one.\nCandidates will be considered for the opening that best aligns with their location, clearance status, experience, and work preferences. Candidates who meet the requirements for multiple openings may be considered for more than one.\nAbout the Role\nOpenTeams builds AI platforms that governments and enterprises own outright: the infrastructure, the data, the models, and the evidence that the whole thing does what it claims. We're hiring several engineers to build and run that infrastructure.\nThe work spans the full depth of an AI platform. The Kubernetes clusters that schedule GPU workloads, move large datasets, and keep tenants isolated from one another. The services that make it a platform rather than a cluster: workflow orchestration, data ingest, model serving, policy enforcement, audit logging. The delivery path that gets released into production reliably and can prove what it shipped. The cloud infrastructure that ties it all together, and the operational practices that keep a distributed system resilient.\nMuch of this has to run where you can't assume normal cloud resources, or even an internet connection. That constraint is the interesting part of the job. Portability, reproducibility, and operability are design inputs from the first commit rather than problems handed downstream.\nWe build on open source and contribute back. Kubernetes, Terraform and OpenTofu, Argo, Prometheus, Nebari, and others. Upstream work is part of the job, not something you do on weekends.\nThis posting covers multiple roles, spanning mid-level through senior. We understand nobody spans every area above, so tell us where you fit. We assign level-based roles based on what you've actually done rather than a year count.\nKey Responsibilities\nBuild and operate Kubernetes-based infrastructure for demanding AI/ML workloads, including GPU scheduling, resource management, and multi-tenant isolation\nDesign and implement platform services for orchestration, data ingest, model serving, and results management behind documented APIs\nWrite infrastructure as code and build GitOps pipelines so environments are reproducible from source\nBuild and operate CI/CD pipelines that produce versioned, signed, scanned release artifacts along with the documentation needed to deploy them\nOwn reliability: capacity planning, upgrade paths, failure-mode analysis, backup and recovery, incident response, and postmortems\nImplement monitoring, logging, tracing, and alerting, and define the service level objectives they're measured against\nDeploy and validate the platform in restricted, disconnected, or limited-connectivity environments, and verify parity after each release\nKeep the platform portable by constraining dependencies to what's confirmed available in target environments\nWrite runbooks and operational documentation that other engineers can execute without you in the room\nContribute to Nebari and other open-source infrastructure, Kubernetes, and MLOps projects\nWork with security engineers, government stakeholders, and other engineers to turn requirements into systems that hold up\nCollaborate asynchronously across a distributed team\nRequired Skills & Experience\nU.S. citizenship, and the ability to obtain and maintain a U.S. security clearance\nFour or more years of hands-on experience building or operating production infrastructure, platforms, or distributed systems\nProduction experience with Kubernetes and containerized workloads\nExperience with at least one major cloud platform: AWS, Azure, or Google Cloud\nExperience with infrastructure as code and CI/CD, using tools such as Terraform, OpenTofu, Pulumi, or Helm\nWorking proficiency in Python, Go, Bash, or a comparable language\nExperience implementing or operating production monitoring and observability\nAbility to write documentation, runbooks, and deployment procedures that other people can actually follow\nAbility to work independently and collaborate well in a remote, distributed team\nNice to Have\nYou will not have all of these, and very few people will. They're the things that would help, not a checklist. If the required list above describes you, apply.\nAn active U.S. security clearance, particularly TS/SCI with CI polygraph\nExperience deploying or operating software in air-gapped, disconnected, or otherwise restricted environments\nExperience with Department of Defense, Intelligence Community, or comparably regulated programsFamiliarity with the Risk Management Framework, NIST 800-53 or 800-171, or similar frameworks, and with producing the evidence they require\nExperience supporting an Authorization to Operate, or with continuous ATO modelsSupply chain security work: hardened images, artifact signing, SBOM generation, dependency and container scanning, policy enforcement\nFamiliarity with cross-domain solutions, guards, data diodes, or similar transfer mechanisms\nExperience with classified cloud environments, including AWS Secret or Top Secret regionsA DoD 8140/8570 qualifying certification such as Security+, CISSP, CASP+, or CISM, or willingness to obtain one after joining\nExperience building MLOps pipelines or infrastructure for AI/ML workloads\nExperience with GPU scheduling, distributed inference, or large-scale data and evaluation pipelines\nExperience with model-serving or gateway frameworks such as KServe, vLLM, or LLM-DExperience designing API-first services and vendor-agnostic platforms that run across multiple environments\nExperience with agentic workflow frameworks or multi-step AI pipeline orchestration\nContributions to open-source Kubernetes, infrastructure, MLOps, or observability projects, and experience with Nebari specifically\nFamiliarity with data sovereignty and privacy requirements for enterprise or government AI systems\nExperience leading technical initiatives, setting engineering standards, or mentoring other engineers\nExperience supporting rapid prototyping programs or defense innovation initiatives\nGrow With Us\nAt OpenTeams, growth isn’t just about the company-it’s about you.\nWe believe the best careers are built at the edge of your potential. That is where new tools, ideas, and technologies change the world. Here, you’ll work alongside pioneers of AI, solving problems that matter: making AI more transparent, more ethical, and more empowering. As your skills grow, our career framework provides a pathway and recognition of that increased impact.\nOpportunities aren’t limited by geography. You’ll collaborate with global experts, contribute to open source projects that power the world’s technology, and stretch your skills daily. That global perspective and diversity makes our solution more universal and robust. We are committed to continuing to celebrate diversity on our team.\nSupported people are successful people. We offer 100% employer paid medical premiums for employees and self-managed PTO with a minimum time off requirement, so that our teams are able to do their best work.\nWe invest in curiosity, creativity, and ownership. That means you’ll be trusted to boldly innovate, supported to learn fast, and celebrated for successful collaboration.\nCommitment to diversity, equity, inclusion, and belonging\nOpenTeams understands that valuing diverse creative practices and forms of knowledge is crucial to and enriches the company’s core mission. We encourage applications from everyone, including members of all equity-seeking communities, such as (but certainly not limited to) women, racialized and Indigenous persons, disabled people, persons of all sexual orientations, gender identities and expressions.\nWe are an equal opportunity employer - all qualified applicants will receive equal consideration for recruitment, interviews, employment, training, compensation, promotion, and related activities. We do not discriminate based on race, religion, gender, gender identity, gender expression, color, national origin, pregnancy, ancestry, domestic partner status, disability, sexual orientation, age, genetic predisposition, medical condition, marital status, citizenship status, military or veteran status, or any other basis covered by applicable laws. OpenTeams will not tolerate discrimination or harassment based on these characteristics or any other unlawful behavior, conduct, or purpose.","description_format":"text","description_chars":10215,"description_truncated":false,"requirements":{"experience_years_min":4,"management_years_min":null,"team_size_min":null,"manages_managers":false,"education":{"level":"phd","optional":false},"security_clearance":true,"languages":[]},"benefits":["Equity"],"hiring_locations":[{"name":"United States","iso":"US","kind":"country"},{"name":"Washington","iso":null,"kind":"city"},{"name":"Colorado","iso":null,"kind":"city"}],"hiring_excludes":[],"relocation_offered":false,"industries":["AI Consulting & Integration","Open Source Projects & Foundations"],"lifecycle":[{"event":"open","at":"2026-09-24T20:47:01Z"}],"liveness":{"score":89,"band":"hot","label":"Hiring now","p_open":1,"p_active":0.891,"p_room":1,"age_days":4,"expected_fill_days":24,"reasons":["conf:4","velocity","win:early"],"computed_at":"2026-09-29T05:45:00Z"},"pay":{"stated_usd_annual":250000,"is_top_pay":true},"html_url":"https://alion.io/job/openteams-platform-engineer-aiml-infrastructure","json_url":"https://alion.io/job/openteams-platform-engineer-aiml-infrastructure.json","meta":{"generated_at":"2026-09-30T00:04:53Z","cache_seconds":300,"methodology":"https://alion.io/methodology","terms":"https://alion.io/terms","contact":"https://alion.io/contact","api":"https://alion.io/developers","usage":{"tier":"search","counted_by":"address","units_charged":0,"used_today":0,"day_limit":null,"remaining_today":null,"minute_limit":null,"resets_at":"2026-10-01T00:00:00Z"}}}