{"id":1270716,"url":"https://alion.io/job/realm-labs-software-engineer-ml-infrastructure","title":"Software Engineer, ML Infrastructure","company":{"id":2203168,"name":"Realm Labs","domain":"realmlabs.ai","url":"https://alion.io/company/realmlabs","size_band":"11-50","is_staffing_agency":false,"employer_type":"direct","is_intermediary":false,"listed_via":null,"ats_vendor":"Gem","truth_index":null},"role":"AI/ML","role_family":"AI/ML","seniority":"senior","employment_type":"full_time","work_mode":"on_site","remote_scope":null,"remote_scope_basis":null,"remote_working_hours":null,"hiring_geo_confidence":"structured","locations":["Sunnyvale, United States"],"countries":["US"],"hiring_countries":[],"hiring_countries_total":0,"salary":{"min":180000,"max":250000,"currency":"USD","period":"year","gross":null,"usd_annual":250000},"salary_estimate":null,"experience_years_min":5,"visa_sponsorship":true,"relocation_package":false,"has_equity":false,"technologies":[{"name":"KV Cache","optional":false},{"name":"Linux","optional":false},{"name":"LLM","optional":false},{"name":"PyTorch","optional":false},{"name":"SGLang","optional":false},{"name":"TensorFlow","optional":false},{"name":"TensorRT","optional":false},{"name":"TensorRT-LLM","optional":false},{"name":"Transformers","optional":false},{"name":"Triton Inference Server","optional":false},{"name":"vLLM","optional":false},{"name":"CUDA","optional":true},{"name":"CUDA Toolkit","optional":true},{"name":"Kubernetes","optional":true},{"name":"NCCL","optional":true}],"status":"live","first_seen_at":"2026-01-16T21:45:33Z","employer_posted_date":"2026-01-19","last_verified_at":"2026-09-26T06:33:54Z","board_verified":true,"closed_at":null,"days_open":253,"trust":{"level":"stale","repost_count":0,"flags":["stale"],"days_open":252},"description":"Role Overview\nWe are hiring a Founding ML Infrastructure Engineer to own the end-to-end deployment, optimization, and operation of our suits of models in production.\nThis is a core founding role focused on building and operating production-grade LLM systems. You will apply deep knowledge of model internals to deploy, optimize, and run modern LLMs at scale, owning performance end-to-end across latency, throughput, and reliability.\nYou will design and operate the full ML serving stack from model artifacts to GPU execution, and work closely with Product and ML teams to ensure our models can support high QPS, strict SLAs, and production correctness.\nThis role is ideal for someone who deeply understands how LLMs work internally, but chooses to specialize in making them fast, stable, and production-ready.\nAbout Realm Labs\nRealm Labs is an AI trust and security startup. We help enterprises detect, debug, and prevent AI’s misbehaviors in production. We are backed by top VCs and serve some of the most iconic global enterprises.\nKey Responsibilities\nOwn the end-to-end LLM inference stack, including:Model loading and execution\nGPU utilization and memory efficiency\nRuntime performance tuning\nProduction deployment and scaling\n\nDesign and operate high-performance LLM serving systems using technologies such as:vLLM, TensorRT / TensorRT-LLM, Triton Inference Server, SGLang\n\nOptimize inference across:Latency\nThroughput (QPS)\nGPU memory footprint\nCost efficiency\n\nWork hands-on with PyTorch and TensorFlow models, including:Model graph understanding\nAttention mechanisms, KV cache behavior, batching strategies\nPrecision tradeoffs (FP16, BF16, INT8, etc.)\n\nBuild and maintain production-grade GPU services:Multi-model serving\nAutoscaling strategies\nFault isolation and graceful degradation\n\nCollaborate with application and platform teams to:Define serving APIs\nEnsure correctness and safety of outputs\nDebug production issues end-to-end\n\nBuild a reproducible model training and versioning system for customer deployments\nEstablish best practices for:Model versioning\nRollouts and rollbacks\nPerformance benchmarking\nProduction validation\n\nExpected Qualifications\n5+ years of professional experience in ML infrastructure, systems engineering, or production ML roles.\nStrong software engineering fundamentals; ability to write robust, maintainable production code.\nDeep hands-on experience with LLM inference infrastructure, including:PyTorch (required)\nTensorFlow (working knowledge)\n\nProven experience with GPU inference optimization, including:TensorRT / TensorRT-LLM\nvLLM\nTriton Inference Server\nSGLang or similar serving runtimes\n\nStrong understanding of LLM internals, such as:Transformer architectures\nAttention and KV caching\nBatching, streaming, and token-level generation\n\nExperience running ML systems in production with high traffic and SLAs\nComfortable working in Linux-based, cloud production environments\nPreferred Qualifications\nExperience deploying LLMs on Kubernetes and GPU clusters.\nFamiliarity with CUDA, NCCL, or low-level GPU performance concepts.\nExperience with:Model sharding and parallelism strategies\nMulti-GPU inference\nStreaming inference systems\n\nKnowledge of observability for ML systems (metrics, latency breakdowns, GPU monitoring).\nExperience working at startups or owning systems with minimal abstraction layers.\nAdditional Information\nThis is a founding, high-ownership role with direct impact on core product capabilities.\nYou will be expected to build, run, and own systems end-to-end.\nThe role may include limited on-call responsibilities aligned with production ownership.\nCompensation & Benefits\nMarket aligned compensation and benefits\nFounding engineer equity (Equity is a significant component of this role and will be discussed)\nMedical, Dental, Vision, Life insurance, 401-K, In-office lunch etc.\nVisa sponsorship: We do sponsor visas! However, we aren't able to successfully sponsor visas for every role and candidate. But if we make you an offer, we will make all reasonable effort to get you a visa, and we retain an immigration lawyer to help with this.\nCompensation\nThe base pay range for this role is $180,000 - $250,000 per year.","description_format":"text","description_chars":4187,"description_truncated":false,"requirements":{"experience_years_min":5,"management_years_min":null,"team_size_min":null,"manages_managers":false,"education":null,"security_clearance":false,"languages":[]},"benefits":["Equity","Life insurance"],"hiring_locations":[],"hiring_excludes":[],"relocation_offered":false,"industries":["Artificial Intelligence","Cybersecurity","Cybersecurity AI"],"lifecycle":[{"event":"open","at":"2026-09-25T23:59:00Z"}],"liveness":{"score":7,"band":"cold","label":"Long shot","p_open":1,"p_active":0.252,"p_room":0.28,"age_days":252,"expected_fill_days":33,"reasons":["conf:5","win:tail","crowd:"],"computed_at":"2026-09-26T05:45:00Z"},"pay":{"stated_usd_annual":250000,"is_top_pay":true},"html_url":"https://alion.io/job/realm-labs-software-engineer-ml-infrastructure","json_url":"https://alion.io/job/realm-labs-software-engineer-ml-infrastructure.json","meta":{"generated_at":"2026-09-27T03:02:08Z","cache_seconds":300,"methodology":"https://alion.io/methodology","terms":"https://alion.io/terms","contact":"https://alion.io/contact","api":"https://alion.io/developers","usage":{"tier":"crawler","counted_by":"address","units_charged":1,"used_today":2780,"day_limit":5000,"remaining_today":2220,"minute_limit":60,"resets_at":"2026-09-28T00:00:00Z"}}}