{"id":1251164,"url":"https://alion.io/job/dsta-ai-service-delivery-manager","title":"AI Service Delivery Manager","company":{"id":247928,"name":"DSTA","domain":"dsta.gov.sg","url":"https://alion.io/company/dsta","size_band":"501-1000","is_staffing_agency":false,"employer_type":"direct","is_intermediary":false,"listed_via":null,"ats_vendor":"Career site","truth_index":null},"role":"Solutions","role_family":"Solutions","seniority":"lead","employment_type":"full_time","work_mode":"on_site","remote_scope":null,"remote_scope_basis":null,"remote_working_hours":null,"hiring_geo_confidence":"structured","locations":["Singapore"],"countries":["SG"],"hiring_countries":[],"hiring_countries_total":0,"salary":null,"salary_estimate":{"min_usd":89000,"max_usd":205000,"period":"year","method":"global_role_cell_scaled_by_country","sample_n":363},"experience_years_min":null,"visa_sponsorship":false,"relocation_package":false,"has_equity":false,"technologies":[{"name":"Grafana","optional":false},{"name":"ITIL","optional":false},{"name":"Kubernetes","optional":false},{"name":"Linux","optional":false},{"name":"OpenTelemetry","optional":false},{"name":"Prometheus","optional":false}],"status":"live","first_seen_at":"2026-09-21T02:25:51Z","employer_posted_date":"2026-09-25","last_verified_at":"2026-09-25T18:34:37Z","board_verified":true,"closed_at":null,"days_open":5,"trust":{"level":"ok","repost_count":null,"flags":[],"days_open":5},"description":"We are seeking an AI Service Delivery Manager to own a portfolio of live AI deployments running in on-premise and private cloud environments. We are looking for someone who believes that keeping AI systems trusted, available and effective in operational use is engineering work in its own right. As the single accountable owner for service health and adoption, you will hold service levels, drive incidents through to resolution, and coordinate upgrades across live operational environments. Working with infrastructure, engineering and user teams, you will ensure AI systems continue to perform reliably in highly controlled environments where internet connectivity is often unavailable, and change must be carefully managed.\nWhat you'll do:\nService Ownership: Own end-to-end service delivery for a portfolio of on-premise and private cloud AI deployments, meeting agreed service levels for availability, inference latency, throughput, and support responsiveness. Maintain an accurate configuration baseline for every environment, and lead regular service reporting and reviews through to senior stakeholders.\nDeployment & Transition to Service: Work with AI products, infrastructure and engineering teams to bring new environments into supported service, from site readiness through cutover, stabilisation, and formal handover. Own the service acceptance process and drive the site-side prerequisites that determine the success of on-premise deployments, including GPU and server availability, power and cooling, network rules, directory integration, and security sign-off.\nIncident & Problem Management: Act as incident manager for major incidents, coordinating engineering, infrastructure, security, and user teams through to resolution. Drive root cause analysis, publish post-incident reviews with tracked corrective actions, and convert recurring failure patterns into permanent fixes or product improvements.\nChange & Release Management: Plan and coordinate AI product and platform upgrades, model version rollouts, patching, and configuration changes across environments with differing maintenance windows and approval forums. Manage the end-to-end release process for on-premise deployments, including offline and air-gapped update bundles, registry mirrors, signed artefacts, alongside hardware dependencies such as GPU drivers, firmware, and version compatibility.\nAI Service Assurance: Monitor model and application performance in production, including latency, throughput, evaluation and regression results, output quality issues, and drift indicators. Identify and escalate service degradation before it impacts users. Ensure evaluation suites are executed before every model or prompt change reaches production, with all changes documented, versioned, and traceable for accreditation and audit.\nAdoption & Continuous Improvement: Monitor adoption and utilisation across the portfolio, identify deployments that stall after go-live, and determine whether the underlying issue relate to training, documentation, capacity or reliability. Own the continual service improvement plan, maintain operational runbooks for restricted environments, and feed structured operational insights back into engineering teams.\nWe are looking for someone who:\nIs organised and meticulous, able to own multiple production environments without sight of the details\nTakes initiative and is comfortable building processes where none yet exist\nRemains calm under pressure and becoming more methodical, not less, during major incidents\nEnjoys collaborating across engineering, security and user organisations\nCommunicates clearly with both engineers and senior stakeholders, building trusted working relationships\nAdapts quickly to ambiguity in a fast-moving environment\nTakes pride in delivering reliable services that others depend on\nRequirements\nDegree in Computer Science, Computer Engineering and AI/ML, or a related discipline.\nOperationally-grounded problem solvers, from backgrounds including but not limited to service delivery, technical account management, implementation, technical programme management, senior technical support, or other operationally focused engineering roles.\nHands-on experience supporting software deployed on infrastructure outside your direct control, with working knowledge of Linux, containers, networking fundamentals, and at least one private cloud or virtualisation platform.\nAble to remain clam under pressure, communicate confidently with senior stakeholders, and drive complex incidents through to resolution.\nResourceful self-starters comfortable with ambiguity and enjoy building processes where none exists.\nStrong collaborator who values clear documentation, honest reporting, and effective communication across technical and operational teams.\nCurious about how AI systems behave in production, with a willingness to continuously learn and influence outcomes across teams without direct authority.\nFamiliarity with Kubernetes, GPU infrastructure, MLSecOps, evaluation tooling, observability platforms (Prometheus, Grafana, OpenTelemetry), infrastructure-as-code, or ITIL v4 will be advantageous.","description_format":"text","description_chars":5135,"description_truncated":false,"requirements":{"experience_years_min":null,"management_years_min":null,"team_size_min":null,"manages_managers":false,"education":{"level":"bachelor","optional":false},"security_clearance":false,"languages":[]},"benefits":[],"hiring_locations":[],"hiring_excludes":[],"relocation_offered":false,"industries":["Military"],"lifecycle":[{"event":"open","at":"2026-09-25T18:34:17Z"}],"liveness":{"score":83,"band":"hot","label":"Hiring now","p_open":1,"p_active":0.83,"p_room":1,"age_days":5,"expected_fill_days":35,"reasons":["conf:9","win:early"],"computed_at":"2026-09-26T03:44:39Z"},"pay":null,"html_url":"https://alion.io/job/dsta-ai-service-delivery-manager","json_url":"https://alion.io/job/dsta-ai-service-delivery-manager.json","meta":{"generated_at":"2026-09-26T03:44:39Z","cache_seconds":300,"methodology":"https://alion.io/methodology","terms":"https://alion.io/terms","contact":"https://alion.io/contact","api":"https://alion.io/developers","usage":{"tier":"crawler","counted_by":"address","units_charged":1,"used_today":4003,"day_limit":5000,"remaining_today":997,"minute_limit":60,"resets_at":"2026-09-27T00:00:00Z"}}}