{"id":1459706,"url":"https://alion.io/job/awtg-devops-engineer","title":"DevOps Engineer","company":{"id":1907331,"name":"AWTG","domain":"awtg.co.uk","url":"https://alion.io/company/awtg","size_band":null,"is_staffing_agency":false,"employer_type":"direct","is_intermediary":false,"listed_via":null,"ats_vendor":"Schema","truth_index":{"grade":"D","score":42,"open_postings":37,"ghost_share":0.973,"stale_share":0,"repost_share":0,"time_to_fill_p50_days":null,"computed_at":"2026-10-03T05:45:00Z"}},"role":"DevOps","role_family":"DevOps","seniority":null,"employment_type":null,"work_mode":"on_site","remote_scope":null,"remote_scope_basis":null,"remote_working_hours":null,"hiring_geo_confidence":"structured","locations":[],"countries":[],"hiring_countries":[],"hiring_countries_total":0,"salary":null,"salary_estimate":null,"experience_years_min":null,"visa_sponsorship":false,"relocation_package":false,"has_equity":false,"technologies":[{"name":"AWS","optional":false},{"name":"Azure","optional":false},{"name":"GCP","optional":false},{"name":"Bash","optional":true},{"name":"Blue-Green Deployment","optional":true},{"name":"CI/CD","optional":true},{"name":"Configuration Management","optional":true},{"name":"Embeddings","optional":true},{"name":"FinOps","optional":true},{"name":"GitHub Actions","optional":true},{"name":"GitLab CI","optional":true},{"name":"Go","optional":true},{"name":"Google Cloud Run","optional":true},{"name":"Grafana","optional":true},{"name":"Hallucination","optional":true},{"name":"IAM","optional":true},{"name":"Kubernetes","optional":true},{"name":"Least Privilege","optional":true},{"name":"Linux","optional":true},{"name":"LLM","optional":true},{"name":"LLMOps","optional":true},{"name":"OpenAI","optional":true},{"name":"Pinecone","optional":true},{"name":"Progressive Delivery","optional":true},{"name":"Prometheus","optional":true},{"name":"Python","optional":true},{"name":"RAG","optional":true},{"name":"Redis","optional":true},{"name":"SLI/SLO/SLA","optional":true},{"name":"Terraform","optional":true},{"name":"Vertex AI","optional":true}],"status":"live","first_seen_at":"2026-03-16T10:53:38Z","employer_posted_date":"2026-06-26","last_verified_at":"2026-10-03T17:36:09Z","board_verified":true,"closed_at":null,"days_open":201,"trust":{"level":"ghost","repost_count":0,"flags":["stale","company_stale"],"days_open":200},"description":"DevOps Engineer Role (Multi-Cloud: GCP primary, Azure/AWS nice to have)You will shape and evolve a DevOps toolchain that enables reliable product delivery across a mixed estate: GCP (VMs and Cloud Run), on-prem, and multi-cloud patterns. You will work closely with delivery teams to design repeatable, secure, scalable deployment strategies, improve operational performance, and reduce manual effort through automation and infrastructure as code.\nYou will also support the deployment and operation of AI-enabled platforms, including LLM-based services, RAG applications, vector database integrations, and AI observability tooling to monitor performance, reliability, latency, cost, and quality across production AI workloads.\nAbout AWTG\nAWTG is a global technology partner delivering secure, scalable, mission-critical SaaS platforms that help organisations innovate with confidence. Established in 2006 and headquartered in London, AWTG brings deep experience across telecoms, smart cities, Industry 4.0, cloud, data, AI, and digital governance, supporting public and private sector clients worldwide.\nQuality, security, and operational excellence are central to how we work. Our services align with recognised international standards and best practice, supported by certifications including ISO/IEC 27001, ISO/IEC 20000-1, ISO/IEC 42001, ISO 9001, and Cyber Essentials Plus, with independent CREST-accredited penetration testing.\nWe operate as a full lifecycle partner, covering advisory, architecture and design, rollout and integration, and long-term support and operations, with proven delivery at enterprise scale. Our platforms are engineered for performance, automation, and insight, incorporating AI-powered analytics, LLM-enabled services, RAG-based knowledge systems, and multi-cloud architectures, supported by robust programme governance and auditable controls.\nWhat you will do\nDesign and implement a coherent DevOps toolchain that enables safe, repeatable delivery across GCP (VMs/Cloud Run), on-prem, and multi-cloud practices (Azure/AWS).\nBuild and improve CI/CD pipelines, such as GitHub Actions or GitLab CI, focusing on deployment repeatability, speed, and risk reduction.\nEstablish and maintain infrastructure as code, including Terraform and related patterns, reducing manual configuration and improving consistency.\nSupport the deployment and operation of LLM-based applications, including APIs, inference services, orchestration layers, prompt/configuration management, and supporting cloud services.\nDeploy and support RAG-based systems, including document ingestion pipelines, embedding workflows, vector databases, retrieval services, and integration with LLM application layers.\nImplement AI deployment observability practices, monitoring key indicators such as model/API latency, token usage, cost, error rates, retrieval performance, hallucination risk signals, user feedback, and end-to-end request tracing.\nImprove reliability and availability through proactive capacity planning, performance tuning, and resilience patterns, such as rollback strategies, blue/green, and canary, where appropriate.\nStrengthen security posture by embedding security controls into designs and pipelines, including IAM, secrets, least privilege, supply-chain controls, and auditability.\nLead incident investigation and fault resolution; improve operational maturity through runbooks, post-incident reviews, and preventative actions.\nPartner with engineering, AI/ML, and product teams to plan and design large groups of stories, translating requirements into delivery and operational work.\nDrive development process optimisation with teams, identifying improvement opportunities and helping implement pragmatic changes.\nImplement and evolve observability practices, including metrics, logs, and traces, using Prometheus/Grafana and cloud-native equivalents to reduce MTTR and improve SLO performance.\nSupport systems design and integration across services, coordinating integration builds and supporting integration testing activities.\nDevelop and maintain scripts/tools of medium-to-high complexity to automate build, release, environment management, AI deployment workflows, and operational tasks.\nMentor and coach junior engineers through pairing, reviews, standards, and knowledge sharing, without line management responsibility.\nWhat you will bring - must-have\nHands-on DevOps experience delivering secure, reliable services in production environments.\nStrong GCP experience, including compute on VMs and serverless, IAM, networking, monitoring, and operational tooling, with the ability to design for scale and availability.\nExperience deploying or supporting AI/LLM-enabled applications in production or pre-production environments.\nUnderstanding of RAG deployment patterns, including document ingestion, embeddings, vector databases, retrieval APIs, and integration with LLM services.\nExperience implementing observability for AI-enabled systems, including logs, metrics, traces, latency monitoring, error monitoring, usage tracking, and cost visibility.\nProven CI/CD capability, using GitHub Actions and/or GitLab CI, including secure pipeline design and automated release strategies.\nStrong Infrastructure as Code experience, especially Terraform, plus experience migrating away from manual/console-heavy estates.\nSolid Linux and networking fundamentals, including hybrid connectivity considerations between cloud and on-prem.\nPractical information security engineering mindset: least privilege IAM, secrets management, secure build/release, and audit-ready change controls.\nObservability experience using Prometheus/Grafana and/or cloud-native monitoring to drive actionable operational insight.\nStrong troubleshooting and service support ability: diagnosing incidents, fixing faults, and improving stability through prevention and automation.\nAbility to design and review systems with medium risk, impact, and complexity, selecting appropriate standards, methods, and tools.\nStrong scripting/programming ability, such as Python, Go, or Bash, with disciplined testing and documentation.\nCollaborative ways of working: able to translate requirements into delivery plans, work across teams, and represent user needs in technical decisions.\nNice to have\nExperience operating or deploying workloads across container and orchestration platforms, such as Kubernetes, as well as serverless and VM estates.\nExperience with LLMOps, MLOps, or AI platform operations, including model gateway patterns, evaluation pipelines, prompt/config deployment, and AI service monitoring.\nExperience with vector databases or retrieval platforms such as Pinecone, Redis, Vertex AI Search, Azure AI Search, or similar.\nExperience supporting AI platforms or services such as Vertex AI, Azure AI Foundry, OpenAI/Azure OpenAI, or comparable LLM providers.\nExperience with policy-as-code and security tooling, such as OPA, SAST/DAST, container scanning, and dependency scanning.\nExperience with release strategies such as canary, blue-green, and progressive delivery patterns.\nCost optimisation experience across cloud platforms, including FinOps mindset, budgeting/forecasting, right-sizing, workload efficiency, and AI/LLM usage cost control.\nExperience building internal developer platforms, golden paths, templates, or paved-road approaches.\nFamiliarity with service management practices, including runbooks, SLAs/SLOs, and incident/problem/change processes.\nExperience coordinating cross-system integration builds and supporting integration testing at scale.\nExperience designing for regulated environments or formal assurance frameworks.\nExperience working with data pipelines, search/retrieval systems, or knowledge-base platforms used by AI applications.","description_format":"text","description_chars":7755,"description_truncated":false,"requirements":{"experience_years_min":null,"management_years_min":null,"team_size_min":null,"manages_managers":false,"education":null,"security_clearance":false,"languages":[]},"benefits":[],"hiring_locations":[],"hiring_excludes":[],"relocation_offered":false,"industries":["Cybersecurity","Information Technology","Science & Engineering"],"lifecycle":[{"event":"open","at":"2026-09-29T11:41:06Z"}],"visa":[{"country":"GB","licensed_sponsor":true,"evidence":"Licensed UK visa sponsor (Skilled Worker)","filings_12m":null,"filings_prev_12m":null,"green_card_filings_12m":null,"median_offered_wage_usd":null,"route":"Skilled Worker","cap_exempt":false,"checked_at":"2026-10-03T21:09:13+00:00","sources":["UK Home Office: register of licensed sponsors (workers)"],"filings_for_role_12m":0}],"liveness":{"score":2,"band":"cold","label":"Long shot","p_open":0.9,"p_active":0.09,"p_room":0.28,"age_days":200,"expected_fill_days":21,"reasons":["conf:83","stale_co","ghost","win:tail","crowd:"],"computed_at":"2026-10-03T05:45:00Z"},"pay":null,"html_url":"https://alion.io/job/awtg-devops-engineer","json_url":"https://alion.io/job/awtg-devops-engineer.json","meta":{"generated_at":"2026-10-04T00:57:50Z","cache_seconds":300,"methodology":"https://alion.io/methodology","terms":"https://alion.io/terms","contact":"https://alion.io/contact","api":"https://alion.io/developers","usage":{"tier":"crawler","counted_by":"address","units_charged":1,"used_today":1246,"day_limit":5000,"remaining_today":3754,"minute_limit":60,"resets_at":"2026-10-05T00:00:00Z"}}}