{"id":1516256,"url":"https://alion.io/job/pearson-specialist-observability-engineering","title":"Specialist, Observability Engineering","company":{"id":34029,"name":"Pearson","domain":"pearson.com","url":"https://alion.io/company/pearson","size_band":"1001-5000","is_staffing_agency":false,"employer_type":"direct","is_intermediary":false,"listed_via":null,"ats_vendor":"Oracle","truth_index":{"grade":"B","score":79,"open_postings":19,"ghost_share":0,"stale_share":0.842,"repost_share":0,"time_to_fill_p50_days":34,"computed_at":"2026-10-01T05:45:00Z"}},"role":"DevOps","role_family":"DevOps","seniority":"senior","employment_type":null,"work_mode":"on_site","remote_scope":null,"remote_scope_basis":null,"remote_working_hours":null,"hiring_geo_confidence":"structured","locations":["Bengaluru, India"],"countries":["IN"],"hiring_countries":[],"hiring_countries_total":0,"salary":null,"salary_estimate":{"min_usd":18000,"max_usd":41000,"period":"year","method":"role_seniority_country_remote_cell","sample_n":42},"experience_years_min":5,"visa_sponsorship":false,"relocation_package":false,"has_equity":false,"technologies":[{"name":"AppDynamics","optional":false},{"name":"Datadog","optional":false},{"name":"Dynatrace","optional":false},{"name":"Grafana","optional":false},{"name":"Loki","optional":false},{"name":"New Relic","optional":false},{"name":"OpenTelemetry","optional":false},{"name":"Platform Engineering","optional":false},{"name":"Prometheus","optional":false},{"name":"Splunk","optional":false},{"name":"AIOps","optional":true},{"name":"Anomaly Detection","optional":true},{"name":"AWS","optional":true},{"name":"Azure","optional":true},{"name":"Azure DevOps","optional":true},{"name":"CI/CD","optional":true},{"name":"Configuration Management","optional":true},{"name":"GCP","optional":true},{"name":"GitHub Actions","optional":true},{"name":"GitLab CI","optional":true},{"name":"IAM","optional":true},{"name":"Incident Management","optional":true},{"name":"ITSM","optional":true},{"name":"Jenkins","optional":true},{"name":"Jira","optional":true},{"name":"Kubernetes","optional":true},{"name":"OpenTofu","optional":true},{"name":"Opsgenie","optional":true},{"name":"PagerDuty","optional":true},{"name":"Python","optional":true},{"name":"ServiceNow","optional":true},{"name":"Terraform","optional":true}],"status":"live","first_seen_at":"2026-09-30T06:06:15Z","employer_posted_date":"2026-09-30","last_verified_at":"2026-10-01T17:45:21Z","board_verified":true,"closed_at":null,"days_open":1,"trust":{"level":"ok","repost_count":null,"flags":[],"days_open":1},"description":"Function: Observability Engineering / SRE / Cloud Engineering\nAbout the Role\nWe are looking for an Observability Engineer to join our Observability Engineering team and help build, standardize, and operate a modern enterprise observability platform across cloud and application environments.\nThe ideal candidate will have strong hands-on engineering experience with New Relic, Grafana, and modern observability technologies, combined with solid foundations in cloud engineering, automation, DevOps, and SRE practices.\nThis role goes beyond dashboard creation and monitoring operations. You will be responsible for engineering scalable observability solutions, defining telemetry standards, building reusable monitoring capabilities, improving application and infrastructure visibility, and enabling engineering teams to adopt observability as part of their software delivery lifecycle.\nYou will work closely with SRE, Cloud Engineering, Platform Engineering, Application Engineering, and Operations teams to build a consistent and scalable observability experience across the organization.\nKey Responsibilities\n1. Observability Engineering\nDesign, implement, and maintain enterprise-grade observability solutions across applications, infrastructure, cloud platforms, and services.\nBuild and maintain monitoring, alerting, dashboards, service health views, and operational telemetry.\nDevelop standardized observability patterns for metrics, logs, traces, events, and application performance monitoring.\nImplement observability solutions using New Relic, Grafana, and other industry-standard tools.\nDevelop reusable dashboards, alerts, instrumentation patterns, and observability components.\nEstablish observability standards and best practices across engineering teams.\nContinuously improve signal quality by reducing alert noise, false positives, and non-actionable alerts.\n2. New Relic Engineering\nHands-on engineering experience with New Relic APM, Infrastructure Monitoring, Browser Monitoring, Synthetic Monitoring, Logs, Distributed Tracing, NRQL, Alerts, Workloads and Dashboards.\nDesign and implement New Relic monitoring and alerting strategies for enterprise applications.\nDevelop complex NRQL queries, alert conditions, dashboards, and operational views.\nConfigure and optimize New Relic agents and integrations.\nImplement application and infrastructure instrumentation.\nDevelop reusable New Relic configurations and automation using APIs/IaC where appropriate.\nParticipate in New Relic platform governance, licensing optimization, and standardization.\nEvaluate and implement emerging New Relic capabilities to improve engineering productivity and reliability.\n3. Grafana & Visualization\nBuild and maintain operational dashboards using Grafana.\nIntegrate Grafana with multiple telemetry and data sources.\nDesign effective dashboards for application health, infrastructure, SRE, NOC, and executive operational visibility.\nDevelop visualization standards and reusable dashboard templates.\nUnderstand the difference between visualization, monitoring, alerting, and observability, and apply each appropriately.\n4. OpenTelemetry & Modern Observability\nExperience with OpenTelemetry and modern telemetry architectures.\nImplement and manage telemetry collection for metrics, logs, and traces.\nUnderstand distributed tracing and service dependency mapping.\nWork with telemetry pipelines, collectors, agents, exporters, and integrations.\nExperience with technologies such as Prometheus, Loki, Elastic, Splunk, Datadog, Dynatrace, AppDynamics, or similar observability platforms is desirable.\nEvaluate new observability technologies and recommend solutions based on scalability, cost, reliability, and engineering value.\n5. Cloud Engineering\nStrong cloud engineering fundamentals are expected, including experience with one or more major cloud platforms:\nAWS\nMicrosoft Azure\nGoogle Cloud Platform\nExperience should include:\nCompute, networking, storage, databases, containers, and cloud-native services.\nCloud monitoring and logging.\nIAM and security fundamentals.\nInfrastructure automation.\nCloud-native architecture and operational best practices.\nTroubleshooting distributed cloud environments.\n6. Automation & Infrastructure as Code\nAutomate repetitive observability and operational activities.\nDevelop scripts and tools using Python, Bash, Go, or similar languages.\nUse Terraform / OpenTofu or equivalent Infrastructure as Code technologies.\nBuild reusable automation for dashboards, alerts, instrumentation, integrations, and configuration management.\nIntegrate observability capabilities into CI/CD pipelines.\n7. DevOps & CI/CD\nExperience with modern CI/CD practices and tools.\nHands-on experience with GitHub Actions, Jenkins, GitLab CI, Azure DevOps, or similar platforms.\nIntegrate observability and quality gates into deployment pipelines.\nImplement deployment markers and release health monitoring.\nEnable automated validation of application and infrastructure health following deployments.\n8. SRE & Reliability Engineering\nApply SRE principles to improve system reliability and operational maturity.\nDefine and monitor SLIs, SLOs, and error budgets.\nParticipate in incident investigation and root-cause analysis.\nDevelop proactive monitoring and reliability solutions.\nIdentify reliability gaps and engineer solutions to eliminate recurring incidents.\nSupport capacity, performance, availability, and resilience engineering.\nRequired Skills & Experience\nMust Have\n5+ years of experience in Observability, SRE, DevOps, Cloud Engineering, Platform Engineering, or a related engineering discipline.\nStrong hands-on experience with New Relic.\nStrong hands-on experience with Grafana.\nExperience building production-grade dashboards, monitoring and alerting solutions.\nStrong understanding of APM, infrastructure monitoring, logging, metrics, tracing, and distributed systems.\nExperience with NRQL and New Relic alerting.\nExperience with OpenTelemetry is highly desirable.\nStrong cloud engineering experience in AWS, Azure, or GCP.\nStrong scripting/programming experience in Python, Bash, Go, or similar.\nExperience with Terraform/OpenTofu or another Infrastructure-as-Code technology.\nUnderstanding of CI/CD and DevOps practices.\nUnderstanding of SRE principles, SLIs, SLOs, incident management, and reliability engineering.\nStrong troubleshooting and analytical skills.\nGood to Have\nNew Relic certifications or equivalent hands-on expertise.\nGrafana/Prometheus experience.\nOpenTelemetry implementation experience.\nKubernetes and container observability.\nExperience with Prometheus, Loki, Elastic, Splunk, Datadog, Dynatrace, AppDynamics or similar platforms.\nExperience designing enterprise observability architectures.\nExperience with observability platform migrations or consolidation.\nExperience with observability cost optimization and licensing governance.\nExperience developing observability-as-code.\nExperience integrating observability with ServiceNow, Jira, PagerDuty, Opsgenie, or similar ITSM/incident platforms.\nExperience with AI-assisted observability, AIOps, anomaly detection, or automated incident investigation.\nWhat Success Looks Like\nIn this role, you will be successful when you can:\nEngineer rather than simply operate monitoring.\nBuild scalable observability solutions that can be reused across hundreds of applications.\nTurn raw telemetry into meaningful engineering and operational insights.\nReduce alert noise and improve signal quality.\nStandardize New Relic and Grafana adoption across engineering teams.\nAutomate observability configuration and reduce manual operational work.\nImprove application reliability through better telemetry and proactive detection.\nEnable engineering teams to own more of their operational health.\nHelp establish Observability as a Platform/Capability, rather than simply another monitoring tool.\nBehavioral & Leadership Skills\nStrong ownership and accountability.\nAbility to work across engineering, application, infrastructure, and operations teams.\nStrong communication and stakeholder management skills.\nAbility to explain complex technical concepts to both engineers and non-technical stakeholders.\nStrong problem-solving and analytical mindset.\nComfortable working in a fast-paced, enterprise engineering environment.\nPassion for automation, engineering excellence, reliability, and continuous improvement.\nAbility to challenge existing approaches and introduce better engineering practices.","description_format":"text","description_chars":8463,"description_truncated":false,"requirements":{"experience_years_min":5,"management_years_min":null,"team_size_min":null,"manages_managers":false,"education":null,"security_clearance":false,"languages":[]},"benefits":[],"hiring_locations":[],"hiring_excludes":[],"relocation_offered":false,"industries":["EdTech Platforms","Publishing","E-learning"],"lifecycle":[{"event":"open","at":"2026-09-30T09:42:33Z"}],"liveness":{"score":63,"band":"ok","label":"Likely open","p_open":1,"p_active":0.632,"p_room":1,"age_days":0,"expected_fill_days":34,"reasons":["conf:10","stale_co","velocity","win:early","comp:brand"],"computed_at":"2026-10-01T05:45:00Z"},"pay":null,"html_url":"https://alion.io/job/pearson-specialist-observability-engineering","json_url":"https://alion.io/job/pearson-specialist-observability-engineering.json","meta":{"generated_at":"2026-10-01T19:17:35Z","cache_seconds":300,"methodology":"https://alion.io/methodology","terms":"https://alion.io/terms","contact":"https://alion.io/contact","api":"https://alion.io/developers","usage":{"tier":"crawler","counted_by":"address","units_charged":1,"used_today":1802,"day_limit":5000,"remaining_today":3198,"minute_limit":60,"resets_at":"2026-10-02T00:00:00Z"}}}