1,059,896open jobs
62,082companies
176,030added this week
Browse all
Salary
≈ $30k – $78k per year (Estimated)
Location
Hybrid (Bengaluru, India)
Seniority
Staff · 8+ years exp

First seen by Alion on Sep 18, 2026.

Overview
Company
Impact
Profile match
Value Lane specializes in the recruitment of highly specialized and strategic talent for mid to senior-level positions for businesses or divisions with a focus on AI, Data, Analytics, Quantum, Digital, Technology, and related fields across industries.

Role Overview :

As a Staff Engineer for Platform Reliability Engineering, you will serve as a technical anchor for our mission-critical infrastructure, ensuring our systems remain resilient, scalable, and performant at massive scale. You will operate at the intersection of software engineering and systems operations, collaborating closely with product engineering teams, security architects, and leadership to define the roadmap for platform stability. By architecting robust backend systems and implementing advanced observability frameworks, you will directly influence the uptime and reliability of services that support millions of users, ultimately driving business growth through operational excellence and technical innovation in our engineering hub.

Key Responsibilities :

- Architect and implement highly available backend systems on AWS to ensure seamless service delivery and fault tolerance for global customers.

- Lead the design and deployment of advanced observability services and monitoring tools to proactively identify bottlenecks and reduce mean time to resolution.

- Drive the evolution of our platform architecture by mentoring senior engineers and establishing best practices for system reliability and performance engineering.

- Optimize cloud infrastructure costs and resource utilization by implementing automated scaling policies and efficient backend architecture patterns.

- Partner with cross-functional teams to conduct deep-dive post-mortems and implement systemic improvements that prevent recurring incidents and enhance overall platform health.

Required Skillset :

- Demonstrated expertise in designing and maintaining complex, distributed backend architectures with a focus on high-concurrency and low-latency requirements.

- Advanced proficiency in Python for building automation tools, custom monitoring agents, and infrastructure-as-code solutions.

- Deep hands-on experience managing and scaling cloud-native environments on AWS, including mastery of VPC networking, IAM, and managed database services.

- Proven ability to implement and manage enterprise-grade observability stacks, translating raw telemetry data into actionable insights for engineering stakeholders.

- Exceptional communication skills, with the ability to articulate complex technical trade-offs to non-technical stakeholders and influence senior leadership decisions.

- Strong collaborative mindset, capable of fostering a culture of reliability and engineering rigor within a high-growth, hybrid work environment.

- A Bachelor's or Master's degree in Computer Science or a related field, complemented by 8 - 15 years of progressive experience in site reliability or platform engineering roles.

Mandatory - Platform Reliability :

- 8+ years of software engineering experience, with substantial time owning production systems and their reliability.

- Demonstrated ownership of SLOs, error budgets and on-call for business-critical systems.

- Proven incident command experience on serious, customer-impacting outages, and a track record of eliminating repeat causes.

- Strong resilience engineering instincts - failure modes, degradation strategies, capacity planning and performance tuning.

- Exceptional debugging ability across application, database, messaging, network and infrastructure layers.

Mandatory - Full Stack Engineering :

- Deep backend engineering expertise in Java, Python, Go or equivalent, writing production-quality code rather than scripting alone.

- Strong, current frontend capability with React/TypeScript or equivalent.

- Proven track record designing distributed systems - APIs, data modelling, event-driven architecture, concurrency and consistency trade-offs.

- Comfortable reading and changing application code across the stack to fix reliability problems at source.

Mandatory - LLM & Agentic AI :

- Hands-on experience building and running LLM-powered or Agentic AI systems in production, including their operational failure modes.

- Practical understanding of prompt and context strategies, RAG, tool use and agent orchestration.

- Experience evaluating and testing AI systems where the output is not deterministic.

- Fluent use of Claude Code, Cursor or equivalent AI development tools in daily engineering work.

Mandatory - Leadership & Ways of Working :

- Track record of technical influence beyond personal output - standards adopted, teams unblocked and engineers levelled up.

- Ability to work independently with ambiguous requirements and bring clarity to others.

- Strong written communication across design documents, RFCs, post-mortems and runbooks.

- Bias towards owning outcomes rather than tickets, and towards fixing causes rather than symptoms.

Desirable :

- IoT, energy, industrial or real-time data experience.

- Kafka, MQTT, time-series databases or stream processing at scale.

- Experience operating AI-first products or Agentic AI systems in production.

- Experience building a platform or reliability function from 0 to 1.

- Security engineering, compliance or disaster recovery and business continuity experience.

- Experience in high-velocity product or startup environments.

Skills

AWS, System Reliability, Observability Services, Python, IT Automation Tools, Full Stack, Native Cloud, Cloud Infrastructure

Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
1,059,896 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account Continue with Google
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Backend
Similar stack
Same company
Bengaluru
≈ $121k – $223k per year (Estimated) • In office • 8+ years exp • Conyers
Go
Java
Kotlin
C#
Mobile
Dependency Injection
DevOps
Rest API
Terraform
GCP
Azure DevOps
OpenTelemetry
Azure
CI/CD
AWS
Kubernetes
Grafana
Bicep
Apply
Software Developer 1 day ago
$49k – $63k per year • Hybrid • Secret • Full-Time • 2+ years exp • Ottawa
DevOps
CI/CD
Management
Agile
Apply
$70k – $88k per year • Hybrid • Secret • Full-Time • 10+ years exp • Ottawa
DevOps
CI/CD
Management
Agile
Apply
Software Architect 9 hours ago
$81k – $88k per year • Remote (Canada) • Secret • Full-Time • 10+ years exp • Bachelor's Degree • Ottawa
DevOps
Kubernetes
Linux
Unix
Management
Agile
Apply
≈ $134k – $248k per year (Estimated) • In office • 8+ years exp • Bachelor's Degree • United States
Go
Databases
PostgreSQL
CockroachDB
Apache Kafka
DevOps
Datadog
AWS
AWS Lambda
Amazon S3
Amazon ECS
API Gateway
Apply
≈ $79k – $210k per year (Estimated) • Equity • In office • Full-Time • 3+ years exp • Bachelor's Degree • Vancouver
Python
Java
C++
C++
TensorFlow C++
PyTorch C++
AI/ML
AI Agents
TensorFlow
PyTorch
RAG
Hugging Face
Recommender Systems
Machine Learning
DevOps
GCP
AWS
Management
Agile
Apply
≈ $121k – $224k per year (Estimated) • In office • 8+ years exp • Palm Beach Gardens
Python
JavaScript
TypeScript
SQL
Databases
Apache Iceberg
DynamoDB
Amazon Redshift
AI/ML
AWS Bedrock
LLM
Frontend
React.js
DevOps
Rest API
Terraform
GCP
AWS CDK
CloudFormation
Azure
CI/CD
AWS
Docker
AWS Fargate
AWS Lambda
FinOps
GitHub
Amazon S3
IAM
Amazon ECS
Amazon CloudWatch
API Gateway
Management
ServiceNow
Apply
$36k – $43k per year • In office • 4+ years exp • Bachelor's Degree • Moscow
Go
JavaScript
TypeScript
Frontend
Zustand
Redux
Tailwind CSS
React.js
Vite
React Query
Radix UI
Storybook
Material UI
Ant Design
shadcn/ui
Mantine
Recharts
Redux Toolkit
Mobile
State Management
DevOps
Rest API
gRPC
WebSockets
GitLab CI
GitHub
VPN
Cybersecurity
ViPNet
Design
Figma
QA
Playwright
Vitest
Apply
$148k – $222k per year • Equity • In office • Full-Time • 7+ years exp • Toronto
Python
JavaScript
Python
Flask
Django
Databases
MySQL
PostgreSQL
Frontend
React.js
DevOps
CI/CD
Git
AWS
AWS Lambda
Amazon EC2
Amazon S3
Management
Agile
Scrum
Apply
≈ $82k – $160k per year (Estimated) • In office • 4+ years exp • Bachelor's Degree • New York
Python
SQL
Databases
Snowflake
AI/ML
Google AI Studio
DevOps
AWS
Analytics
Power BI
Apply
Python Developer 8 days ago
≈ $10k – $30k per year (Estimated) • In office • 2+ years exp • Bengaluru
Python
SQL
Python
Flask
FastAPI
Django
DevOps
Rest API
Azure
Git
Apply
≈ $17k – $45k per year (Estimated) • Hybrid • 5+ years exp • Bengaluru
Python
JavaScript
Node JS
AI/ML
Prompt Engineering
LLM
RAG
Frontend
React.js
Apply
Full Stack Developer 10 days ago
≈ $17k – $44k per year (Estimated) • In office • 5+ years exp • Bengaluru • Hyderabad
JavaScript
Node JS
Frontend
React.js
DevOps
Azure
CI/CD
Jenkins
AWS
Docker
Kubernetes
Azure AKS
Apply
≈ $17k – $44k per year (Estimated) • In office • 4+ years exp • Bengaluru
Python
Python
Flask
Databases
Redis
ElasticSearch
AI/ML
LLM
RAG
DevOps
CI/CD
Git
Docker
Linux
Management
Agile
Apply
≈ $17k – $44k per year (Estimated) • In office • 6+ years exp • Bengaluru • Chennai • Pune
Python
JavaScript
SQL
Node JS
Python
Flask
FastAPI
Django
Databases
PostgreSQL
DevOps
GitHub Actions
Azure
CI/CD
GitHub
Management
Agile
Apply
Hybrid • Full-Time • 4+ years exp • Bengaluru
Java
Java
Spring Boot
Databases
RabbitMQ
Apache Kafka
AI/ML
Copilot
Cursor
Windsurf
OpenAI Codex
DevOps
Podman
Docker
Twelve-Factor App
Apply
≈ $26k – $57k per year (Estimated) • In office • Full-Time • 12+ years exp • Bachelor's Degree • Bengaluru
JavaScript
TypeScript
SQL
C#
C#
ASP.NET Core
Entity Framework Core
Blazor
xUnit
Databases
PostgreSQL
Redis
Weaviate
Milvus
Pinecone
RabbitMQ
MS SQL
ElasticSearch
AI/ML
Copilot
LangGraph
LangChain
MLFlow
Embeddings
Prompt Engineering
AI Agents
Llama
Kubeflow
LLM
RAG
Anomaly Detection
Amazon SageMaker
Human-in-the-Loop
Frontend
Redux
GraphQL
Next.js
Angular
React.js
DevOps
Terraform
Azure DevOps
WebSockets
Azure
CI/CD
Jenkins
AWS
Docker
Kubernetes
Bicep
Cybersecurity
GDPR
QA
Jest
Apply
≈ $9.5k – $24k per year (Estimated) • Hybrid • Full-Time • 5+ years exp • Bachelor's Degree • Bengaluru
Python
Java
C#
DevOps
Azure DevOps
Azure
CI/CD
Jenkins
Management
Outlook
Agile
Scrum
Microsoft Office
QA
Selenium
Cucumber
Apply
≈ $24k – $60k per year (Estimated) • Hybrid • Full-Time • 5+ years exp • Bachelor's Degree • Bengaluru
Python
C
C
Valgrind
Embedded C
AI/ML
Copilot
LLM
DevOps
CI/CD
Git
BGP
MPLS
Web3
Layer 2
Apply
≈ $24k – $57k per year (Estimated) • In office • Full-Time • 4+ years exp • Bachelor's Degree • Bengaluru
Python
C++
Chips/EDA
LTspice
Cadence Allegro
Ansys SIwave
HyperLynx
Apply
See all jobs
This is one of many
1,059,896 more open roles from verified company boards, updated every day.