{"id":1252416,"url":"https://alion.io/job/high-impact-talent-devops-cloud-infrastructure-engineer","title":"DevOps & Cloud Infrastructure Engineer","company":{"id":3800232,"name":"High Impact Talent","domain":"highimpacttalent.com","url":"https://alion.io/company/high-impact-talent","size_band":null,"is_staffing_agency":false,"employer_type":"direct","is_intermediary":true,"listed_via":null,"ats_vendor":null,"truth_index":null},"role":"DevOps","role_family":"DevOps","seniority":"senior","employment_type":null,"work_mode":"on_site","remote_scope":null,"remote_scope_basis":null,"remote_working_hours":null,"hiring_geo_confidence":"structured","locations":["Kolkata, India"],"countries":["IN"],"hiring_countries":[],"hiring_countries_total":0,"salary":{"min":3000000,"max":4000000,"currency":"INR","period":"year","gross":true,"usd_annual":41928},"salary_estimate":null,"experience_years_min":6,"visa_sponsorship":false,"relocation_package":false,"has_equity":false,"technologies":[{"name":"Agile","optional":false},{"name":"Amazon CloudWatch","optional":false},{"name":"Amazon EC2","optional":false},{"name":"Amazon ECS","optional":false},{"name":"Amazon EKS","optional":false},{"name":"Amazon S3","optional":false},{"name":"Apache HTTP Server","optional":false},{"name":"Apache Kafka","optional":false},{"name":"AWS","optional":false},{"name":"AWS Lambda","optional":false},{"name":"Azure","optional":false},{"name":"Blue-Green Deployment","optional":false},{"name":"Caddy","optional":false},{"name":"CI/CD","optional":false},{"name":"Claude Code","optional":false},{"name":"CloudFormation","optional":false},{"name":"Cursor","optional":false},{"name":"DigitalOcean","optional":false},{"name":"DNS","optional":false},{"name":"Docker","optional":false},{"name":"GCP","optional":false},{"name":"Gemini","optional":false},{"name":"Git","optional":false},{"name":"GitHub Actions","optional":false},{"name":"Grafana","optional":false},{"name":"IAM","optional":false},{"name":"Kubernetes","optional":false},{"name":"Least Privilege","optional":false},{"name":"LLM","optional":false},{"name":"Nginx","optional":false},{"name":"PagerDuty","optional":false},{"name":"PM2","optional":false},{"name":"PostgreSQL","optional":false},{"name":"Prometheus","optional":false},{"name":"Pulumi","optional":false},{"name":"Python","optional":false},{"name":"Redis","optional":false},{"name":"Scrum","optional":false},{"name":"Terraform","optional":false},{"name":"Traefik","optional":false},{"name":"JavaScript","optional":true},{"name":"Node JS","optional":true}],"status":"live","first_seen_at":"2026-09-15T09:13:24Z","employer_posted_date":null,"last_verified_at":"2026-09-15T09:13:24Z","board_verified":false,"closed_at":null,"days_open":15,"trust":{"level":"not_scored","repost_count":null,"flags":[],"days_open":15},"description":"Role Summary :\n\nYou will own Mynts infrastructure and delivery lifecycle end-to-end: architect and provision our AWS footprint (EC2, ECS, Lambda, RDS, S3, ECR, Route 53), containerise and orchestrate services with Docker and Kubernetes, and run the data-and-messaging backbone (PostgreSQL, Redis, Kafka).\n\nYoull build the CI/CD pipelines that ship our Node/React/Python stack safely, tame the runtime and web tier (NVM, Pyenv, nginx/Apache, Caddy/Traefik, Certbot), and keep every environment Dev, QA, Staging, and Production fast, observable, and cheap to run.\n\nAbove all, youll bring seasoned, opinionated judgement from multiple cloud ecosystems to recommend the most cost-effective, seamlessly-scalable path for every decision.\n\nThe AI-Augmented Edge :\n\nAt Pyvot, AI is your primary workforce. Use Claude Code, Cursor, and Gemini to draft IaC, generate pipeline configs, review infra diffs, and reason about failure modes targeting a 3x5x efficiency gain.\n\nPractise Cross-LLM Validation: one model proposes the architecture, another stress-tests it for cost, blast-radius, and scaling limits. Your value is measured by uptime, unit economics, and how gracefully the platform scales not tickets closed.\n\nCore Responsibilities :\n\nCloud Architecture & Infrastructure-as-Code : Be the Architect :\n\n- Multi-Service AWS Footprint : Design, provision, and harden EC2, ECS, Lambda, RDS (PostgreSQL), S3, ECR, and Route 53 as reproducible infrastructure not hand-crafted snowflakes.\n\n- Infrastructure-as-Code : Own the estate in Terraform / CloudFormation (or Pulumi / CDK) with modular, reviewed, version-controlled definitions and least-privilege IAM baked in.\n\n- Cloud-Agnostic Judgement : Bring real experience from more than one provider (AWS plus GCP / Azure / DigitalOcean / etc.) to pick the right primitive for the job and avoid needless lock-in.\n\n- Cost-Effective by Design : Right-size compute, exploit spot / reserved / savings plans, and continuously trim the bill without sacrificing reliability treat the cloud invoice as a product metric.\n\nContainers, Orchestration & the Data Backbone :\n\n- Docker & Kubernetes : Containerise every service and run it on Kubernetes (EKS or self-managed) with sane resource limits, autoscaling, health checks, and zero-downtime rollouts.\n\n- Data & Messaging Layer : Operate PostgreSQL (replication, backups, PITR, tuning), Redis (cache / queues), and Kafka (event streaming) as first-class, monitored, recoverable services.\n\n- Runtime & Web Tier : Manage Node (NVM) and Python (Pyenv) runtimes and front the platform with nginx / Apache and Caddy / Traefik, with automated TLS via Certbot / ACME.\n\n- Seamless Scaling : Build horizontal-scaling paths (load balancing, autoscaling groups, HPA) so a 100x traffic jump is a config change, not a fire-drill.\n\nCI/CD, Delivery & Automation :\n\n- Pipelines That Ship Safely : Co-own GitHub Actions (build, test, scan, deploy) with reproducible builds, artifact versioning in ECR, and blue-green / canary rollouts.\n\n- Git-Based, Push-Button Deploys : Standardise deployment across every service and box no manual scp, no editing prod by hand with fast, auditable rollbacks.\n\n- Secrets & Config : Manage secrets and environment config safely (SSM / Secrets Manager / Vault) from a single source of truth, with no credentials in git.\n\n- Environment Parity : Keep Dev, QA, Staging, and Production consistent and disposable so what passes staging behaves in production.\n\nReliability, Observability & Site Reliability Engineering :\n\n- Monitoring & Alerting : Stand up metrics, logs, and traces (CloudWatch, Prometheus / Grafana, ELK, PagerDuty) with actionable alarms not alert fatigue.\n\n- Availability & Recovery : Define and defend SLOs, RTO / RPO targets, backups, and disaster-recovery / failover drills you have actually rehearsed.\n\n- Incident Response : Be first responder for infrastructure incidents diagnose, mitigate, and run blameless post-mortems that harden the system.\n\n- Boot & Persistence Hygiene : Ensure services survive reboots, scaling events, and deploys (pm2 / systemd / docker restart policies) with no silent drift.\n\nAdvisory, Standards & Team Leadership : Be the Trusted Voice :\n\n- Architecture Recommendations : Proactively bring cost / scaling / reliability trade-off proposals to the CTO youre hired for judgement, not just execution.\n\n- Runbooks & Standards : Maintain infrastructure runbooks, deployment standards, and on-call playbooks the whole team can follow.\n\n- Mentoring : Level up engineers on deployment, containers, and operational best-practice; make the platform something anyone can safely ship to.\n\n- Security-Aware Ops : Partner with the CyberSecurity Lead on hardening, IAM, network segmentation, and audit-friendly, deny-delete logging.\n\nRoles Requirements : Intersection of Cloud, Delivery & Reliability\n\nWere looking for a full-spectrum infrastructure engineer equal parts cloud architect, platform / DevOps engineer, and reliability practitioner: someone at ease with a Terraform diff, a Kubernetes manifest, a Postgres failover, and a 2 AM scaling alert and seasoned enough to tell us the smartest, cheapest way to run all of it.\n\nMust-Have Skills & Experience :\n\n- DevOps / Cloud Engineering Experience : 6 - 10 years running production infrastructure for a SaaS / cloud-native product end-to-end.\n\n- AWS Depth : Hands-on with EC2, ECS, Lambda, RDS, S3, ECR, Route 53, IAM, and VPC networking in real production.\n\n- Multi-Cloud Exposure : Genuine experience across more than one cloud / ecosystem (AWS plus GCP / Azure / DO / etc.) able to compare and choose, not just operate one.\n\n- Containers & Orchestration : Strong Docker and Kubernetes (EKS or self-managed) autoscaling, rollouts, and resource management.\n\n- IaC & CI/CD : Terraform / CloudFormation and pipeline ownership (GitHub Actions or equivalent) with automated, rollback-safe deploys.\n\n- Data & Runtime Ops : PostgreSQL operations (backup / replication / tuning), Redis, and the web / runtime tier (nginx / Apache, Caddy / Traefik, Certbot, NVM, Pyenv).\n\n- Cost & Scale Judgement : A track record of cutting cloud spend and scaling systems smoothly with the numbers to prove it.\n\nShould-Have Skills & Experience :\n\n- Event Streaming : Production experience with Kafka (or equivalent) for high-throughput, ordered event pipelines.\n\n- Observability & SRE : CloudWatch / Prometheus / Grafana / ELK / PagerDuty, SLOs, RTO / RPO, and rehearsed DR / failover.\n\n- Networking & TLS : Solid grasp of DNS, load balancing, reverse proxies, and certificate automation.\n\n- Scripting & Automation : Comfortable in Bash plus Python / Node to automate anything that repeats.\n\n- Agile & Cross-Functional Collaboration : Comfortable in Agile / Scrum delivery and working closely with engineering, product, and leadership.\n\nGood-To-Have Skills & Experience :\n\n- Education & Certifications : Degree from a premier institute (IITs / NITs / BITS / IIITs) and/or AWS Solutions Architect / DevOps / CKA / CKAD a strong plus.\n\n- Security-Adjacent Ops : Familiarity with IAM hardening, secrets management, and compliance friendly logging (works alongside our Security Lead).\n\n- AI-Augmented Tooling : Experience using AI / LLM tools (Claude Code, Cursor, Gemini) for IaC drafting, config review, and incident reasoning.\n\n- Startup / Scale-Up Experience : Prior ownership of infrastructure through rapid growth on a lean budget.\nSkills\nDevOps, Cloud Infrastructure, AWS Lambda, Docker, Kubernetes, IAC Terraform, Site Reliability, CloudFormation, Redis, Prometheus","description_format":"text","description_chars":7531,"description_truncated":false,"requirements":{"experience_years_min":6,"management_years_min":null,"team_size_min":null,"manages_managers":false,"education":null,"security_clearance":false,"languages":[]},"benefits":[],"hiring_locations":[],"hiring_excludes":[],"relocation_offered":false,"industries":["Recruiting & Staffing","Job Boards & Aggregators"],"lifecycle":[{"event":"open","at":"2026-09-25T18:04:06Z"}],"liveness":{"score":59,"band":"ok","label":"Likely open","p_open":0.85,"p_active":0.775,"p_room":0.9,"age_days":14,"expected_fill_days":23,"reasons":["seen:14","velocity","win:mid"],"computed_at":"2026-09-30T05:45:00Z"},"pay":{"stated_usd_annual":41928,"is_top_pay":false},"html_url":"https://alion.io/job/high-impact-talent-devops-cloud-infrastructure-engineer","json_url":"https://alion.io/job/high-impact-talent-devops-cloud-infrastructure-engineer.json","meta":{"generated_at":"2026-10-01T04:18:23Z","cache_seconds":300,"methodology":"https://alion.io/methodology","terms":"https://alion.io/terms","contact":"https://alion.io/contact","api":"https://alion.io/developers","usage":{"tier":"crawler","counted_by":"address","units_charged":1,"used_today":3619,"day_limit":5000,"remaining_today":1381,"minute_limit":60,"resets_at":"2026-10-02T00:00:00Z"}}}