368,611open jobs
9,439companies
50,719added this week
Browse all
Salary
$28k – $69k per year (Estimated)
Location
Remote (India)
Seniority
Senior · 5+ years exp
Employment
Full-Time
Overview
Company
Impact
Profile match
Jobgether is an AI-powered job platform focused on remote and flexible work. It matches candidates with relevant roles using skills and preference-based algorithms, and also offers career coaching and job-search guidance.

This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Senior Software Engineer - Infra Agent Systems based in India.

This is a high-impact engineering role focused on building production AI agents that operate and automate large-scale GPU infrastructure. You’ll design systems that diagnose hardware failures, investigate incidents, gather evidence, and support remediation across complex infrastructure environments. The role spans AI agent systems, distributed services, knowledge graphs, retrieval, orchestration, and developer tooling. You’ll own systems end to end, from architecture and implementation through deployment, observability, and production operations. You’ll collaborate across infrastructure, datacenter, and engineering teams to turn operational knowledge into reliable automation. This is an opportunity to help shape how autonomous AI systems can safely and intelligently operate real-world infrastructure at massive scale.

Accountabilities:

    • Design and build production AI agent systems capable of diagnosing, investigating, and supporting remediation of infrastructure issues across large-scale GPU environments.
    • Develop the distributed services, orchestration frameworks, knowledge graphs, retrieval systems, and supporting infrastructure that power AI agents.
    • Build fleet intelligence capabilities that combine telemetry, infrastructure state, operational knowledge, and historical incidents to improve agent decision-making.
    • Integrate agent systems with observability, incident management, ticketing, fleet inventory, source control, communication platforms, and internal infrastructure through reliable APIs.
    • Own services throughout their lifecycle, including architecture, implementation, testing, deployment, monitoring, reliability, and production support.
    • Improve agent quality and reliability through evaluations, retrieval optimization, better tools, and continuous feedback from production environments.
    • Convert insights and knowledge generated through production use into reliable, reviewed software, workflows, and automation.
    • Contribute to the evolution of platform architecture and engineering practices as autonomous infrastructure capabilities scale.
    • Requirements

      • Bachelor’s degree or equivalent professional experience in Computer Science, Engineering, or a related technical field.
      • 5+ years of professional experience building production backend systems, distributed systems, infrastructure platforms, or similarly complex software.
      • Strong systems design capabilities and demonstrated experience taking significant systems from initial architecture through production.
      • Deep expertise in at least one relevant area, such as AI agent systems, orchestration, tool use, evaluation, grounding, knowledge graphs, graph data modeling, search, retrieval, ranking, RAG, or semantic search.
      • Strong backend engineering skills, including API design, service boundaries, data modeling, and integrations across complex technical environments.
      • Experience with Kubernetes, GitOps practices such as ArgoCD, infrastructure-as-code, and cloud platforms.
      • Proficiency in one or more relevant programming languages, such as Go, TypeScript, Python, or Rust, with the ability to work across multiple languages when required.
      • Strong analytical and problem-solving abilities, with an interest in solving ambiguous and technically challenging infrastructure problems.
      • Ability to own systems in production, balancing engineering quality, reliability, operational requirements, and delivery speed.
      • Experience with GPU infrastructure, datacenters, bare-metal environments, hardware failure modes, BMC/IPMI, or cluster schedulers is a plus.
      • Experience with graph databases, event-driven systems, messaging platforms such as NATS or Kafka, or observability tools such as Prometheus and Grafana is advantageous.
      • Experience building evaluation frameworks or improving the reliability and quality of LLM-powered systems is also a plus.
      • Comfortable working in a remote, highly collaborative engineering environment with strong ownership and autonomy.
      • Benefits

        • Fully remote position based in India.
        • Opportunity to work on AI agents operating at significant infrastructure scale.
        • Exposure to the intersection of artificial intelligence, distributed systems, GPU infrastructure, knowledge systems, and automation.
        • Opportunity to build foundational systems and influence the architecture of emerging autonomous infrastructure technologies.
        • End-to-end ownership of production software, from design and development through deployment and operations.
        • Collaborative environment with technically ambitious engineers and researchers working on challenging AI infrastructure problems.
        • Significant opportunities for learning, experimentation, and professional growth in a rapidly evolving technology space.
        • Opportunity to contribute to systems that have direct, measurable production impact.
Free account
Stop reading job ads. Get the ones that fit.
One free account turns this page into a shortlist built around your stack, your level and your pay.
Match on every job. Stack, seniority, pay and location, scored against your profile.
368,611 open roles. Read straight off company career pages, refreshed every day.
Unlimited applications. Every one you send is tracked in one place, on-site or on a company board.
3 tailored CVs a month. Rewritten for the exact job you are applying to. Included free.
Create a free account
Free forever. No card. Under a minute.

Your match

How well do you fit this role?
Two answers are enough for a real match. No account needed.
Check my fit
Answers stay in this browser until you create an account.

Recommended for you based on this role

Similar stack
Same company
In your city
$109k – $118k per year • In office • Contractor • 5+ years exp • Plano
Bash
JavaScript
Node JS
Python
Python
FastAPI
Flask
Gunicorn
Uvicorn
Databases
Meilisearch
pgvector
PostgreSQL
Redis
AI/ML
AI Agents
AWS Bedrock
Claude
LLM
RAG
Model Context Protocol
DevOps
AWS
CI/CD
Docker
Git
Grafana
Incident Management
Jenkins
Kubernetes
Nginx
Prometheus
Rest API
Ubuntu
Amazon CloudWatch
Apply
$36k – $81k per year (Estimated) • In office • Gurgaon
AI/ML
ElevenLabs
LangChain
LangGraph
LLM
Prompt Engineering
RAG
AI Agents
DevOps
AWS
Azure
GCP
Apply
$173k – $250k per year • In office • Full-Time • San Francisco • New York • Seattle • Austin
TypeScript
Databases
PostgreSQL
AI/ML
AI Agents
Claude
Claude Code
Cursor
LLM
LLM Guardrails
Model Context Protocol
Apply
$224k – $279k per year • In office • Full-Time • 3+ years exp • Bachelor's Degree • San Francisco • New York • Seattle • Austin
Python
JavaScript
Databases
PostgreSQL
Redis
AI/ML
Time Series Forecasting
Frontend
Bootstrap
DevOps
Ansible
CI/CD
Docker
Grafana
Incident Management
OpenTelemetry
Platform Engineering
Prometheus
Terraform
Robotics
Digital Twin
Apply
$31k – $61k per year (Estimated) • In office • Full-Time • 5+ years exp • Bengaluru
Python
SQL
Databases
Presto
AI/ML
AI Agents
Fine-tuning
Kubeflow
LangChain
LangGraph
LangSmith
LLM
LoRA
MLFlow
PEFT
Prompt Engineering
Spark
Transformers
Amazon SageMaker
LLM Guardrails
DevOps
Git
GitHub
Apply
$126k – $201k per year • Equity • Remote • Full-Time • 5+ years exp • Bachelor's Degree
Analytics
A/B Testing
Apply
$84k – $166k per year (Estimated) • Remote • Full-Time • 7+ years exp • Bachelor's Degree
SQL
Apply
$80k – $190k per year • Remote • Full-Time • 2+ years exp
Apply
$134k – $223k per year (Estimated) • Remote • Full-Time • 5+ years exp • Bachelor's Degree
Bash
Python
AI/ML
Claude
Claude Code
Copilot
OpenAI Codex
DevOps
Azure
Azure DevOps
CI/CD
Gerrit
Git
Jenkins
KVM
QEMU
RTOS
VMWare
Xen
Cybersecurity
Tcpdump
Wireshark
IoT
FreeRTOS
Management
Confluence
Jira
Apply
$165k – $301k per year (Estimated) • Equity • Remote • Full-Time • 12+ years exp
AI/ML
AI Agents
Apply
See all jobs
This is one of many
368,611 more open roles from verified company boards, updated every day.