ZenteiQ. AI Platform Engineering Backend Engineer Mid-Level / Senior Full-time Bengaluru (Onsite) You are an engineer who gets uncomfortable when a service does something it cannot explain, when there is no trace, no metric, no clear reason a job took four times longer than expected. You build backends that are boring in the best sense: predictable under load, observable when something goes wrong, and straightforward for the next engineer to operate. You do not reach for complexity until simplicity has genuinely failed.
The Problem You Will Own: ZenteiQ is building AI infrastructure for scientific and domain-expert users, a platform where models get registered, evaluated, promoted, and served under real production conditions. Right now, the core APIs, data layer, and async job infrastructure need an engineer who can take them from working to production-grade: reliable schema design, event-driven workflows that handle failure gracefully, and services that remain operable as load increases. You will own the backend platform layer end-to-end from the database schema to the Kafka consumer to the Kubernetes deployment, and be the person who sets the quality bar for how the rest of the team writes and ships services.
Responsibilities:
- Design and build the REST APIs that power the model registry, evaluation execution, and promotion workflows with clean OpenAPI specs, explicit versioning, and error handling that gives callers enough information to recover.
- Own the PostgreSQL data layer: schema design with proper indexing and constraints, migration strategy, and query optimisation that holds up as data volume grows.
- Build the async job infrastructure: Celery workers on Redis, Airflow DAGs for orchestration, and make them sufficiently observable that on-call is not a guessing game.
- Design and maintain the event-driven components (Kafka / Pub/Sub) with the reliability properties that distributed systems demand: idempotency, backpressure handling, deadletter queues, and retry logic.
- Ensure every service is containerised, deployable on Kubernetes (GKE), and ships with the logging, metrics, and tracing that lets you debug a production incident in under thirty minutes.
- Write architecture decision records for non-obvious choices and API documentation that a new team member can follow without asking you to walk them through it.
Requirements:
- You have shipped Python backend services to production and can point to specific decisions: schema design, caching strategy, async pattern that you made and stand behind.
- You understand PostgreSQL well enough to look at a slow query, explain why it is slow, and fix it, not just add an index and hope.
- You have built or maintained an event-driven system (Kafka, Pub/Sub, or equivalent) and can articulate what happens when a consumer falls behind or a message is processed twice.
- You understand the difference between synchronous and asynchronous failure modes and know when to use Celery, when to use Airflow, and when neither is the right answer.
- You treat observability as part of the feature, not a follow-up ticket; your services ship with structured logs, metrics, and health endpoints from day one.
Good to have:
- Production experience with Go (goroutines, channels, context propagation); we use it for performance-sensitive microservices and would like the overlap.
- Hands-on experience with Airflow or a comparable orchestration framework for multistep ML or data workflows.
- Familiarity with distributed tracing (Jaeger, OpenTelemetry) and a full observability stack (Prometheus, Grafana, Loki).
- Experience in a compliance-conscious environment (SOC2 HIPAA, GDPR); awareness of what needs to be auditable and why.
- This is not a role for someone who considers deployment someone else's problem.
- If you draw a hard line between writing code and running it, the scope here will be uncomfortable.
- Technology Area Required Preferred Languages: Python (FastAPI, async/await, Pydantic) Go (goroutines, channels) Databases PostgreSQL (schema, indexing, migrations) Firestore, MongoDB Area Required Preferred Caching / Queues Redis (caching, rate limiting, Celery) Messaging Kafka or Google Pub/Sub Both Orchestration Docker, Kubernetes (GKE) Airflow, Prefect Cloud GCP (Cloud SQL, GCS, GKE) Multi-cloud patterns Observability Prometheus, structured logging Grafana, Loki, Jaeger / OpenTelemetry API Design OpenAPI/Swagger, versioning, auth patterns GraphQL Quality Unit + integration tests, ADRs, code review SOC2 / HIPAA awareness Why ZenteiQ Selected by the IndiaAI Mission (MeitY) to build India's sovereign Scientific Foundation Model backend infrastructure you build will underpin nationally significant AI research.

