{"id":1527863,"url":"https://alion.io/job/cloudera-graphrag-engineer-2","title":"GraphRAG Engineer","company":{"id":58503,"name":"Cloudera","domain":"cloudera.com","url":"https://alion.io/company/cloudera","size_band":"1001-5000","is_staffing_agency":false,"employer_type":"direct","is_intermediary":false,"listed_via":null,"ats_vendor":"Workday","truth_index":{"grade":"A","score":95,"open_postings":12,"ghost_share":0,"stale_share":0.417,"repost_share":0,"time_to_fill_p50_days":21,"computed_at":"2026-10-01T05:45:00Z"}},"role":"Industrial Engineering","role_family":"Industrial Engineering","seniority":null,"employment_type":"full_time","work_mode":"remote","remote_scope":"stated_countries","remote_scope_basis":"board_field","remote_working_hours":null,"hiring_geo_confidence":"structured","locations":["Barcelona, Spain","Madrid, Spain","Hungary","Poland","Spain","Czech Republic"],"countries":["ES","HU","PL","CZ"],"hiring_countries":["HU"],"hiring_countries_total":1,"salary":null,"salary_estimate":{"min_usd":23000,"max_usd":67000,"period":"year","method":null,"sample_n":9274},"experience_years_min":null,"visa_sponsorship":false,"relocation_package":false,"has_equity":false,"technologies":[{"name":"AI Agents","optional":false},{"name":"Amazon EKS","optional":false},{"name":"Apache Kafka","optional":false},{"name":"AWS","optional":false},{"name":"CI/CD","optional":false},{"name":"Datadog","optional":false},{"name":"DLP","optional":false},{"name":"Docker","optional":false},{"name":"Embeddings","optional":false},{"name":"FinOps","optional":false},{"name":"GCP","optional":false},{"name":"Git","optional":false},{"name":"GitOps","optional":false},{"name":"Google GKE","optional":false},{"name":"Grafana","optional":false},{"name":"GraphRAG","optional":false},{"name":"HashiCorp Vault","optional":false},{"name":"Hybrid Search","optional":false},{"name":"Jira","optional":false},{"name":"Knowledge Graph","optional":false},{"name":"Kubernetes","optional":false},{"name":"LangChain","optional":false},{"name":"LlamaIndex","optional":false},{"name":"LLM","optional":false},{"name":"Neo4j","optional":false},{"name":"OpenTelemetry","optional":false},{"name":"pgvector","optional":false},{"name":"PostgreSQL","optional":false},{"name":"Prompt Engineering","optional":false},{"name":"RAG","optional":false},{"name":"SQL","optional":false},{"name":"Terraform","optional":false},{"name":"Zero Trust","optional":false}],"status":"live","first_seen_at":"2026-09-29T00:00:00Z","employer_posted_date":"2026-09-29","last_verified_at":"2026-10-02T00:39:30Z","board_verified":true,"closed_at":null,"days_open":3,"trust":{"level":"ok","repost_count":null,"flags":[],"days_open":3},"description":"Business Area:\nITSeniority Level:\nMid-Senior levelJob Description:\nAt Cloudera, we empower people to transform complex data into clear and actionable insights. With as much data under management as the hyperscalers, we're the preferred data partner for the top companies in almost every industry. Powered by the relentless innovation of the open source community, Cloudera advances digital transformation for the world’s largest enterprises.\nAbout the Team & Role\nWe are engineering an enterprise-grade Everything-as-Code (EaC) AI-First Platform that transforms modern enterprise operations through automated delivery pipelines, machine-readable specifications, and agentic intelligence. As a GraphRAG Engineer, you will own the semantic, vector, and graph storage layer powering the core context engine for our enterprise AI utilities and Internal Developer Portal.\nOperating at the intersection of modern database administration, distributed event streaming, and generative AI pipelining, you will bridge our AWS MSK event mesh with downstream knowledge graphs and vector engines across AWS and GCP. You will lead the deployment of our SDLC Context Graph and GraphRAG Engine, enabling automated Change Advisory Board (CAB) compliance, semantic code/schema lineage tracking, and enterprise LLM proxy integrations.\nAs a GraphRAG Engineer you will:\n Graph & Vector Database Infrastructure: Provision, tune, and maintain production-grade clusters for Graph databases (Neo4j using Cypher, APOC, and causal clustering) and Vector storage engines (pgvector on PostgreSQL / AWS/GCP managed storage). Engineer high-throughput index structures, cosine similarity vector indexes, and query optimizations for sub-second responses.\nSDLC Context Graph & Lineage Pipelines: Build automated ingestion pipelines to parse Git repositories, ASTs, Jira issue links, Apache Avro schemas, and CI/CD metadata into a unified enterprise knowledge graph. \n GraphRAG Orchestration & Agentic Search: Connect distributed pipeline engines to hydrate hybrid retrievers (combining structured SQL, Cypher graph traversals, and dense vector embeddings) for AI-driven developer workflows and autonomous coding agents.\nCyclic Agent Safeguards & Governance: Configure circuit breakers, confidence scoring thresholds, and step-limit constraints to restrict autonomous cyclic agent execution, protect token budgets, and prevent runaway execution loops.\nPrompts-as-Code & Enterprise LLM Gateway Integration: Integrate microservices and knowledge stores with the central Enterprise AI Gateway, maintaining version-controlled system prompt structures inside localized .ai/spoke directories while adhering to DLP PII scrubbing rules and token rate limits.\nHigh Availability & FinOps: Implement automated failover, backup restoration, and multi-cloud storage tier cost controls across AWS and GCP environments.\n We are excited if you have (Required Technical Expertise):\nGraph Databases: Deep operational and development experience with Neo4j (Cypher, APOC, causal clustering) or enterprise Knowledge Graphs.\nVector Search & RAG: Proven expertise with pgvector (PostgreSQL), embeddings management, hybrid search techniques, and framework integrations (LangChain, LlamaIndex, or custom RAG pipelines).\n Database Administration & Cloud Storage: Hands-on experience managing relational (PostgreSQL) and graph databases across AWS and GCP cloud environments.\n Data Pipelining & Streaming: Proficiency in consuming Apache Avro payloads, streaming Kafka events (AWS MSK), and parsing structured/unstructured code and JSON artifacts.\nAgentic AI & Prompt Engineering: Practical understanding of Prompts-as-Code patterns, few-shot prompt optimization, and agent tool specification.\n You may also have:\nExperience with Infrastructure-as-Code (Terraform) primitives, Kubernetes (EKS/GKE), Docker, and pull-based GitOps workflows.\nExposure to HashiCorp Vault Transit encryption, OIDC keyless authentication, and zero-trust workload identities.\nFamiliarity with OpenTelemetry (OTel) instrumentation for tracking vector search query latencies and LLM inference performance in Datadog or Grafana.\nWhat you can expect from us:\nGenerous PTO Policy\n\nSupport work life balance with Unplugged Days\n\nFlexible WFH Policy\n\nMental & Physical Wellness programs\n\nPhone and Internet Reimbursement program\n\nAccess to Continued Career Development\n\nComprehensive Benefits and Competitive Packages\n\nPaid Volunteer Time\n\nEmployee Resource Groups\n\nEEO/VEVRAA","description_format":"text","description_chars":4483,"description_truncated":false,"requirements":{"experience_years_min":null,"management_years_min":null,"team_size_min":null,"manages_managers":false,"education":null,"security_clearance":false,"languages":[]},"benefits":["Internet reimbursement","Wellness"],"hiring_locations":[{"name":"Hungary","iso":"HU","kind":"country"}],"hiring_excludes":[],"relocation_offered":false,"industries":["Artificial Intelligence","Data & Analytics","Professional Services","MLOps"],"lifecycle":[{"event":"open","at":"2026-09-30T14:41:34Z"}],"liveness":{"score":90,"band":"hot","label":"Hiring now","p_open":1,"p_active":0.903,"p_room":1,"age_days":2,"expected_fill_days":21,"reasons":["conf:3","velocity","win:early"],"computed_at":"2026-10-01T05:45:00Z"},"pay":null,"html_url":"https://alion.io/job/cloudera-graphrag-engineer-2","json_url":"https://alion.io/job/cloudera-graphrag-engineer-2.json","meta":{"generated_at":"2026-10-02T00:41:56Z","cache_seconds":300,"methodology":"https://alion.io/methodology","terms":"https://alion.io/terms","contact":"https://alion.io/contact","api":"https://alion.io/developers","usage":{"tier":"crawler","counted_by":"address","units_charged":1,"used_today":661,"day_limit":5000,"remaining_today":4339,"minute_limit":60,"resets_at":"2026-10-03T00:00:00Z"}}}