Cloudera
Cloudera
Cloudera is an enterprise data cloud company that helps organizations manage, govern, and analyze data across hybrid environments. Its platform supports modern data architectures, including open data lakehouses, data meshes, and unified data fabrics, with capabilities for artificial intelligence, machine learning, and data engineering. Cloudera serves sectors such as financial services, telecommunications, healthcare, and logistics, enabling teams to work with data at scale and turn complex information into operational insight.

GraphRAG Engineer - Cloudera Spain Remote

Build Cloudera’s enterprise AI knowledge graph and GraphRAG capabilities across Neo4j, vector storage, and automated ingestion pipelines. Operate supporting AWS and GCP infrastructure for developer workflows and enterprise AI services.

Description

  • Provision, optimize, and support production Neo4j graph databases and pgvector storage clusters
  • Design high-throughput indexes, cosine-similarity search, and query improvements for sub-second performance
  • Develop ingestion pipelines that transform Git repositories, ASTs, Jira links, Avro schemas, and CI/CD metadata into an enterprise knowledge graph
  • Connect distributed pipeline engines to hybrid retrieval systems using SQL, Cypher traversals, and dense embeddings
  • Set circuit breakers, confidence thresholds, and execution limits for autonomous agents
  • Integrate microservices and knowledge repositories with the Enterprise AI Gateway
  • Manage version-controlled prompt structures in localized .ai/ directories while applying DLP, PII-scrubbing, and token-rate controls
  • Automate failover, backup recovery, and multi-cloud storage cost management across AWS and GCP
  • Own the semantic, vector, and graph storage layer supporting the context engine, enterprise AI utilities, and Internal Developer Portal
  • Lead delivery of the SDLC Context Graph and GraphRAG Engine for CAB compliance, code and schema lineage, and enterprise LLM proxy integrations

Requirements

  • Extensive hands-on experience developing and operating Neo4j, including Cypher, APOC, and causal clustering, or comparable enterprise knowledge graph systems
  • Demonstrated expertise with pgvector and PostgreSQL, embedding management, hybrid search, and LangChain, LlamaIndex, or custom RAG pipeline integrations
  • Practical experience managing relational and graph databases in AWS and GCP environments
  • Ability to consume Apache Avro payloads, process Kafka events through AWS MSK, and parse structured or unstructured code and JSON artifacts
  • Working knowledge of Prompts-as-Code, few-shot prompt refinement, and agent tool specifications
  • Experience with Terraform infrastructure primitives, Kubernetes on EKS or GKE, Docker, and pull-based GitOps processes
  • Familiarity with HashiCorp Vault Transit encryption, OIDC keyless authentication, and zero-trust workload identities
  • Understanding of OpenTelemetry instrumentation for measuring vector-search latency and LLM inference performance in Datadog or Grafana

Benefits

  • Generous paid time off
  • Unplugged days designed to support work-life balance
  • Flexible work-from-home policy
  • Mental and physical wellness programs
  • Phone and internet reimbursement
  • Ongoing career development opportunities
  • Comprehensive benefits with competitive packages
  • Paid time for volunteering
  • Employee resource groups

Related Jobs

StreSERT Integrated Limited | Consulting | BPO | Talent Acquisition | SILution (L&D)

Digital Marketing and IT Executive (Hybrid, Nigeria)

StreSERT Integrated Limited | Consulting | BPO | Talent Acquisition | SILution (L&D)

Lead digital marketing and IT activities for a Nigerian consulting group, covering websites, social media, content, SEO, and paid campaigns. Help grow brand visibility, business enquiries, executive profiles, and measurable digital performance.

Open