#4188 Hybrid/GraphRAG
#4188 Hybrid/GraphRAG
#4188 Given a knowledge graph, including embedded documents with indexed document chunks, describe the prompt sequence for answering a user's question about the knowledge graph
Answering a user's question using a hybrid Knowledge Graph (KG) and vector database (embedded document chunks) is a powerful approach often referred to as Graph Retrieval-Augmented Generation (GraphRAG). Because you are combining structured data (graph relationships) with unstructured data (text chunks), a multi-step prompt sequence is usually required to accurately extract, retrieve, and synthesize the information.
Here is the standard prompt sequence used in this architecture.
1. Query Analysis and Entity Extraction Prompt
Before searching the database, the system must understand what to look for. The first prompt is sent to the LLM to break down the user's natural language question into actionable search parameters for both the Knowledge Graph and the vector database.
System Role: You are a data extraction assistant. Analyze the user's query and extract key entities, relationships, and semantic search terms.
Input: The user's raw question.
Output Format: JSON containing a list of graph entities (nodes), potential relationships (edges), and a refined search query for the vector database.
Example Prompt:
"Given the user query: '{user_query}', extract the primary entities to match against a Knowledge Graph node index. Also, generate a refined semantic search string to query a vector database of document chunks. Output your response in strictly formatted JSON."
2. Context Translation Prompt (Intermediate/Optional)
Once the system retrieves the relevant graph data (e.g., a subgraph of nodes and edges) and the top-k document chunks, the raw data can be difficult for the final LLM to process if the graph is too large. This intermediate prompt translates raw graph triples (Subject --> Predicate --> Object) into a natural language summary.
System Role: You are a context summarizer.
Input: Retrieved Knowledge Graph triples and structured metadata.
Output Format: A concise natural language paragraph describing the relationships.
Example Prompt:
"Translate the following Knowledge Graph relationships into a clear, concise paragraph: '{retrieved_triples}'. Focus only on facts relevant to the entities mentioned."
3. Synthesis and Final Generation Prompt
This is the core RAG prompt. It combines the original user question, the structured context retrieved from the Knowledge Graph, and the unstructured context from the document chunks. The LLM is instructed to synthesize these two data sources, resolve any conflicts, and generate the final answer.
System Role: You are an expert knowledge assistant. You answer questions based strictly on the provided graph data and document chunks.
Input: The user query, the graph context (raw or translated), and the text chunk context.
Instructions: Synthesize the information. Prioritize facts confirmed by both sources. If the sources conflict, note the discrepancy. Cite your sources based on the document chunk IDs or Graph node IDs.
Example Prompt:
"Answer the user's question using only the provided context.
Graph Context: {graph_context}
Document Context: {document_chunks}
User Question: {user_query}
Instructions: Combine insights from the graph and documents. If the answer is not contained in the context, state 'I do not have enough information to answer this.' Include citations to specific document chunks or graph nodes where applicable."
4. Verification and Formatting Prompt (Self-Correction)
In advanced systems, a final prompt is used to evaluate the generated answer against the original query to ensure no hallucinations occurred and that the formatting meets the user's expectations.
System Role: You are a quality assurance auditor.
Input: The generated answer and the retrieved context.
Output Format: A boolean pass/fail, or a corrected version of the answer.
Example Prompt:
"Review the following answer: '{generated_answer}'. Verify that every claim is supported by the context: '{context}'. If there are unsupported claims, rewrite the answer to remove them."
Would you like me to write out a specific, copy-pasteable JSON prompt template for one of these specific steps?
How does the LLM know the knowledge graph schema?
The LLM does not inherently know the schema of your knowledge graph. Because LLMs are trained on generalized data, they will readily hallucinate node labels, relationship types, or property names if asked to generate graph queries (like Cypher or SPARQL) or extract entities without strict guidance.
To bridge this gap, the schema must be explicitly provided to or discovered by the LLM during the pipeline execution. Here are the best-practice architectural approaches to solving this, ranging from the simplest to the most robust.
1. Static Schema Injection (The Standard Approach)
For small to medium-sized graphs, the most common and straightforward method is to inject the exact schema directly into the system prompt of the Query Analysis or Query Generation step.
How it works: You query your graph database for its schema (e.g., using
CALL apoc.meta.schema()in Neo4j) and format it into a concise string. This string is prepended to the LLM's instructions.Example Prompt: > "You are an expert graph database query writer. Generate a valid query based on the following schema:
Node Properties: User {name: STRING, age: INTEGER}, Document {id: STRING, content: STRING} Relationships: (User)-[:WROTE]->(Document), (Document)-[:CITES]->(Document) User Question: {user_query}"
Pros: Simple to implement, highly accurate for smaller domains.
Cons: Consumes context window tokens and falls apart if the schema is too large or complex.
2. Retrieval-Augmented Schema (Dynamic Injection)
When the schema is too massive to fit efficiently into a context window, you treat the schema itself as a RAG problem.
How it works: You embed the definitions of your nodes, relationships, and properties into a vector database. When a user asks a question, you first run a semantic search against the schema vector store to retrieve only the relevant parts of the ontology. You then inject only that sub-schema into the LLM prompt.
Example: If the user asks about "employee hierarchies," the system retrieves the schema for
Employee,Manager,Department, and theREPORTS_TOrelationship, ignoring the massive schema detailingIT_InfrastructureorCustomer_Orders.Pros: Highly scalable and token-efficient.
Cons: Requires maintaining a secondary index specifically for the ontology/schema.
3. Agentic Introspection (Tool-Calling)
In an agentic framework, instead of giving the LLM the schema upfront, you give the LLM the tools to explore the database itself.
How it works: The LLM is equipped with a tool like
get_graph_schemaorsample_node_properties. When it receives the user's question, it first decides to query the database to discover what labels and relationships exist. Once it understands the landscape, it formulates the final extraction or query generation prompt.Pros: Extremely resilient to schema changes; the LLM handles the logic autonomously.
Cons: Adds latency (multiple round-trips to the LLM) and requires higher-tier reasoning models to reliably execute the tool sequence without getting stuck in a loop.
4. Few-Shot Prompting with Validated Examples
Relying purely on the schema definition often isn't enough, especially for complex analytical queries. The LLM needs to see how the schema is practically applied.
How it works: Alongside (or instead of) the raw schema, you provide a curated list of user questions mapped to the correct, highly-optimized graph extraction JSON or graph queries.
Pros: Drastically reduces syntax errors and hallucinated relationships by providing ground-truth patterns.
Cons: Requires manual curation of the few-shot examples as the application evolves.
For an enterprise-grade semantic search or extraction tool, combining Dynamic Injection (for the ontology) with Few-Shot Prompting (for syntax accuracy) yields the most reliable results.
Would you like to explore how to format a semantic representation of a schema so it embeds cleanly into a vector database?
Consider that the documents and document chunks with indexes and raw text are nodes of the knowledge graph
When your documents, text chunks, and their vector embeddings are stored natively as nodes within the same knowledge graph (e.g., a (Document)-[:HAS_CHUNK]->(Chunk) structure), the architecture becomes significantly more unified. You are no longer managing a split-brain system between a vector database and a graph database.
Because the raw text and embeddings are part of the graph hierarchy, the way the LLM learns the schema and generates the retrieval strategy changes. To get the most accurate and reliable results—rather than relying on the LLM to guess the traversal paths—you must tightly couple the schema injection with the structural logic of the graph.
Here is the best approach for passing this unified schema to the LLM.
1. Abstracting the Schema via a Domain-Specific Language (DSL)
Forcing an LLM to write raw, complex hybrid queries (e.g., a Cypher query that performs a kNN vector search on a chunk, then traverses to connected entity nodes, then fetches adjacent chunks) often leads to syntax errors or hallucinations.
The most robust solution is to abstract your schema into a lightweight Domain-Specific Language (DSL) or a strict JSON schema definition. You provide the LLM with this DSL, and it generates a structured intent, which your application code then compiles into the actual database query.
How it works: You define the valid "paths" the LLM can take in the system prompt.
Prompt Injection:
"You are a query router. Our knowledge graph contains both conceptual entities and raw document chunks.
Available Schema Paths:
VECTOR_SEARCH(Chunk): Finds text chunks semantically similar to the user's query.TRAVERSE(Chunk -> Entity): Finds the concepts mentioned within a specific text chunk.TRAVERSE(Entity -> Chunk): Finds all raw text chunks that mention a specific concept.TRAVERSE(Chunk -> NEXT_CHUNK): Retrieves the subsequent text chunk to provide expanded reading context.
Output your retrieval strategy using only these predefined paths."
2. Injecting the "Graph Shape" (Ontology Patterns)
Instead of just giving the LLM a flat list of node labels (Document, Chunk, Person, Technology) and properties, you must inject the specific patterns (triples) that represent your document model.
If the LLM knows the "shape" of the data, it understands that a Chunk is the bridge between a Document and an Entity.
Prompt Injection:
"Our knowledge graph follows this strict structural ontology:
(Document)-[:CONTAINS]->(Chunk)(Chunk)-[:NEXT]->(Chunk)// Forms a linked list of text for continuous reading(Chunk)-[:MENTIONS]->(Entity)// Links raw text to semantic concepts
The
Chunknode contains the propertyraw_textand a vector indextext_embedding. Ensure any query designed to extract raw information targets theraw_textproperty of theChunknodes connected to the relevantEntity."
3. Contextualizing the Hybrid Retrieval Prompt
Because everything is in one graph, the prompt sequence described previously collapses into a more streamlined process. The LLM's query generation step now dictates a single, coordinated extraction.
The synthesis prompt (the final step where the LLM answers the user) also changes, because the retrieved context is now a unified subgraph.
The Synthesizer Prompt:
"Answer the user's query based on the following retrieved subgraph.
Retrieved Subgraph:
Entity: [Concept Name]
Connected Document: [Document Title]
Relevant Chunks:
Chunk 14: "..."
Chunk 15: "..." (via NEXT_CHUNK relationship)
Synthesize your answer relying primarily on the raw text chunks, using the Entity relationships only to establish context."
Why this is the best solution:
By treating the chunks as nodes and abstracting the traversal paths into a DSL, you completely remove the LLM's burden of knowing the underlying database query language syntax. It only needs to understand the logic of the schema (e.g., "I need to find a concept, then pull the chunks connected to it, then read the next chunk for context").
Would you like me to outline how the parsing logic in your application code would translate this LLM-generated DSL into an executable hybrid graph query?
Yes
Translating an LLM-generated Domain-Specific Language (DSL) into an executable hybrid graph query requires building a lightweight query compiler in your application layer. Since you are mapping a conceptual retrieval strategy to a concrete database syntax (like Cypher for Neo4j), treating the LLM's output as an Abstract Syntax Tree (AST) or a sequence of pipeline operations is the most robust software engineering approach.
Here is the architectural breakdown of how your application code parses the DSL and compiles the query.
1. The LLM Output (The DSL Payload)
First, enforce a strict JSON schema for the LLM to output. Instead of writing queries, the LLM outputs a sequential retrieval pipeline.
Example LLM Output:
{
"pipeline": [
{
"step": 1,
"operation": "VECTOR_SEARCH",
"target": "Chunk",
"query_string": "quantum error correction",
"top_k": 3
},
{
"step": 2,
"operation": "TRAVERSE_TO_CONCEPT",
"relationship": "MENTIONS",
"target_node": "Entity"
},
{
"step": 3,
"operation": "EXPAND_CONTEXT",
"relationship": "NEXT",
"target_node": "Chunk",
"depth": 1
}
]
}
2. The Application Parser (The Compiler)
Your application code (e.g., Python) reads this JSON array and uses a mapping function or the Strategy pattern to translate each operation into a fragment of your target database language.
For a graph database like Neo4j, the compiler strings together Cypher clauses, passing the variables down the chain using WITH.
class DSLCompiler:
def __init__(self):
self.query_parts = []
self.parameters = {}
def compile(self, dsl_pipeline: list) -> tuple[str, dict]:
for index, step in enumerate(dsl_pipeline):
operation = step.get("operation")
if operation == "VECTOR_SEARCH":
self._compile_vector_search(step, index)
elif operation == "TRAVERSE_TO_CONCEPT":
self._compile_concept_traversal(step)
elif operation == "EXPAND_CONTEXT":
self._compile_context_expansion(step)
else:
raise ValueError(f"Unknown DSL operation: {operation}")
# Final return statement to aggregate the synthesized subgraph
self.query_parts.append(
"RETURN chunk.text AS primary_text, "
"collect(DISTINCT entity.name) AS connected_entities, "
"collect(DISTINCT next_chunk.text) AS surrounding_context"
)
return "\n".join(self.query_parts), self.parameters
def _compile_vector_search(self, step, index):
# Maps to a semantic index call. Assume we embed the query_string externally first.
self.parameters["query_embedding"] = get_embedding(step["query_string"])
self.parameters["top_k"] = step.get("top_k", 5)
cypher = (
f"CALL db.index.vector.queryNodes('chunk_embeddings', $top_k, $query_embedding) "
f"YIELD node AS chunk, score\n"
f"WITH chunk, score"
)
self.query_parts.append(cypher)
def _compile_concept_traversal(self, step):
# Traverses from the active 'chunk' variable to connected entities
rel = step.get("relationship", "MENTIONS")
target = step.get("target_node", "Entity")
cypher = (
f"OPTIONAL MATCH (chunk)-[:{rel}]->(entity:{target})\n"
f"WITH chunk, entity"
)
self.query_parts.append(cypher)
def _compile_context_expansion(self, step):
# Follows the linked list of chunks to fetch the next sequential chunk for reading context
rel = step.get("relationship", "NEXT")
cypher = (
f"OPTIONAL MATCH (chunk)-[:{rel}]->(next_chunk:Chunk)\n"
f"WITH chunk, entity, next_chunk"
)
self.query_parts.append(cypher)
3. The Compiled Executable Query
When the compile() method runs against the LLM's JSON payload, it generates a perfectly syntactically correct, highly-optimized hybrid query that looks like this:
CALL db.index.vector.queryNodes('chunk_embeddings', $top_k, $query_embedding) YIELD node AS chunk, score
WITH chunk, score
OPTIONAL MATCH (chunk)-[:MENTIONS]->(entity:Entity)
WITH chunk, entity
OPTIONAL MATCH (chunk)-[:NEXT]->(next_chunk:Chunk)
WITH chunk, entity, next_chunk
RETURN chunk.text AS primary_text, collect(DISTINCT entity.name) AS connected_entities, collect(DISTINCT next_chunk.text) AS surrounding_context
4. Execution and Subgraph Assembly
Finally, your application executes the query using the database driver. The resulting payload is a clean, structured dictionary representing a targeted subgraph.
Because you abstracted the schema into the DSL and handled the query logic in the application layer, the LLM never had to worry about Cypher syntax errors, OPTIONAL MATCH chaining, or variable scoping. You just inject this structured return data directly into the final Synthesizer Prompt.
Would you like to look at how to implement an exact-match regex fallback within this compiler for instances where the LLM's semantic search criteria are too fuzzy for precise entity extraction?
From <https://gemini.google.com/app/> Google Gemini 3.1 Pro
Comments
Post a Comment