ResearchPod Summary
Standard Retrieval-Augmented Generation (RAG) systems are designed to answer isolated questions by retrieving documents based on semantic similarity. However, enterprise users typically interact with systems in sessions—coherent episodes where they need information from multiple, often semantically distant, parts of a knowledge base. Because standard retrieval only surfaces documents that look like the current query, it often fails to retrieve the full set of documents required to resolve a user's entire session, forcing the user to issue multiple follow-up queries.
The authors propose a pre-retrieval solution: reorganizing the knowledge base (KB) offline to reflect how documents are actually used together. By training a Word2Vec model on document co-occurrence sequences—derived from expert-labeled queries, synthetic QA pairs, and random walks through the KB—the system learns a co-occurrence embedding space. Documents that are functionally related are clustered together in this space, even if they are semantically distant. At query time, the system performs standard retrieval to find the most relevant documents and then expands the candidate pool to include all documents from the clusters of those initial results.
Testing on the WixQA benchmark, the authors demonstrate that this cluster-expanded hybrid retrieval raises single-query session coverage from 41% to 58%. This improvement is consistent across four different embedding models and six functional domains, confirming that the learned cluster structure provides a signal orthogonal to standard semantic embeddings. Furthermore, the method reduces the number of retrieval calls required to reach a 70% coverage threshold by 34%, directly lowering latency and API costs for enterprise applications. As a bonus, the clustering process compresses the effective knowledge base to 20% of its original size.
AI-generated third-party summary by ResearchPod. Not official content or an endorsement by the paper authors or affiliated organizations.