Self-Knowledge Retrieval Augmented Generation Framework for Patent Matching
Jian Zhang, Songlin Lei, Zhuohao Yang, Bangli Liu, Ziwei Wang, Xufeng Weng, Gehan Amaratunga, Yu Lin, Hongwei Wang
Zhejiang University · ZJU-UIUC Institute · Shaoxing K3i Technology Co. Ltd · State Key Laboratory of CAD&CG
cs.IR, cs.CL
Submitted: 2026-08-11
Updated: 2026-08-12
Comments: Accepted by IEEE CSCWD 2026
DOI: 10.1109/CSCWD68734.2026.11582454
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 75/100
The gist: This paper proposes a self-knowledge retrieval-augmented generation (RAG) framework for patent matching.
Terminology
Summary
This paper proposes a self-knowledge retrieval-augmented generation (RAG) framework for patent matching. The authors note that patent documents have complex structures, dense technical terminology, and multi-modal information, which makes it difficult for traditional methods to accurately identify subtle differences between patents. Existing LLM-based patent matching approaches typically rely on domain-specific pretraining or instruction tuning, which entail high manual labeling costs and catastrophic forgetting. While RAG methods introduce external knowledge, they fail to fully leverage the LLM's capability to automatically parse patents and mine deep semantic relationships.
To address these limitations, the proposed framework guides LLMs to autonomously extract key technical entities and construct hierarchical ontological structures from patent matching queries, thereby enabling query expansion and precise retrieval. The method integrates FAISS retrieval with a generative matching mechanism, leveraging self-knowledge to enhance the model's understanding of patent innovations and significantly improve retrieval and matching accuracy.
The framework operates in a three-stage paradigm of Preparation-Retrieval-Reasoning.
In the first stage, self-knowledge mining, the LLM extracts entities and ontologies from the query patent abstract. Entities represent concrete technical elements extracted as keywords, while ontological information is derived from the International Patent Classification (IPC) system, organizing abstract categories, attributes, and relationships in a tree structure. In the second stage, knowledge-guided patent retrieval, the entity information is concatenated to form an expanded query string combined with the original abstract text, which is then encoded using a pre-trained language model (BGE model) and searched via the FAISS framework using cosine similarity to identify the top-K candidate patents. In the third stage, context-aware generation, all available information—including the original abstract, self-mined entity and ontology information, and retrieved similar patents—is integrated into an instruction set that enables the LLM to perform deep reasoning and generate the final matching result.
The method was validated on the PatentMatch dataset, which contains 1,000 patent matching instances (500 Chinese and 500 English) covering 8 IPC categories. The authors compared against several baselines including vanilla LLMs (Qwen2-Instruct-7B, GLM-4-Chat-9B, Qwen2.5-Instruct-14B), domain-specific LLMs (MoZi-7B, PatentGPT-1.5B, PatentGPT-1.0-Dense-70B), Chain-of-Thought (CoT) reasoning, and conventional RAG. Experimental results show that the proposed method consistently outperforms all baselines. For example, with GLM-4-Chat-9B, the method achieved 83.6 on the English dataset and 79.0 on the Chinese dataset, resulting in an overall accuracy of 81.3, significantly surpassing all other backbone models. With Qwen2.5-Instruct-14B, the method achieved an overall accuracy of 80.7.
A case study demonstrates that the vanilla LLM relies solely on limited information from the query patent, leading to misinterpretations and incorrect matching results, while the proposed method leverages internally mined knowledge to enhance the model's comprehension and guide it to correctly identify the most accurately matched patent.
The main contributions of the paper are: (1) designing a patent self-knowledge mining strategy based on LLMs that leverages the model's capability to automatically extract entity and ontology information from patents; (2) introducing a self-knowledge guided RAG framework that enhances FAISS vector index retrieval and generative matching through self-mined knowledge and query expansion, surpassing traditional methods such as CoT and conventional RAG; (3) validating the effectiveness of the proposed method on real-world patent datasets, with case studies demonstrating the advantages of the approach.
Future work will focus on dynamic ontology generation and integrating multimodal patent data (e.g., images and text) to further improve the model's adaptability to real-world patent retrieval and matching scenarios.
Improvements for AI systems
Improvements to AI Systems:
-
Autonomous Domain Knowledge Extraction: Implement a self-knowledge mining module that enables LLMs to automatically extract technical entities and build hierarchical ontological structures (e.g., using IPC taxonomies) from input queries, eliminating the need for manual labeling or domain-specific fine-tuning. This reduces catastrophic forgetting and lowers deployment costs.
-
Query Expansion via Structured Semantics: Enhance retrieval pipelines by concatenating extracted entity keywords and ontological relationships with the original query text before embedding. This improves vector search precision (e.g., via FAISS with cosine similarity) by capturing subtle technical distinctions that raw text alone misses.
-
Three-Stage
Preparation-Retrieval-Reasoning
Pipeline: Integrate a staged reasoning framework where (a) the LLM parses and structures input, (b) a hybrid retriever (dense vector + expanded query) fetches top-K candidates, and (c) the LLM performs context-aware generation using all retrieved and self-mined information. This improves final matching accuracy by combining internal knowledge with external evidence. -
Cross-Lingual and Multi-Category Robustness: Apply the framework to multilingual patent datasets (e.g., Chinese and English) and diverse IPC categories, enabling AI systems to generalize across languages and technical domains without retraining, as demonstrated by consistent accuracy gains (e.g., 81.3 overall with GLM-4-Chat-9B vs. baselines).
-
Dynamic Ontology Adaptation: Extend the system to support dynamic ontology generation—updating entity hierarchies in real-time as new patents or technical terms emerge—improving adaptability to evolving patent landscapes.
-
Multimodal Patent Understanding: Integrate multimodal data (e.g., patent figures, diagrams, and text) into the self-knowledge extraction and retrieval stages, allowing the AI to reason over visual and textual cues jointly, which is critical for patents with complex schematics.
What the Improved AI System Can Do:
-
Automatically analyze a patent query, extract key technical entities and their hierarchical relationships, and use them to retrieve highly relevant patents with higher precision than traditional RAG or CoT methods.
-
Operate across multiple languages and technical fields without manual annotation, reducing human effort and avoiding model forgetting.
-
Provide explainable matching results by grounding decisions in both self-mined ontological knowledge and retrieved similar patents, improving trust and interpretability.
-
Adapt to new patent domains and multimodal inputs (e.g., images + text) in real-world scenarios, enabling robust patent search, novelty assessment, and prior-art analysis with minimal reconfiguration.
Related papers
- The Price of Isolation: Estimating the Ecosystem Cost of Symmetric Two-Sided A/B Testing
- SCAR: Semantic Continuity-Aware Retrieval for Efficient Context Expansion in RAG
- MixLoRA-DSI: Dynamically Expandable Mixture-of-LoRA Experts for Rehearsal-Free Generative Retrieval over Dynamic Corpora
- RRCM: Ranking-Driven Retrieval over Collaborative and Meta Memories for LLM Recommendation
- Right Family, Wrong Skill: Evaluating Risk Exposure in Agent Skill Retrieval
- UltRAG: a Universal Simple Scalable Recipe for Knowledge Graph RAG