[2502.09956] KGGen: Extracting Knowledge Graphs from Plain Text with Language Models
Abstract:Recent interest in building foundation models for KGs has highlighted a fundamental challenge: knowledge-graph data is relatively scarce. The best-known KGs are primarily human-labeled, created by pattern-matching, or extracted using early NLP techniques. While human-generated KGs are in short supply, automatically extracted KGs are of questionable quality. We present a solution to this data scarcity problem in the form of a text-to-KG generator (KGGen), a package that uses language models to create high-quality graphs from plaintext. Unlike other KG extractors, KGGen clusters related entities to reduce sparsity in extracted KGs. KGGen is available as a Python library (\texttt{pip install kg-gen}), making it accessible to everyone. Along with KGGen, we release the first benchmark, Measure of of Information in Nodes and Edges (MINE), that tests an extractor's ability to produce a useful KG from plain text. We benchmark our new tool against existing extractors and demonstrate far superior performance.
View PDF HTML (experimental) Abstract:Recent interest in building foundation models for KGs has highlighted a fundamental challenge: knowledge-graph data is relatively scarce. The best-known KGs are primarily human-labeled, created by pattern-matching, or extracted using early NLP techniques. While human-generated KGs are in short supply, automatically extracted KGs are of questionable quality. We present a solution to this data scarcity problem in the form of a text-to-KG generator (KGGen), a package that uses language models to create high-quality graphs from plaintext. Unlike other KG…
saved by
related reading
- The GraphRAG manifesto: Adding knowledge to GenAIneo4j.com
- [2003.02320] Knowledge Graphsarxiv.org
- [2005.11401] Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasksarxiv.org
- [2303.13948] Knowledge Graphs: Opportunities and Challengesarxiv.org
- Andrej Karpathy on X: "LLM Knowledge Bases Something I'm finding very useful recently: using LLMs to build personal knowledge bases for various topics of research interest. In this way, a large fraction of my recent token throughput is going less into manipulating code, and more into manipulating" / Xx.com
- [2308.06374] Large Language Models and Knowledge Graphs: Opportunities and Challengesarxiv.org
- GraphRAG: New tool for complex data discovery now on GitHub - Microsoft Researchmicrosoft.com
- knowledgator/gliner-multitask-large-v0.5 · Hugging Facehuggingface.co
- llm-wikigist.github.com
- title.txtsybrandt.com
- Knowledge graph vs. vector database for grounding your LLMneo4j.com
- Datacurve | The data engine for frontier AIdatacurve.ai