Abstract
Tabular data is one of the most common data sources on the internet and is widely used in various data analytics tasks. Identifying semantic concepts within tables is often a critical component of these pipelines, yet it remains a challenging task to automate. To address this problem, we present RAGDify, a large language model (LLM)-based system designed for the Cell Entity Annotation (CEA) task. Our system employs a three-step pipeline inspired by Retrieval-Augmented Generation (RAG) and advanced reasoning techniques: (1) retrieving context-aware candidate entities, (2) engaging in a debate-like evaluation to compare top candidates, and (3) applying chain-of-verification-inspired prompting to validate the final entity match. We propose RAGDify as a solution for the SemTab’25 challenge, targeting the key challenges inherent in automating the CEA task.
| Original language | English |
|---|---|
| Pages (from-to) | 221-228 |
| Number of pages | 8 |
| Journal | CEUR Workshop Proceedings |
| Volume | 4144 |
| State | Published - 2025 |
| Event | 20th International Workshop on Ontology Matching, OM 2025 - Nara, Japan Duration: 2 Nov 2025 → 2 Nov 2025 |
Bibliographical note
Publisher Copyright:© 2025 Copyright for this paper by its authors.
Keywords
- Cell Entity Annotation
- Entity Matching
- Large Language Models Reasoning
- Retrieval Augmented Generation
- Table to Knowledge Graph Matching
ASJC Scopus subject areas
- General Computer Science
Fingerprint
Dive into the research topics of 'LLM-Driven Retrieval, Debate, and Verification for Robust Table‐to‐Knowledge‐Graph Matching'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver