FIZ Karlsruhe
MathD2 Dataset for Mathematical Term Disambiguation
Pages
14
Time to read
33 mins
Publication
Language
English
Pages
14
Time to read
33 mins
Publication
Language
English
This technical report presents the MathD2 dataset aimed at addressing the challenge of disambiguating mathematical terms within scholarly literature. The report outlines the complexities involved in understanding how contextualized textual representations can aid in distinguishing between multiple meanings of mathematical terms. It describes the construction of the MathD2 dataset, which is derived from ProofWiki's disambiguation pages, and details three methodologies for automatic disambiguation: supervised classification, zero-shot prediction based on semantic textual similarity, and zero-shot LLM prompting. The effectiveness of these approaches is evaluated, with results indicating that the first two methods achieve an accuracy greater than 0.9 on the ground truth dataset. Furthermore, the report discusses the limitations of existing initiatives in mathematical term extraction and disambiguation, emphasizing the need for improved methodologies in this domain. The findings contribute to the ongoing research in mathematical information retrieval and knowledge discovery.