What is missing from this picture? Persistent homology and mixup barcodes as a means of investigating negative embedding space
Himanshu Yadav, Thomas Bryan Smith, Peter Bubenik, Christopher McCarty
TL;DR
This paper addresses how negative embedding space in topic model embeddings can reflect missing context and potential interdisciplinarity. It combines persistent homology to identify holes in the embedding landscape with mixup barcodes to quantify how unobserved publications fill those holes, using top2vec embeddings of dimensions.ai UF data. The study finds $252$ persistent holes, and shows that older pretraining documents more often fill these holes than newer posttraining documents, with hole specific patterns revealing interdisciplinary fillings. The results suggest negative space encodes historical context and selective recombination, offering a framework for understanding conceptual landscapes, guiding pretraining strategies, and informing science policy and funding decisions about interdisciplinarity and missing context.
Abstract
Recent work in the information sciences, especially informetrics and scientometrics, has made substantial contributions to the development of new metrics that eschew the intrinsic biases of citation metrics. This work has tended to employ either network scientific (topological) approaches to quantifying the disruptiveness of peer-reviewed research, or topic modeling approaches to quantifying conceptual novelty. We propose a combination of these approaches, investigating the prospect of topological data analysis (TDA), specifically persistent homology and mixup barcodes, as a means of understanding the negative space among document embeddings generated by topic models. Using top2vec, we embed documents and topics in n-dimensional space, we use persistent homology to identify holes in the embedding distribution, and then use mixup barcodes to determine which holes are being filled by a set of unobserved publications. In this case, the unobserved publications represent research that was published before or after the data used to train top2vec. We investigate the extent that negative embedding space represents missing context (older research) versus innovation space (newer research), and the extend that the documents that occupy this space represents integrations of the research topics on the periphery. Potential applications for this metric are discussed.
