Visualizing Data Uncertainty in the PZSBL
Historical databases, by their very nature, contain information that is incomplete, ambiguous, imprecise, or internally inconsistent. The Database of the Slovenian Biographical Lexicon (PZSBL) discussed here brings together data on individuals, including biographical events such as births and deaths, name variants, occupations and activities, related persons, and the sources from which this information was obtained. A more in-depth consideration of uncertainty in such a database not only contributes to a better understanding of historical events, but can also facilitate the generation of new knowledge and data through auxiliary tools such as visualisations.
Previous research on uncertainty in cultural-historical collections (Windhager et al., 2019) has shown that uncertainty can occur across various data dimensions, including temporal, relational, spatial, and provenance dimensions. Building on this approach, we propose a two-level model for defining uncertainty in the PZSBL database: provenance uncertainty and data uncertainty.
Provenance uncertainty describes uncertainty associated with the characteristics of data sources. We consider whether a source is primary or secondary, as well as its type, such as a manuscript, interview, questionnaire, lexicon, etc. We also record uncertainty at the level of the individual source, focusing on the quality of the data obtained from that source and assessed to date. This makes it possible to adjust uncertainty estimates interactively.
Data uncertainty describes uncertainty arising from the characteristics of the data themselves. At the level of an individual person, it takes into account factors such as the completeness of the record, the amount of available biographical information, missing attributes, and similarities or ambiguities between different entries. This enables us to identify individuals who are more difficult to identify with confidence due to the limited amount of information available about them. It also allows us to define uncertainty arising from a high degree of similarity between the data associated with two or more individuals, where it cannot be established with certainty whether the records refer to the same person or to several different people.
We also consider data uncertainty at the level of individual attributes. In the case of names, we distinguish less certain forms of representation, such as abbreviations and pseudonyms. For locations, uncertainty may arise from ambiguous historical place names or from the specification of a broader region rather than a precise location. Dates may be represented only by a year, a time span, or an approximate temporal designation. Attributes such as occupation or activity may likewise be less certain due to non-standardised or inconsistent definitions. A particular case of data uncertainty arises when multiple sources provide conflicting information. Its assessment therefore depends on the provenance certainty of the sources, the number of sources, and the degree of agreement between them, while also accounting for the possibility that inaccurate information may be repeated across the majority of identified sources.
The two dimensions are complementary: provenance uncertainty describes the reliability of individual sources, while data uncertainty describes how comprehensively and unambiguously an individual is represented in the database. Together, they provide a basis for the structured assessment of uncertainty at the level of individual persons and for the further development of computational and visualisation methods for quantifying and representing uncertainty.
*
Ahac Meden is a researcher at the Institute of Cultural History (IKZ) at ZRC SAZU, the Research Centre of the Slovenian Academy of Sciences and Arts in Ljubljana. He serves as subject editor for popular culture, media arts, and new media, as well as technical editor of the online edition of The New Slovenian Biographical Lexicon. He graduated in Cultural Studies from the University of Ljubljana in 2007 and obtained a master’s degree in Social Communication from Pompeu Fabra University in Barcelona in 2009. His doctoral research in anthropology focuses on the phenomenology of data in the context of digital humanities and open science.
E-mail: ahac.meden@zrc-sazu.si
*
Uroš Šmajdek is a doctoral student, early-stage researcher, and teaching assistant at the Faculty of Computer and Information Science, University of Ljubljana. He received his bachelor’s degree in Computer and Information Science in 2020 and his master’s degree in 2023, with a thesis on interactive physically based rendering. His current research focuses primarily on visualisation, digital humanities, and language technologies, while his broader interests include history, computer graphics, and game technology.
E-mail: uros.smajdek@fri.uni-lj.si