Robust metadata in multiple environments
We live in a time of wondrous technologies, but plain old academic articles remain hamstrung by clinging to the affordances of the paper media of the past. The affordances offered by digital media are not present in any depth; there is simply copy and paste and blue hyperlinks. The Visual-Meta approach makes even basic PDF documents contain rich metadata, which allows them to be active rather than passive knowledge components, and this can be done in any media in which the document can be render

Frode Hegland

Published in: HT ’22: Proceedings of the 33rd ACM Conference on Hypertext and Social Media · DOI: 10.1145/3511095.3536360

License: © 2022 Copyright held by the owner/author(s).

ACM source attribution. Complete text transcribed from the ACM Hypertext 2022 proceedings HTML, verified against the companion proceedings PDF, under ACM’s confirmed authorization. The terminal Visual-Meta appendix is publisher metadata and excluded from the scholarly article body.

Robust metadata in multiple environments

Frode Hegland Web and Internet Science Group (WAIS), University of Southampton · Southampton, Hants, United Kingdom · frode@hegland.com

ABSTRACT

We live in a time of wondrous technologies, but plain old academic articles remain hamstrung by clinging to the affordances of the paper media of the past. The affordances offered by digital media are not present in any depth; there is simply copy and paste and blue hyperlinks. The Visual-Meta approach makes even basic PDF documents contain rich metadata, which allows them to be active rather than passive knowledge components, and this can be done in any media in which the document can be rendered. This paper presents Visual-Meta and explains what it is, what the basic benefits are, and how it can augment knowledge in augmented environments such as AR/VR, often today referred to as the ‘metaverse’.

CCS CONCEPTS

• Applied computing → Document metadata.

KEYWORDS

Metadata, document, infrastructure VR, AR, physical media, Visual-Meta, Metaverse

ACM Reference Format: Frode Hegland. 2022. Robust metadata in multiple environments. In Proceedings of the 33rd ACM Conference on Hypertext and Social Media (HT ’22), June 28-July 1, 2022, Barcelona, Spain. ACM, New York, NY, USA, 4 pages. https://doi.org/10.1145/3511095.3536360

1 INTRODUCING VISUAL-META

Visual-Meta was first introduced at HUMAN ’19 in the paper ‘Visual-Meta: An approach to Surfacing Metadata’ [ 2 ]. In 2020, Visual-Meta was further discussed in HT ’20 in ‘Addressing the Skies of the Future of Text: A Call for Continuous Improvement in Infrastructures’ [ 3 ]. Visual-Meta 1.1 was implemented in the HT ’21 Proceedings, with Visual-Meta data being added to all published conference items 1 . Vint Cerf introduced Visual-Meta in his October 2021 ACM Communications column ‘Cerf's Up’ [ 1 ].

2 VISUAL-META IS AN APPENDIX FOR METADATA

Visual-Meta is an approach to metadata where metadata for citing the document, the document's structure, what the document cites, and any included glossary terms, etc. is ‘printed’ in an appendix, visually, in a format inspired by BibTeX. For this document it is:

author = {Frode Hegland},

title = {Visual-Meta in the Metaverse},

year = {2022}

This is then parsable by viewer software to allow the user to ‘fold’ the document into an outline (based on the headings being known), view and follow citations, and instantly (by clicking on citations in the body of the document to see full reference information in a popup) see glossary terms, for example, as has been implemented in Author, the macOS word processor, and in Reader, the PDF reader.

This is also parseable at scale so that large volumes of documents can instantly provide what sources they cite in order to build citation networks, glossary definitions, and information about what data is held in charts and graphs. This means queries such as ‘show me data from January-September 2014 in this corpus of documents’ can easily happen without the use of Machine Learning to ‘guess’ which data is where.

Already now Reader, our PDF reading software, can create and append a Visual-Meta appendix if the document the user opens has a DOI. This means that the basic citation metadata can be easily applied to legacy documents. This is important, and it helps illustrate the flexibility of Visual-Meta, since the system which appends the Visual-Meta announces itself in the header as either ‘generator’, if it was the software which produced the Visual-Meta, or ‘appended by’, if it was appended by software systems later. This is also how further Visual-Meta appendixes will be produced when leaving a VR environment where specific views may have been created, based on the metadata which only the specific VR environment can support. A further Visual-Meta appendix will then be appended with all the VR-related view data as well as the VR environment's name and version number on export, so that it may scan for such information when opening the document again. This method also allows for appending errata if there is a reason to update the document non-destructively in the future.

3 BASIC BENEFITS OF VISUAL-META

This simple addition to the document confers the advantage of documents becoming active components in a network and being able to offer readers rich views and interactions to augment their understanding, while also saving themselves time. Specific benefits of Visual-Meta include:

    Surfacing of Data : Visual-Meta can include glossaries in a structured manner, as well as descriptions of any graphs or tables, to enable server or client-side software to easily draw out specific data for use without relying on fragile ML techniques. Visual-Meta can also include full References in structured form in order to make citation analysis simple. Furthermore, it allows encoding author names and affiliations in an unambiguous manner, something basic citation formats can have issues with.

    Rich Interactions & Views : With the document being able to communicate what it is and how it relates to the reading software, it becomes possible to read documents in new ways; folding documents into outlines, performing more useful Find and semantic analysis, have citation information and endnotes appear on click, and much more, as is shown on the https://visual-meta.info website. This is only the beginning of what can be done; we are now working on optimal ways to use Visual-Meta to take rich metadata into and out of VR/AR environments.

    Open & Low Cost : Because it is human readable, it is easy for anyone to understand and implement. In traditional manuscripts much of the metadata already exists, but it is normally stripped on publishing. Citation data is frequently stored by a publishing website as BibTeX or similar, and this is something which can easily be converted to Visual-Meta (Visual-Meta style is based on BibTeX). This also makes Visual-Meta extensible.

    Robust : Underpinning this is its inherent robustness. Since it is a human and computer readable appendix, it does not interfere with any PDF reader software. The document can have the underlying encoding changed, and the document can even be printed, scanned, and have Optical Character Recognition (OCR) performed on it, and none of the metadata will be lost.

4 VISUAL-META IN THE METAVERSE

4.0.1 Limits in the room where the Metaverse happens. In the physical, real world we can take a piece of paper and go into any office, any room and do whatever can be done with a piece of paper. We can fold it, annotate it, cut it, paste it, crumple it and even turn it into a paper airplane should we so wish. In a software environment we can only do what the software system affords. Even something as simple as annotating a PDF in different readers can give different results, and many advanced functions such as seeing only the names in the document to help you find something is only available in a few pieces of software, not all.

In a virtual environment this also holds true, but it takes on a new significance because we can easily tunnel in a feed of our traditional computer screens so that we can work on one or many monitors in the virtual environment. What we cannot do, however, is take a document out of that screen and fully into the virtual environment. As far as the software which produces the environment is concerned, the screen is simply an image, a texture, not text.

In terms of how to think about connecting different environments and media types, it is worth considering the perspective that we are already living in the Metaverse (also called cyberspace and other names) but we predominantly access it through flat screens. It is time to acknowledge the multi-dimensionality of digital information, something which VR and AR headsets help lay bare.

Visual-Meta in the Metaverse can be parsed simply by appearing, since it can be read as OCR even if it is just presented as an image, a texture, and all the metadata will then become available to the virtual environment. Visual-Meta in the Metaverse can be parsed since it can be read as OCR even if it is just presented as an image, as a texture, and all the metadata will then become available to the virtual environment. This simplicity allows documents to retain their rich metadata value across analog and digital formats and even into future environments.

4.1 The Metaverse will not replace flat screens

The physiology of the human visual system, with foveal vision for high resolution at the centre of eyes, nested in a highly movable eye-socket within a limited range of motion themselves, in a human head with more potential for movement but at greater energy cost, suggests that close reading will remain optimised at what a high-resolution computer screen can provide today. Seeing connections for comprehension or composing, however, can easily benefit from a radically increased field of information display. It is for these two different text interaction modes that Visual-Meta serves as conduit. Currently, it is not difficult to sit down with a commercially available VR headset, such as the Meta Quest 2, and choose to have computer screens appear as passthrough video. However, what appears on the screen is only a texture in a virtual environment; it does not impart semantic or interactive information. With Visual-Meta attached to the document, however, even if there is no native import of the document into VR space, the VR system could read the Visual-Meta and provide the user with all the benefits of the metadata in this rich environment.

For example, the user may choose to ‘pull out’ the metadata into VR space and put up a horizontal table of contents above the document, have all the glossary terms appear in a constellation on the user's right and all the references as lines going into the background, able to pull forth any cited document, Ted Nelson style.

Furthermore, as the user pulls out different pieces of the document and maybe even ‘throws’ the whole document up against a wall so that they may walk up and down on it and have lines appear on request, showing connections within the document.

4.1.1 The Metaverse has already taken over, we will just experience more of it in VR/AR soon. When the user is done working in the rich environment, a new Visual-Meta appendix can automatically be appended with all the spatial information that was involved. When the user exits the VR space and continues to interact with the document in ‘flatland’ none of that augmented environment data and work is lost. When the user goes back into VR, a flick of the hand can bring the document back into the same rich state.

5 ‘OWNING’ KNOWLEDGE

With Visual-Meta we will be able to truly take our knowledge with us and interact with it in ways which suit us, and which whatever media access device we choose to use, can be used to its full potential.

Reading on a tablet, laptop, in full VR, or pulling sections out in AR will become possible. We will not have information silos, we will have a truly blue sky of possibilities for how to deal with our knowledge.

Ownership of digital information in terms of intellectual property and privacy will be important, but even the most basic aspects of the information—being able to access and open/view the information, and interact with it—will not be a given. Information is power, so this ownership must be taken; no corporate body will give up this power automatically. We must develop the technical means to hold on to our knowledge; we must develop ‘vehicles’ of knowledge, free to roam on any highway if they follow basic rules; we should not build computer game worlds where we are not free to take our knowledge with us.

6 ARTIFICIAL DAY

Computers have long used their GPUs to augment their CPUs. We have used our occipital lobes, where we primarily processes vision, to augment our prefrontal cortex, where our executive functions lie for as long as we have been around. Computers have so far only had a small role in this for most knowledge workers, since the visual presentation from computers are confined to relatively small rectangular screens.

When the visual computer representation becomes all encompassing, the equation changes and we need to look at changing our mental perspective on this visual perspective.

The term ‘artificial’ now means that something is not genuine, whereas King Charles II used the phrase “very artificial” while praising Wren's St Paul's Cathedral.

We should now look at what we mean by the terms ‘virtual’ and even ‘reality’ in light of this new day. Similar to the evolution of the word ‘artificial’, ‘virtual’ originally referred to the virtue, of inherent qualities of moral character. Taking etymological history maybe a bit far, and including the fact that the earliest known use of ‘artificial’ in English was in the phrase ‘artificial day’, referring to the part of the day from sunrise to sunset, as opposed to the 24 hour day, we can ask ourselves not only how we can change our perception in the Metaverse, but also ask ourselves if we have the virtue to do what is moral and right, in this new environment. It can be what we want it to be; us humans will shape the Metaverse and it can be a human heaven or hell. It is truly up to us to shape this new day, and we are now at the new dawn.

7 THE POWER OF IMAGINATION

Before an explorer sets out on an expedition, such as the Polynesian Islanders, the vikings who were my forebears, or the space travel I hope and expect our descendants will be able to carry out, there must be a goal in the mind of the explorer. It would be suicidal to set out in a vessel with no expectation of reaching a shore, a new world, a goal.

In the early 1950s a young Doug Engelbart used his imagination to dream of a truly augmented world where computer systems and social systems would work together to increase our ability to understand and communicate. He gave the world his vision in 1968 in the ‘mother of all demos’ and we are still struggling to catch up to what he foresaw. Once companies sold us desktop publishing, word processing, spreadsheets and web browsers, the public's imagination was satisfied. We were impressed by ever faster computers and lacked the collective imagination (and commercial and academic vehicles) to demand ever more powerful augmentations. The PC and smartphone revolutions augmented us, without question, but still never came close to Engelbart's vision. With the Metaverse we now have hope—maybe humanity's last—of inspiring people to realise how advanced technology can augment our ability to think, communicate and become better human beings, not just to play better games or have more visually attractive social experiences.

8 WHAT ARE WE DOING ABOUT IT?

How we will work in the Metaverse during this dawn will determine how we will exist during the long day ahead. Will we be lucid and involved or drugged and distracted by the whiz bang theatrics of it all? I believe we can each draw our own conclusions. To borrow from Lin-Manuel Miranda's Hamilton in two linked quotes: We can put our pencil to our temple and connect it to our brain, and write our way out. This is why we must all think about the future of text and the infrastructures which can enable us to rise up from our early digital slumber and write our future systems into existence, where we robustly own our connected knowledge. Where we think is where we go. We must encourage further thinking through dialogue and demonstrations. And that is exactly what we are doing.

We are a small, virtual (naturally) group including people from Meta, Apple and academia, with the founder of the modern Library of Alexandria, Ismail Serageldin, and the co-inventor of the Internet, Vint Cerf, as advisors.

We have hosted The Future of Text Symposium for over a decade, published two volumes on The Future of Text [ 4 ] and we are now publishing a monthly Journal. Our Lab, the Future Text Lab, holds Open Office Hours twice a week, every Monday and Friday, focused on working with text in VR and AR environments: https://futuretextpublishing.com

We are now seeking to expand our efforts and invite you to join us.

9 NOTE: READING THIS VISUAL-META ENABLED PDF IN ‘READER’

This document is distributed as a PDF which will open in any standard PDF viewer. If you choose to open it in our free Reader’ PDF viewer for macOS, you will get useful interactions because of the inclusion of Visual-Meta, including the ability to fold into an outline, click on citations, select text and Cmd-F to ‘Find’ all the occurrences of that text–and if the selected text has a Glossary entry, that entry will appear at the top of the screen–and more:

https:// www.augmentedtext.info for free download of Reader for macOS. http://visual-meta.info to learn more about Visual-Meta.

ACKNOWLEDGMENTS

I would like to thank the following for their support and encouragement on the journey of realising Visual-Meta as part of my PhD work and beyond; Dame Wendy Hall, Les Carr and David Millard who mentored and advised me through my PhD on this topic. Mark Anderson and Christopher Gutteridge of the University of Southampton for invaluable dialogue. Also thank you to Doug Engelbart who was my mentor so many years ago and is greatly missed. Thank you also to my advisors Vint Cerf and Ismail Serageldin. Further thank you to the Future Text Lab team for helping extend this thinking into VR, who are, in addition to those listed above: Peter Wasilko, Rafael Nepô, Alan Laidlaw, Brendan Langen, Fabien Benetou, Brandel Zachernuk, Keith Martin and Adam Wern. I would like to thank Wayne Graves and Craig Rodkin of the ACM Digital Library for his support in making the pilot of Visual-Meta real, and for the chairs of Hypertext ’22, Alejandro Bellogín and Ludovico Boratto for supporting its inclusion in their conference proceedings for the second time running. Jacob Hazelgrove implanted the ideas presented here. This paper is dedicated to the memory of Enamul Hoque.

REFERENCES

    [1] Vinton G. Cerf. 2021. The Future of Text Redux. Commun. ACM 64, 10 (2021), 5. https://doi.org/10.1145/3483539

    [2] Frode Hegland. 2019. Visual-Meta: An Approach to Surfacing Metadata. In Proceedings of the 2nd International Workshop on Human Factors in Hypertext . Association for Computing Machinery, New York, NY, USA Hof, Germany, 31–33. https://doi.org/10.1145/3345509.3349281

    [3] Frode Hegland. 2020. Addressing the Skies of the Future of Text: A Call for Continuous Improvement in Infrastructures. In Proceedings of the 31st ACM Conference on Hypertext and Social Media . Association for Computing Machinery, New York, NY, USA, 127–129. https://doi.org/10.1145/3372923.3404779

    [4] Frode Hegland (Ed.). 2020. The Future of Text . Future Text Publishing, London, UK. https://doi.org/10.48197/fot2020a

FOOTNOTE

FOOTNOTE ⁎ Corresponding author & responsible for creation of primary dataset. 1 The HT ’21 papers are online at https://dl.acm.org/doi/proceedings/10.1145/3465336 Permission to make digital or hard copies of part or all of this work for personal or classroom use is granted without fee provided that copies are not made or distributed for profit or commercial advantage and that copies bear this notice and the full citation on the first page. Copyrights for third-party components of this work must be honored. For all other uses, contact the owner/author(s). HT '22, June 28–July 01, 2022, Barcelona, Spain © 2022 Copyright held by the owner/author(s). ACM ISBN 978-1-4503-9233-4/22/06. DOI: https://doi.org/10.1145/3511095.3536360

Do you like what you are reading? Subscribe to receive updates.

Unsubscribe anytime