What Degree of Freedom for the Reader of Patrimonial Digital Editions?
Marie Bizais-Lillig (bizais@unistra.fr) and Xinmin Hu (xinmin.hu@unistra.fr), University of Strasbourg, UR1340-GÉO, USIAS, Strasbourg, France
Published in HT '23: 34th ACM Conference on Hypertext and Social Media · DOI: 10.1145/3603163.3609073 · License: © Copyright held by the owner/author(s). Publication rights licensed to ACM.
Keywords: Digital Scholarly edition, Digital Twins and Scientific Legitimacy, Freedom of reading, Influence of Editorial Framework, Interactive Reading, Knowledge Network, Textual Network
Session: Authoring, Reading, Publishing: Hypertext Authoring
Conference: HT '23
Abstract
In this paper, we present a practical project of building up an in-terconnected body of Ancient and Medieval Chinese texts, whichassociates records in a database, marked up texts, and a set of de-signed architectures (useful to structure individual texts and alsonecessary to set up a website). We first explain how we delineated,acquired and structured the corpus under study. We then explainwhy and how we set up a database which plays an important rolein the editorial pipeline. We finally present our editorial choices,and more specifically why we have decided to limit the explorationtools available on our website to the possibilities offered by hyper-text. The edition of this very large scholarly corpus is intimatelytied to a research project which also builds a knowledge networkand employs a variety of methods – from traditional text analy-sis to computational text mining – to understand how snippetsof knowledge circulated through texts. We ask: Does an editionproduced in the course of a research project necessarily reflect oreven bear the traces of this research project? To what extent shouldthe characteristics – thematic, methodological, etc. – of this projectinfluence future access to such an edition? To what extent should areader’s freedom be preserved?
1 Introduction
It has been argued that (enriched) digital editions interfere with linear reading habits and thus constitute a space of larger freedom for readers who can define their own individual paths and stroll through books.[19] This paper is meant to reassess the impact of hyperlinks on our ways of reading based on some experiments in the process of publication of the "Chinese Knowledge and Poetry" (CKP) Corpus. We argue that digital editions enriched through hyperlinks shape reading in an enlightening nonetheless highly intrusive way, and we identify easy to implement mechanisms meant to broaden possibilities for individual readers. In other words, we acknowledge the effects of research on digital editions and suggest that the framework that research projects impose on texts should be both more visible and to a larger extent optional. In order to operate in concrete terms, we first present the CKP Corpus, its origins and shaping. Then we explain the editorial pipeline and the main devices used for an enriched publication of the corpus. In the third and last part of this paper, we examine how the editorial choices for the CKP publication among many other projects shape the specific lenses through which texts are to be read and we present some of the adjustments that we have decided to implement as a means to soften our influence on readers.
2 Structuring A Body Of Ancient
CHINESE TEXTS
The CKP Corpus is meant to provide a large body of texts for alarger research project entitled "The role played by poetry in theeconomy of knowledge in Medieval China" (CKP). This researchproject draws on the observation that Chinese literati were ex-pected to be cultivated, to master poetry composition, to workas civil servants, to carry administrative or political tasks, to beknowledgeable when it comes to agriculture, taxes and the like,and also, to understand the ways of war and to be proficient armyleaders. Modern categories of knowledge which bring us to lookat Chinese literati culture through the lenses of disciplines henceappear incongruous. The CKP project, which was launched by oneof the authors and which leans on the work of a larger team,1 aimsto shed light on the way knowledge circulated through texts be-longing to or at least familiar to the Chinese literati world during
1 See list of contributors in the appendix.
the first millennium.2 The team computationally digs through avery large set of texts for co-occurrences and cases of text reuse toreach this purpose.
2.1 Sources
This corpus brings together a foundational poetic anthology of the Antiquity along with texts of the Middle Ages (i.e. 1st millennium)which belong to 3 different spheres:
• Belles Lettres with 4 large anthologies or historical collec-tions of texts appreciated for their formal and aesthetic qual-ities,
• Scholarship consisting of commentaries (on all the othertexts, including multiple layers of commentaries), 2 dictio-naries and 3 large early encyclopedias,
• Technical Knowledge with 4 essays or treatises and 1 verylarge collection of texts for practical use.3
The complete corpus would fit in approximately 100,000 pages oftraditional Chinese books.Because of the automated aspect of the project, this corpus needsto be digital and to be structured in a somewhat regular way. Also,because the output consists of either numbers (statistics) or excerptsfrom the texts, the CKP team needs to have the necessary rightson the texts to show the excerpts that establish demonstration.Unfortunately, part of the corpus is either not available in full textor licensed in such ways that the project could draw and publishstatistical results on it, but could not legally publish excerpts. Theseobservations led to the decision to produce an open corpus and toshare it.A collaboration with 3 libraries in France4 has hence been setup with the purpose to
• 1. digitize old editions,• 2. publish the images of the digitized books,• 3. acquire the full text of the books based on the images,• 4. edit and structure the texts, and finally• 5. publish all texts (and allow emendations if needed).
To acquire this part of the corpus in full text, the CKP team is train-ing a handwritten text recognition (HTR) model with the help oflinguists and engineers specialized in Artificial Intelligence withinthe Calfa team.5 Another part of the corpus is available under anopen license in the online Wikisource library. It needs emendationfor accurate transcription and systematic mark up. This editorialwork allows us to tag and structure our entire corpus based on theneeds of the CKP research project. All the files produced through
2The distinction is more specifically important when it comes to texts of referenceto heal diseases, to cultivate the land, or to come up with military strategies, textswhich might not have been authored by literati as such, but which were most probablyused by them in the course of their work as civil servants. The Chinese Empire isestablished around 200 BC. It is commonly agreed that a general transformation ofsociety and culture occurs around 200 AD (see for instance [8]). From this time downto 1000 AD, many conflicting dynamics are at stake, hence the use of "Middle Ages" todefine the period. Afterwards, a shift occurs, with a self-conscious cultural identity,as shown in [3]. The CKP project focuses on the so-called medieval period withoutignoring the legacy of Antiquity.3For a detailed presentation of the corpus: https://gitlab.huma-num.fr/chi-know-po/all-about-plants/-/wikis/home (accessed July 11th 2023).4Namely, the Bibliothèque universitaire des langues et Civilisations (BULAC, Paris),the Bibliothèque de l’Institut des hautes études chinoises (BIHEC, Paris) and the Bibliothèque nationale et universitaire de Strasbourg (BNU, Strasbourg).5URL: https://calfa.fr/ (accessed March 30th 2023)
this pipeline will be openly available in the Nakala repository. Awebsite is under construction to display the enriched edition.
2.2 Analyzing and Rearranging Textual Structure
To structure the CKP corpus, we decided to rely on the resourcesprovided by the Text Encoding Initiative (TEI) and produce XMLfiles.6 TEI is widely used in our discipline,7 it offers a very rationalenvironment for text edition, and it is flexible enough to adapt toour editorial needs.While the corpus is diverse in the sense that the text structurediffers depending on the genre of each text, it is crucial for theproject that we can easily identify segments of texts based on theirnature or characteristics (e.g. main text, commentaries, titles). Wehave thus defined one XML schema for each genre (i.e. anthologies,essays, dictionaries and encyclopedias). The complete corpus henceconforms to identical editorial rules. With sets of predefined rulesfor anthologies, essays, dictionaries and encyclopedias, we can callspecifically tagged segments of texts to display them in a browser.
Figure 1: 1 odd file per genre, the overall editorial structureand predefined set of tags
The digitized contents of a book, no matter if it is illustrationsor full text, follow the storage logic of computational data andgive us the freedom to break through the limitations of actualprinted materials. Instead of the general linear organization ofprinted materials which is reflected in page count, digital editionsprovide us with the opportunity of putting forward the tabular orrhizomatous structure expressed in the printed tables of contents.In other words, printed materials combine in the three dimensionsof space both the functions of storage and that of presentation,whereas digitized contents make it possible to dissociate the twoaspects.We use tree-branched tables of contents to transform a somewhatflat and linear printed text into a large database of XML files8 whichincorporates edited full-text units. A text unit is defined differently
6URL: https://tei-c.org/ (accessed March 30th 2023). We refer to the P5 Guidelines andcomply to version 4, which constantly evolves. The last released version, in April 2023,is version 4.6.0.7See for instance Hilde de Weerdt’s work and the MARKUS project.8These files are stored in an online Base X database and ’called’ to be displayed on thewebsite.
depending on the characteristics of each edited book. In the caseof an anthology such as the Yuefu shiji (Collection of Poems ofthe Music Bureau) where some prefaces, commentaries and musicnotations cover a series of poems, separate XML files store moregeneral layers of metatext. Starting with the smallest element, each XML file is attached to a part if relevant, each part to a section,each section to a volume, and in turn each volume to a specifictype. This tree is built using xi:include in the XML files which store,by means of indexing, all the elements of the level below that itincludes. Hence, the XML file at the level of a section lists all theparts it comprises.All in all, this device not only displays in an immediate repre-sentation the traditional organization of the material book,9 it alsotransforms the CKP corpus into a large network of books, whichare, in turn, structured as smaller networks of titles, textual ele-ments and references. The third part of this paper will present therationale behind the construction of these smaller networks.10
2.3 Editorial Pipeline: Step 1
The pipeline set up by the CKP team to take up this task willbe described using a concrete example, that of the Collection of Poems of the Music Bureau. It has been suggested in the previoussection of this paper that this anthology has a multiplex structuredifficult to navigate. Indeed, the editor of the book, Guo Maoqian(1041–1099), collected poems with a more or less loose connectionwith an imperial institution called the Music Bureau and organizedthem in a five-level structure book beginning with twelve types ofpoems, each type itself containing several volumes. In each volume,there could be sections or direct poems, or under the sections, directpoems or poem groups, and then the poems. The Collection of Poemsof the Music Bureau hence very well illustrates the complexity ofthe editorial work.This work starts with text in either HTML (extracted from Wik-isource) or XML-Alto (output of the HTR processing on images)format, which both provide us with uncertain but useful tagging oftitles and commentaries. Our first task consists of a double struc-turation of the text: a/ through its distribution in as many XML filesas necessary, b/ through its transformation to fit in the TEI schemathat corresponds to its internal characteristics.For the Collection of Poems of the Music Bureau, texts are col-lected in HTML from Wikisource. All the collected material is firstdistributed and stored in CSV files. Each of the 100 volumes of the Collection of Poems of the Music Bureau is saved in a separate CSVfile and each poem is saved as one line in each CSV file. CSV filesconsist of columns for poem types in volume, type number (1 to12), volume number, volume title in Chinese, section title, grouptitle, poem ID, poem title, commentary (944), music notation (88),poem text (5,410), author (597), dynasty (24), and page number.This structure gives us a clear visual layout for checking materialsand for potential revisions over time, carried out both manuallyand computationally. Based on the number of levels identified inthe text, we create a tree like the one presented in figure2. Except
9It represents the table of contents of the material book and gives prompt access tospecific parts of the text within just one click.10The characteristics of the large network which constitutes the overall corpus willnot be described in this paper, as they mostly make use of SQL queries and would, assuch, necessitate complex explanations alien to the hypertext topic at stake here.
for the last 4 columns, each cell is to become an XML file, somefiles serving mostly as containers thanks to the xi:include pointer,some others mostly consisting of tagged text. Each of these files isidentified by a set of metadata and by a unique identifier stored ina CKP database – which will be presented later.
Figure 2: Structure of the Collection of Poems of the Music Bureau
We also pay attention to the inner structure of each textual unit.For the poems included in the Collection of Poems of the Music Bureau, the Wikisource edition does not allow us to easily identifylines and stanzas. When punctuation marks separate lines andwhen the word jie isolates stanzas, we automatically segment thepoems into stanzas and lines. For poems with irregular punctuation,segmentation is produced manually by the expert in our team. Thissegmentation is manually added directly in the XML files.11 Thisstep of the work ensures that the poems are displayed in the mostlegible way on the readers’ screen. The tags are also useful forresearch purposes since they break out all chunks of text in a similarmanner, help to easily position these chunks within a very largebody of texts, and hence, open possibilities to connect fragments oftexts in the process of analysis or representation.At the end of this step, the digital edition of each book is encodedwith an emphasis on multilevel structures. Thanks to the hierarchi-cal structure built using xi:include, the different elements whichconstitute the poetic anthology are all interconnected within thenetwork-like anthology and more largely with all the texts editedby the CKP team.
3 Text Enrichment Based On Adatabase
The CKP Corpus is a collection of tens of thousands of files, eachcorresponding to one text unit whose size varies from one singlepoem to one long essay. Given that there is no determined point
11We automated the workflow as much as possible. This pipeline allows us to workon a large volume of texts to produce, within a few years, relatively complete andsearchable project corpora for both academic usages and the general public. However,to perfect the accuracy and reliability of the corpora and of the database, and tofinally obtain satisfactory text with good quality for further scientific research orother potential explorations, a lot of human work is still required. We validate thedigital edition by a comparison of it with a printed version, focusing on titles, poemnumbers and corresponding pages. The verification part for the Collection of Poemsof the Music Bureau takes months of intensive human work, including corrections oftitles, completion of missing text and information, and checking the missing charactersdue to encoding or other problems.
of entry inside the corpus, the website provides different ways ofaccessing the text.12 Most of the techniques provided to accesstexts use indexed elements stored in an SQL database.13 This SQLdatabase stores information on diverse types of data, namely: 1.people (names, period, date of birth and death); 2. titles of texts(alternative title, genre, author); 3. time periods; 4. words (the se-mantic categories they belong to, some texts they were found in,and words considered synonyms or equivalents based on specificdictionaries or encyclopedias).
3.1 Stored information
The CKP database was first set up because the exploration of the CKP corpus entailed a base of knowledge related to this set of texts.In order to filter the corpus under study to identify series of co-occurrences by a specific author or during a certain period of historyfor example, it is indeed necessary to either include this informationin the metadata of the XML files (i.e. in the header) or to tag theseelements within the body of these files. Each item in our base ofknowledge is attributed a unique identifier to avoid any ambiguity.Nonetheless, the length of the texts edited by the CKP team is suchthat many kinds of information could be worth identifying. The CKP database focuses on two aspects: bibliographical knowledgeand lexical information. However, the lexical part of the databaseis not used for editing purposes for reasons that will be discussedbelow. We shall hence focus on a few bibliographical entities toillustrate and justify our choices.People can either be the authors of texts, mentioned in titles(e.g. “Poem offered to my friend Wang Wei”) or referred to in anessay or a commentary. We decide to label all people in the editedcorpus inside XML files using their unique CKP project identifier.The format for a person is ‘ckpp-xxxxxx’, where the p letter afterthe CKP acronym indicates the ‘people’ category. For example, thepoet Li Bai is assigned a CKP id: ‘ckpp-1608’. We include this id asa reference to identify Li Bai within labels in the XML files whichstates whether, in this part of the text, Li Bai is the author (thisinformation would then be part of the metadata) or whether Li Baiserves as a reference. This id system helps us formulate a knowledgenetwork of the target corpus as well as with other encoded texts ofthis project through the connections documented by the database.We can then retrieve all other texts where Li Bai is an author (filter1) or a reference (filter 2). Full-text search would surely make itmuch more difficult to find all cases where Li Bai appears in thecorpus since Li Bai can also be called Li Taibai, Qinglian Jushi, Li Qinglian, Li Hanlin, but should not be mixed up with any otherhistorical figure whose name could sound similar. The CKP databasestores this kind of information.The same goes with titles whose identifiers are defined in theform ‘ckpw-xxxxxx’, where the w letter after the CKP acronymindicates the ‘writing’ category. The ckpw id is more specificallyuseful when it comes to identifying different versions of one text
12It must be underlined that all the XML-TEI files produced within the CKP projectare also openly available for upload and for reuse from the Nakala repository.13The database is also archived in the Nakala repository, free to upload and modify. Thedatabase is enriched and modified on a server using the Maria DB database managementsystem. A public version is produced based on this database. This public version is theone available on Nakala. The very same public version of the database is duplicated inthe Heurist environment to simplify public access and queries in the database.
present in different anthologies or collections. It also makes count-ing references and quotations within a set of texts extremely easy.If it is sometimes tricky to attribute the right identifier since manytexts may share the same title, such phenomenon, particularly com-mon in the Collection of Poems of the Music Bureau, supports thenecessity of relying on identifiers rather than raw titles.
3.2 Editorial Pipeline: Step 2
The CKP database was first set up using the SQL database Tilman Schalmey had designed in the course of his Ph D [15]. It was refinedand enriched based on the type of information we wanted to includein it. One of our main concern throughout the edification of the CKP database was to make sure that it would echo major open-source biographical databases in the field of sinology for the sake ofinteroperability and to profit from the time and energy which hasalready been spent by others to supply information on people ofthe past14. However, because of the specificities of these databases,it happens so that they include little information about the firstmillennium. As a consequence, our task is two-fold: a/ we need tomake sure that all the information needed is documented in ourdatabase or to supply this information, b/ we use this set of data totag the CKP corpus.We collect all the information structured in our CSV files (booktitles, poem titles, authors). We also use regular expressions andpackages in Python to work on XML files to dissociate people’snames and titles mentioned within the texts – mostly in prose. If theauthors or titles referred to in our text are not already present in ourdatabase, we add them by importing these references from anotheropen database (in less than 10 percent of the cases) or by enteringthe information manually. The database is hence continuously en-riched alongside the editorial work. It is very time-consuming worksince the documentation of authors requires supplying at least theirvarious names (family name, given name, social names, etc.), theirdates of birth and death, and the dynasty during which they lived.Each record within the database is given a unique identifier whichis then retrieved to tag all named entities within the XML files inthe next step.
Figure 3: The pipeline of the complete edition creation
Once all necessary information is available in the database andattributed a unique identifier, we can enrich the structure of the
14See for instance databases such as CBDB (https://cbdb.fas.harvard.edu/), DDBC/DILA(https://authority.dila.edu.tw/) and DNB (https://newarchive.ihp.sinica.edu.tw/), allaccessed July 20th 2023.
XML files with additional tags. At this step, texts are marked up withmore details: with labels such as <cit>, <bibl>, <title>, <author>,<quote>, within the body of each XML-TEI file, and by associatingthe tagged elements with identifiers from the database. In the CKPschema, a paragraph would thus be enriched as the figure4 shows:
Figure 4: Example of XML-encoded text
3.3 The power of indexes
The enrichment of the database and the parallel labelling of the XML files are mainly meant to build connections between differentparts of the corpus. Tables on time periods, people and writings areall interconnected thanks to the SQL database which allows us todocument them as precisely as needed in the project and to identifythem within the CKP corpus due to the use of unique identifiers.The CKP pipeline transforms all edited texts into a connectedbody of texts where all texts pointing to a specific record in the SQL database are hence linked together. This architecture makes itpossible not only to identify a unique title by a unique author ina type of text, but also to filter the whole corpus looking for theoccurrence of a person, either as the author of the text or as anindividual mentioned in the text. As a consequence, the websiteallows readers to scroll through the organization of books of interestrepresented in the horizontal tree of figure 2. The website alsomakes it possible to search for identified titles within part or theentirety of the corpus. It finally offers a set of filters to list all sourcesfrom a period, and a genre, linked to a concept. This also meansthat, once a reader has entered a text unit, this reader may focuson a person’s name or a title mentioned in this text. By clicking onthis element, the reader will access a two-folded page. The top ofthe page displays all the information stored in the different fieldsof the database in relation to this element. The second part of thepage lists all the other texts within which this element also shows –identifying a short passage from this text along with its title, authorand source. A simple click brings the reader to one of the selectedpassages in its larger context.As it is now, the concepts identifiable by these queries are limitedto time, people and title categories. The lexical part of the databases(which records words belonging to semantic categories such asplants, animals, topography, weather, feelings and emotions) hasnot been used to tag the corpus.
4 Digital Editions for What Kind of Reading?
The transformation of the very act of reading triggered by the trans-formation of the medium we use is at the heart of many studieson digital texts and hypertext. Some scholars underline the factthat this transformation is still in progress and that its outcomeis impossible to predict.[18][9] At the same time, some scholarshighlight that the impact varies depending on the nature of thetext, the approach of the reader (leisurely versus scholarly reading
for example), and that the dissemination of texts through intercon-nected blocks of textual units radically modifies the way we read15.Although one may see it as the freedom of the reader to “decidewhether to return to [the author’s] argument, pursue some of theconnections [he] suggest[s] by links, or, using other capacities ofthe system, search for connections [he has] not suggested”[9, p. 6],we would like to argue here that suggested links orient reading ina specific way. Based on the CKP editing experience and reflection,we argue that editorial projects impose their own agenda on thereader despite the transformation of the digital age and the factthat interaction and freedom are part of these editorial projects.
4.1 Digitized vs. Material Books at first sight
The ways material books present themselves to us are very diversethrough time and space. This diversity might be even more strikingwhen we think of the ways in which we connect to material books,as they also change from one individual to another. We can think ofdigitized books as collections of ordered pictures that are offered asa substitute for the original book. Such digitized display of materialbooks generally provides limited new ways of exploring books.16
However, when we think of enriched full-text editions, which mayinclude pictures of material books, what is at stake exactly? Towhat extent do digitized editions mimic characteristics of materialbooks? Do they belong to a distinct paradigm?One way of reflecting upon these questions consists of approach-ing them from the perspective of informal interactions with books.It is very common to assume that in the era of material printedbooks, reading is linear, beginning at one end of the book andending at the other end, only interrupted by pauses imposed bythe time reading demands. This oversimplified assumption is mostprobably the result of the development and success of novels sincethe end of the 18th century.17 However, when we first encounter abook, we often read its back cover, flick through it to get a senseof how it is organized or how it is written. We explore the book inits physical volume, so to speak, to find out whether we will delvefurther into it. Then, if we decide to read it, the way we will carryout this action will vary according to the genre of the book (thinkof the differences between reading a novel, a collection of shortstories or an encyclopedia for instance) or according to the kind ofreader we are. Finally, we could read it in a so-called passive way,or refer to its index to identify specific sections of interest in thetext, and even use its margins to write one’s own comments.18
15See for instance [5].16To say the least. See for example the images available on the wiki of the Chinese Text Project (https://ctext.org/, accessed March 30th 2023). To scroll from page 2 topage 3 in a Chinese traditional book, one needs to turn the page on the left-handside. Nevertheless, the arrows function in the very same way as for a European book,meaning that one has to hit the arrow to the right to turn pages to the left. Althoughcounter-intuitive, this setting is deemed secondary and thus not adapted to the materialdisplayed online.17Building upon [17].18Note that this mode of interaction is not specific to modern times. See how easeof reference could be part of the explanation for the development of the codex in Europe according to [13], see also the well-known representation of Erasmus at workby Dürer for instance. Philological studies and corpus linguistics all the more illustratea non-linear approach of texts albeit both disciplines emphasize the importance ofclose reading.
When it comes to digitized texts, the situation slightly differs.19
As we have noticed earlier, the organization of the content of thebook proves at the same time easier to grasp and a more immediatelink to dive into the actual text. Nonetheless, the efficiency of whatwe earlier called the tabular organization of text through files andhyperlinks seems to obliterate most of the informal and very muchidiosyncratic contact that we know of when it comes to materialbooks. This makes the flatness of material books relative, since,how are we to nonchalantly leaf through a screen?Far from trivial, this question reveals how digital editions differfrom paper editions and how they call for new experiences. We can,of course, look for substitutes to cope with what we identify as thelimits of digital editions. We can, for instance, randomly display apart of one text to mimic what happens when we haphazardly stopleafing through a book.20 We can also offer readers the possibilityto add their own bookmarks21 or their own commentaries.22
Such techniques, as clever and interesting as they are, do notseem to really touch the public. They might be the equivalent of thetechniques used in Europe when printing was introduced to pro-duce printed books that looked like manuscripts.23 The revolutionprinting brought about24, some may argue, was made possible pre-cisely because these new printed books looked familiar to the readerwho hence intuitively knew how to "read" them.[6] All things con-sidered, what is specific to digital editions is the way they invite usto explore reading paths, connecting one text to another throughhypertext. Many elements in the digital files are indexed and pro-duce an almost infinite way of exploring and reading through acorpus.
4.2 Hypertext and the Editorial Line of the CKPProject
In the perspective of an Open Science project such as CKP,25 itquickly became obvious that the corpus preparation carried outby the team needed to be shared with a larger community, hencethe elaboration of an editorial line and an online publication. Fiveessential principles governed this side of the project, knowing that
19The physical dimension of reading, the simple presence of hands holding a bookfor example, has been demonstrated since Michel Picard. [12] This physical aspect ofreading on screens is however outside the scope of this paper.20This casual access to books is what prompted us to display random files on thehomepage of the CKP website as some kind of weekly spotlight (we call it à la une).21See for instance this feature in the Roman de la Rose Digital Library (https://dlmm.library.jhu.edu/en/digital-library-of-medieval-manuscripts/, accessed April 19th 2023),documented by Stephen Mc Cormick.[10]22The tool called Hypothesis offers such possibility. It is present in the Stylo Environ-ment used for example in the editorial project [sens public] which is presented bymembers of its team.[16][20] Comments are added to online publications in multipleformats, including videos, dialogues. For a more theoretical approach, see [14] and [2].23Characterized by handwriting, diverse inscriptions and annotations, along withillustrations. Note that we are not saying that these devices were invented to allow anexperience with digital editions that would to some degree look like the experience ofencountering and reading a material book should disappear. They are most probablynecessary for digital editions to find their place in our culture(s). They have theirorigins in intellectual practices and are thus justified. Also, let us remember that evenvery recent typesets did imitate illuminated capital letters for example.[7]24On the very concept of ’revolution’ and what it implies, see [1].25Open Science includes the idea of opening the results of one’s work to students andnon-experts of one’s field. In the case of text publication, it makes a huge differencewith a limited consumption of energy, time and money.
1/ we were eager to share our work with specialists26 and non-specialists27, 2/ we were conscious that our editorial work wasincomplete and needed improvement over the years28, 3/ we wantedto make full use of the possibilities offered by the database and the XML corpus without considering the corpus as the result of theongoing research29
Principle n°1: Freedom of reading journeys. Online editions com-bine two separate functions: 1/ they store, organize and give accessto a number of files, and 2/ they constitute a tool set to explorethe corpus. We can hence dissociate information from display andexperiment with a diversity of tools to present our corpus. Thearrangement of the CKP corpus provides the reader with almostunlimited possibilities to roam through this body of texts. The nu-merous connections established between texts and the informationenclosed in hypertext offer a new perspective on these texts. Thelinks established within a commentary with previous texts andauthors may reveal, for instance, the weight of certain texts onthe intellectual framework of a commentator. The index of citedworks within a book will reveal this weight and give access toall mentions of the text. However, other approaches are possible.Readers might all the same prefer to ignore sections of a book orthe indexed elements, they may even choose to examine the digitaltwin of a material book. Such is the freedom we want to guaranteethe readers while offering them new ways of looking at the sources.
Principle n°2: Freedom of reading filters. We did not follow thehighly inspiring Re Nom project30 which relocates Rabelais’s mas-terwork by giving access to excerpts of his books through pointingat specific geographic places. Our choice was to keep separate thecorpus from the methodologies or computational tools used by the CKP team to explore it. We hence chose not to tag words in thecorpus based on the lexical records of the CKP database, because itwould bring our readers to identify the corpus to these semanticcategories rather than see it as a representative body of texts froma large period. Nonetheless, this rationale has its limits: taggingtitles, people and citations within texts and associating texts withhistorical periods and genres also shapes the way readers enter thecorpus. The enriched CKP corpus is indeed connected through hard,author-determined links which, objectively, facilitate circulationthrough multiple texts, but do not offer as many possibilities asone could imagine for the reader.31 Still, our website allows readers
26These specialists can upload XML files from the Nakala repository and further editthem to meet their own needs.27Such readers will be able to read part of the corpus online, copy and paste what is ofinterest to them, dig into the material we provide.28Due to limited resources, this short-term project could not fully proofread thismassive corpus.29It is, indeed, an output of our work, but the philological analysis carried on thecorpus stands as one out of many possibilities of exploration of this corpus, which ispresented in research papers and in a Git Lab repository.30URL: https://renom.univ-tours.fr/fr (accessed March 29th 2023)31Maybe because a scholarly edition is specific, the CKP team will not create anenvironment such as the ones set up by the editors of the [sens public] project (https://ateliers.sens-public.org/exigeons-de-meilleures-bibliotheques/index.html, accessed April 23rd 2023). The texts displayed in this collection combine links pointing todifferent sets of explanations and indexes, with retractable sections. It is possible tocirculate in many different ways in these editions whose chapters may be available indifferent languages and juxtaposed with explanations, commentaries, and annotationsusing different mediums (mainly text or video). The [sens public] digital explorationclearly transforms one’s sensory relationship to text.
Figure 5: The different components assembled in the website
to query or filter the CKP corpus and to export the results, i.e. toproduce their own corpus according to their needs.32
Principle n°3: Stability. Parallel to its editorial work, the CKPteam developed text mining tools to find co-occurrences and fuzzyresemblances in different sets of texts. These scripts were all writtenusing Python language rather than XSLT because of the possibili-ties and plasticity it gave us. We thought of refining our publicationinterface by offering the reader access to specific texts from thevisualization of co-occurrences of selected words. Such a projectwould require computer development to transform these scriptsinto online applications. In the long run, such tools would need tobe adapted and updated. The CKP project is a time-limited one andis not able to ensure the durability of tools of this type. As a con-sequence, circulation through the CKP website is strictly foundedon the possibilities offered by hypertext. Its stability is fortifiedby it being part of a larger publication platform called Estrades.33
The texts can be displayed on a computer, a tablet or a smartphone.Technically, the CKP website is grounded on a straightforwardarchitecture as figure5 shows.
Principle n°4: Frugality and Reusability. Fostering new ways ofapproaching the CKP corpus through the mediation of innovativetools could certainly bring new ’readers’ to consider this set ofancient texts. We could imagine a smartphone application whichwould scan plants in real life and connect them not only to bio-logical information but also to elements of cultural history suchas ancient poems, prose works, artworks, early encyclopedias ordictionaries where these plants are mentioned. Our database – in-cluding its lexical part – and our corpus could both be used forsuch an encyclopedic project since all the textual material we gath-ered and structured are part of linked open data. We could evenimagine such an encyclopedic project profiting from the possibil-ities of virtual space with Virtual Reality and Augmented Realitytechnologies. However, the CKP project focuses on the study of theliterati culture of the first millennium in China and will not furtherseek ways of modifying the appearance of our corpus along withthe very logic of reading. All the material produced through this
32We do not want to simply offer readers the possibility to upload XML files of interestand play with them at will: it is important that we provide them with pre-set SQLqueries to facilitate their work. Readers can nonetheless ignore our website and designtheir own SQL and XSLT queries. They can of course enrich the tagging in the uploadedfiles.33URL: https://estrades.hypotheses.org/(accessed July 19th 2023)
research project is available for reuse, including in fields such asthe one described above.
Principle n°5: Interactivity and Constant Improvement. What ismore central to the CKP project is the possibility to modify thefull text provided by the CKP editors. This possibility is crucialsince the corpus displayed on the CKP website is not fully curatedthe way the Collection of Poems of the Music Bureau is. The resultsof automatic text recognition and automatic text structuring callsfor collaborative emendation. An editorial pipeline will connecta mirror-version of the website content on the TACT platform34
for regular updates. Texts of the CKP projects on TACT will beaccessible to all users ready to sign up. These users will be allowedto modify the content of the text (associated with the image of thebook it transcribes) and to intervene where mark up is missing orinaccurate. Online collaboration with readers will help graduallyimprove the quality of the edition displayed on the website.
5 Conclusion
Hypertext has been understood in two different ways: as a technicalterm in the domain of computation and as a literal concept basedon the creation of links between stories, texts, and knowledge.[4]Although there is still a long way to go, it is encouraging to see that,over the past few decades, research and practical implementationsto which we hope that our project will contribute bridge these twoaspects. This progress means that imaginative projections are be-coming quasi-visible-realities.[11] It also proves that, by building upthese knowledge networks, we can more easily uncover knowledgestructures and intertextual links, and hence assist further readingexperience and research.We wonder about the status of a scholarly edition with regardsto a research project in the sense that if the edited body of texts isused by a team for research purposes, this does not imply that whenthe team shares this corpus with other readers, the corpus needs tobe colored by this team’s research objectives. However, because ofthe circumstances under which the edition was made, the readingexperience is very much influenced by the research project in theframe of which the corpus was prepared. The so-called freedom ofthe active reader of a virtual corpus should as a consequence bedownplayed.While human interactions with computers are actually and shouldbe very different from those with physical books, the imitation ofreal-world habits in the process of designing interfaces for humanuse bridles inventiveness when it comes to ‘reading’ online sources.In other words, at this stage of development of digital editions,we are more and more conscious of the possibilities offered by thenew media. It is nonetheless too soon to envision the whole range ofpossibilities. It is also crucial to reassess what is at stake in editorialwork and what differentiates it from research on a body of texts. The CKP project is fundamentally a research project on knowledge andtextual circulation in China during the first millennium. Because theresearch project depends on the structuration of a very large bodyof texts which could be of broader use and interest, the team hasdecided to share the data it produces and to open it to readers withlittle technical knowledge. This choice motivated the creation of a
34 For more information, see: https://tact.demarre-shs.fr/ (accessed April 20th 2023).
website. However, since the website is not reduced to an illustrationof a research project but stands as an editorial project by itself, someof the interventions in the texts – e.g. looking for plants throughoutthe corpus — have been separated from the corpus displayed onthe website. The experience of different users of this still under-construction website will tell us if the limits we have identified areproblematic or if they characterize patrimonial scholarly editions.35
They shall also reveal to what extent the very complex structure ofthe corpus based on hypertext – linking together texts and elementswithin texts – inspires our readers without influencing them toomuch.
References
[1] Sabrina Alcorn Baron, Eric N. Lindquist, and Eleanor Shevlin (Eds.). 2007. Agentof Change. University of Masachusetts Press, Amherst and Boston.
[2] Hélène Beauchef. 2014. Pratiques de l’édition numérique. Les Ateliers de [senspublic], Montréal (Canada), chapter 13: Concevoir un projet éditorial pour leweb, 205–219.Retrieved 2023-03-16 from https://www.parcoursnumeriques-pum.ca/1-pratiques/chapitre13.html Institution: Les Ateliers de [sens public].
[3] Peter K. Bol. 1992. "This Culture of Ours" Intellectual Transitions in Táng and Sung China. Stanford University Press, Stanford (California).
[4] Aurélie Cauvin. 2001.La Littérature Hypertextuelle, analyse et typologie.Retrieved 2023-03-21 from https://www.memoireonline.com/03/07/405/mla-litterature-hypertextuelle-analyse-et-typologie13.html Maitrise de lettres Mod-ernes.
[5] Emile De Rosnay. 2012. Le Coup de dés numérisé : modèles, défis, perspectives.Synergies Canada 4 (March 2012). https://doi.org/10.21083/synergies.v0i4.1689
[6] Elizabeth Eisenstein. 1979. The Printing Press as an Agent of Change. Cambridge University Press, Cambridge.
[7] Lucien Febvre and Henri-Jean Martin. 1958. L’apparition du livre. Albin Michel,Paris.
[8] Jacques Gernet. 1990. Le monde chinois (3rd ed.). Armand Colin, Paris. ISSN:0418-7784.
[9] George P. Landow. 2006. Hypertext 3.0: Critical Theory and NEw Media in an Eraof Globalization. John Hopkins University Press, Baltimore.
[10] Stephen P. Mc Cormick. 2016. A Guide to Digital Medieval Studies in North America. Perspectives médiévales. Revue d’épistémologie des langues et littératuresdu Moyen Âge 37 (January 2016). https://doi.org/10.4000/peme.9655
[11] Mélinda Palombi. 2018.Giacomo Leopardi et Italo Calvino : pour une œu-vre rhizome. In L’hypertexte et l’hypertextualité entre humanités numériques etjeux vidéo. Aix-en-Provence. Retrieved 2023-21-07 from https://hal.science/hal-01840791/document Daniela Vitagliano et Martin Ringot (org.).
[12] Michel Picard. 1986. La lecture comme jeu. Minuit, Paris.[13] C.H. Roberts and T.C. Skeat. 1983. The Birth of the Codex. Oxford University Press, London.
[14] Nicolas Sauret. 2020. De la revue au collectif : la conversation comme dispositifd’éditorialisation des communautés savantes en lettres et sciences humaines. Ph. D.Dissertation. Advisor(s) Zacklad, Manuel and Vitali-Rosati, Marcello. Retrieved2023-04-19 from https://www.theses.fr/2020PA100146
[15] Tilman Schalmey. 2022. Computerlinguistische Datierung schriftsprachlicher chi-nesischer Texte. Ph. D. Dissertation. Universität Trier, Trier, Germany. Heidelberg Asian Studies Publishing.
[16] Michael Sinatra and Marcello (eds) Vitali-Rosati. 2014. Pratiques de l’éditionnumérique. Retrieved 2023-03-16 from http://parcoursnumeriques-pum.ca/1-pratiques/index.html
[17] Peter Stallybrass. 2012. Books and Readers in Early Modern England. University of Pennsylvania Press, Philadelphia, Chapter 2: Books and Scrolls: Navigating the Bible, 42–79. Retrieved 2023-04-19 from https://www.degruyter.com/document/doi/10.9783/9780812204711.42/html Section: Books and Readers in Early Modern England.
[18] Geoffrey Turnovsky. 2023. Lectures imprimées, lectures numériques. Faculté des Lettres de Lorient. Retrieved 2023-03-05 from https://www-actus.univ-ubs.fr/fr/index/articles-chroniques/scvc/planete-conferences/planete-conferences-2022-2023/planete-conferences-lectures-imprimees-lectures-numeriques.html February 28th 2023.
[19] Christian Vandendorpe. 1999. Du papyrus à l’hypertexte: Essai sur les mutationsdu texte et de la lecture. La Découverte, Paris. https://doi.org/10.3917/dec.vande.1999.01
35The information on readers’ experience of the CKP corpus might be difficult togather. It is nonetheless planned to ask some of our visitors to fill a form as a way oftesting the results of our work.
[20] Marcello Vitali-Rosati, Nicolas Sauret, Antoine Fauchié, and Margot Mel-let. 2020.Écrire les SHS en environnement numérique. L’éditeur de texte Stylo.Revue Intelligibilité du numérique 1 (2020).Retrieved 2023-04-19from http://intelligibilite-numerique.numerev.com/numeros/n-1-2020/18-ecrire-les-shs-en-environnement-numerique-l-editeur-de-texte-stylo Institution:Bruno Bachimont.
Acknowledgments
• The UR1340-GÉO supported the CKP projects in its firststeps (2020).
• The Distam consortium made possible first hands-on ex-periments for full text acquisition using HTR technologies(2021-2022).
• The USIAS (University of Strasbourg Institute for Advanced Studies) is the main funder of the CKP project (2021-2024).
• The COLLEX-Persée is the main funder for digitization andfull text acquisition using HTR technologies aspects of the CKP project (2022-2024).
AAPPENDIX: CKP TEAM
The contributors to this project include (alphabetically ordered):
• Marie Bizais-Lillig (Principal Investigator),• Tara Cooke (Script testing),• Elsa Cuillé (HTR annotation and database enrichment),• Xinmin Hu (XML transformation and text editing, visualiza-tion, HTR annotation),
• Shueh-Ying Liao (Database management and enrichment,HTR annotation),
• Tilman Schalmey (Setting up of the database),• Ilaine Wang (XML transformation and text mining tool de-velopment),
• Mariana Zorkina (Text mining tool development).They are assisted by the members of the Calfa team for HTRtraining and models.The publication of the CKP corpus benefits from the Estradesproject, and more specifically from the experience and contributionof Guillaume Porte (for the overall publication infrastructure), andfrom Derek Salmon of Pikselkraft (for web-design).
Received 20 April 2023; revised 26 July 2023; accepted 5 June 2023
Do you like what you are reading? Subscribe to receive updates.
Unsubscribe anytime