Abstract
This study explores how students of Ancient Greek interact with digital commentaries enhanced by inline reference resolution. Comparing static and hover-based interfaces, we identify distinct behavioral patterns and preferences shaped in part by prior tool use. Findings inform hypertext design strategies that balance cognitive load with access to contextual scholarly information.
CCS Concepts: • Applied computing → Digital libraries and archives; Annotation; • Information systems → Digital libraries and archives.
Keywords: Digital Libraries, Classical Commentaries, User studies, Hypertext reading
1 Introduction
Digital hypertext is relatively young in comparison to its many precursors, which we may call analog linked data 38. Commentaries, one such tradition, have existed for millenia, and are linear lists of annotations on a foreign language source text, which we argue is a construction approaching the definition of a hypertext. Many commentaries attempt to provide full clarity on the unfamiliar historical, cultural, and linguistic contexts these texts were written in, yet limitations of print commentaries exist in commentary-as-digital-hypertext projects 1. For example, in commentaries adapted for Perseus Digital Library's Hopper 10, references such as "emphatic by place: cp. 46." are hyperlinked so that clicking on "cp. 46" opens that note in a new web page. Factoring in observations on cognitive load in hypertext reading environments 39, we wondered if such context-switching was close to switching between notes in physical commentaries, and whether there were alternate modes of consuming the hypertext structure of source text, commentary, and text(s) referenced that mitigated cognitive load. Note that Perseus Digital Library has two hypertext reading environments: Perseus Hopper, an older website featuring alongside its source texts notes and references containing links to the material referenced 10, and Perseus Scaife, a newer website also featuring interactive sidebars for vocabulary, morphology, and more 12.
References made more sense when ink and space were limitations, yet they still exist in born-digital commentaries, often with scant explanation of their relevance to the source text, assuming the reader is curious enough to open the link and switch between sources to better understand the source text. Some language learners may be confused by such unexplained links, which could lessen the chance of clicking similar links in the future. A strategy we thought would make these references more approachable is if references were accompanied by a relevant excerpt of the text being referenced. We call this feature inline reference resolution.
In this paper, we propose inline reference resolution as an alternative to linked references, with two different implementations. This work was also an opportunity to learn more about different kinds of language learners, so we ran a user study surveying students from two different undergraduate level Greek classes about their hypertext experience, and tracked their interactions reading Sophocles' Antigone with a version of Richard Jebb's commentary on the play augmented by inline reference resolution. We also wondered what different commentaries expected different levels of language learners to benefit from most 1, so we coded notes from Jebb's commentary in the user study with our taxonomy from 1 and analyzed interaction data through that and other lenses. Here are our hypotheses for this study:
H1. Since inline reference resolution will lower the cognitive load of having to actively build and follow a reading plan that comes with hypertext 39, students will prefer it over the hyperlink norm present on platforms such as Perseus Hopper.
H2. Second-semester Greek students are more likely to seek out syntactic aid and translation and return to notes previously viewed, compared against their more advanced counterparts.
2 Background
2.1 Commentaries as Precursors to Hypertext
Scholars have been annotating to explain texts in unfamiliar languages and from obscure cultural settings for millenia (for ancient Greek scholarship, for example, see 15). The potential value of these annotations changed as canonical citation schemes emerged, based on units such as book/line (e.g., line 34 of book 6 of Homer's Iliad) and book/chapter/section (e.g., section 2 of chapter 42 of book 2 of Thucydides' History of the Peloponnesian War). Canonical citations do not depend on page breaks in a given edition. While a work's precise text will vary in minor ways across editions, citations such as "Hom. Il. 6.34" and "Thuc. 2.42" have described the same chunks of text for generations. When an annotation includes one or more canonical citations, the note becomes one node in a network branching out into text chunks that the citations cite.
Figure 1 shows an excerpt from the first volume of Edward Gibbon's Decline and Fall of the Roman Empire 19. He describes the religious life of ancient Persia by translating a short section of Herodotus' History of the Persian Wars. The footnote for this passage is overlaid onto the text. Citing this passage is problematic: Gibbon's quote appears on page 217 of a particular printing of one specific edition, but that passage will appear on completely different pages across editions. We cannot use chapter and footnote (e.g., chapter 7, footnote 17), since the editor of this Gibbon edition added his own footnotes – our footnote 17 is footnote 16 in other editions. Conversely, when Gibbon cites Herodotus with the string "l. 1, c 131", that reference designates liber (book) 1, caput (chapter) 131. That citation describes the same chunk of text in 2025 as it did in 1776. Figure 2 shows a still widely used commentary on Sophocles' Antigone, published by Sir Richard Jebb in 1891. The Greek text is on the upper left with an English translation on the facing page. There are two levels of annotation, each containing abbreviated citations. The annotations right below the text show information about alternate readings and corrections suggested by scholars. The lower level of annotations, which we center in this paper, comment on a wider range of topics for a more general audience.
Students of Greco-Roman culture developed canonical citations for many widely read sources. Projects such as Perseus Digital Library 9, New Alexandria Commentaries 17, Dickinson College Commentaries 18, and the Ajax Multi-Commentary Project 42 presented digital reading environments using commentaries and other helpful translation tools alongside digitized source texts. Often these environments are websites where usually the left panel contains source text, and the right panel contains a combination of interactive digital aids specific to the text and platform, including selected relevant commentary that includes links to relevant material that can only be seen in a separate window. Some of these projects use born-digital commentaries, but by digitizing print commentaries and the print resources they cite, the graph of information needed to understand a text is more visible, though understanding these graphs is a broad area of study yet to be fully explored. Hypertext is commonly defined as a graph of textual nodes where edges exist as navigational mechanisms to other nodes, and we argue that the conventions behind digital representation of commentaries make these commentaries eligible to be considered hypertext.
2.2 Cognitive Load and Hypertext
Cognitive load in reading hypertext comprises working memory limitations and cognitive difficulties processing the nonlinear structure that is hypertext. It has two types: germane load, or the difficulty of understanding the hypertext's content, and extraneous load, or the difficulty in understanding the hypertext's presentation 39. The latter is what we intend to affect with inline reference resolution. A few factors increase the extraneous load hypertext readers experience. First, unlike linear texts, hypertexts demand the reader to more actively develop and follow a reading plan 30. This effect contributes to extraneous load, yet some readers learn better and are more actively involved in reading when reading this way 7. Completing tasks in these reading plans can also add extraneous load. Reading plans often have navigational tasks (deciding what to read next) and informational tasks (reading the text and synthesizing coherence with previously read texts). Navigational tasks like scrolling and switching between different windows can increase extraneous load 33.
3 Related Work
Citation augmentation is not new. 5 exists as one example, augmenting citations to help scholars track them, yet our work centers improving the referential hypertext experience for students. 44 turns a linear text into a hypertext and tracks cognitive load, but our work augments a pre-existing hypertext structure. Work by 39 presents a variety of investigations into cognitive overload when reading hypertexts, and some themes from this guiding our work are 1) links demanding more decision-making in the reading plan, 2) switching between windows as additive to hypertext reading's cognitive load 33 and that 3) though these are invaluable insights, there is a lack of more recent research on how students interact with hypertext reading environments, especially since students interact with more digital hypertext than they did decades ago. Elsewhere, user studies 6, 35 covered annotations' impact on students' learning and retention of the source material, and the former of these two inspired us to consider the role of regressions in user behavior for this work. 3, 14, 22, 43, 48 are adjacent user studies on making hypertext reading environments more effective, but these are studies done from a cognitive science perspective as opposed to a human-computer interaction mindset. Conversely, 2, 28, 46 are a few of many computer science studies using metrics closer to ours to track engagement, but like many other studies, they are more geared towards advertising-adjacent engagement. 41, 45 are both studies that track clicks and hovers in Wikipedia navigation. While both the methodologies practiced and the hypertext itself resemble commentary-as-hypertext, we are testing alternate presentations of information with the goal of reducing cognitive load. Commentary-as-hypertext environments abound, as mentioned earlier 9, 17, 18, 42, but there are few user studies 24, 34 and fewer recent entries in this niche. 29 analyzes the role of commentaries and annotations in student learning that agrees with our view of commentaries, but this work is a user study building on those thoughts. 37 also strongly argues for rendering hypertext annotations in a way that makes them accessible to more readers, and we attempt to do so by showing excerpts of the text referenced in this user study.
This work also contributes to the small amount of research studying commentaries as hypertext and their usage patterns. In the last 20 years, there have been several monographs on printed classical commentaries 8, 20, 31, yet there is still a lack of research on how users actually use commentaries, or how hypertext reading environments can best help commentary consumers. An early broach of this matter can be seen in 36, but more recently 21, 23, 26 presented pros and cons of commentaries as hypertext, though most assume academics as the primary commentary consumer. 32 mentions rendering the linked data of medieval Arabic commentary, though they focus more on OCR. Some studies center students as commentary consumers, as 47 explores an alternate commentary structure where students comment on the source text in Google Documents and submissions are vetted, and 25 boldly argues against the use of traditional commentaries in the classroom. Somewhere in the middle is 1, where a non-exclusive taxonomy was devised for types of aid provided by commentaries and applied across many commentary types to reveal both differences in aid distribution and references provided, as well as the cost of and other barriers to accessing the full insight of commentaries. 1 represented references in commentaries as a unique barrier to synthesizing full insight that this work intends to ameliorate.
4 Methodology
4.1 Participants
Participants were 16 students at a university in Medford, MA, who were enrolled in second-semester, or intermediate (n = 8) and fourth-semester, or advanced (n = 8) level classes in ancient Greek literature. Informed consent was obtained for all participants. Participants were alternatingly assigned to one of two experimental groups. Due to user error, we were unable to capture activity records for 2 participants, though all survey responses were valid.
4.2 Materials
4.2.1 Selecting A Hypertext
We chose lines 100-120 of Sophocles' Antigone as our source text for this study, along with Richard Jebb's commentary on the play as our source text companion 11, 27. The intermediate class was already reading Antigone with Jebb's commentary as one of a few translation aids for the text. Jebb's commentaries on Sophocles' works are reference-heavy 11, and we wanted to pick a section that was interesting and dense enough to apply inline reference resolution to without completely disorienting students. The references are frequent and diverse enough regarding what they refer to, but it was important to select enough for a student to interact with without giving them too much to read, and this section of Sophocles' Antigone fit our criteria. However, since translation is more of a germane load than reading in one's native language, both classes reviewed the lines together in class before participating in the user study. Our goal was to study what happens during an initial reading of the lines, so participants had up to 10 minutes for reading.
4.2.2 Augmenting Jebb's Commentary
To rigorously test inline reference resolution, we needed a commentary that presented references often so that resolving them outside of the test environment added to cognitive load, whether that meant clicking on a link, opening another website to see the referenced material, or physically opening a copy of the same material. As elucidated in 1, commentaries from more than a century ago were more likely to have more references in their items than more recent commentaries, so one may assume that Jebb's work has more references, which is supported by 16, 27. In fact, of Jebb's 32 notes covering the 21 sample lines, we found 92 references to specific source text lines, reference work items, or commentary notes, with each surface-level note having at least one reference. That means each note in this section of Jebb's commentary has a little over 4 references on average. This hypertext subset of 32 notes-as-nodes is larger than hypertexts used for similar studies by Blom et al. (15 nodes, 3) and Vörös et al. (19 nodes, 48). From there, we wanted to find which references needed more cognitive load to understand, and then apply inline reference resolution on them. There were two categories of reference use that warranted inline reference resolution. The first category of reference use was a reference deployed without a quotation in Greek or English of what was being referenced (i.e. "emphatic by place: cp. 46." 27), and the second category of reference was a reference deployed with a quotation of what it was referencing while omitting all but the leading and ending words to save ink and space in print (i.e. "Eur. Phoen. 1491 'στολίδα...τρυφᾶς') 27). Of the 92 references found in the range of commentary coverage, 35 references warranted inline reference resolution. We augmented entries containing these references by presenting the relevant excerpt of the text referenced closely after the reference itself. Yet resolving references did not stop there, as we found three references in Jebb's commentary recursively making their own references, and one challenge was hunting those references down and appropriately representing them for consumption. References made inside of inline reference resolution will be known as metareferences from now on, and while any trees of references made using the notes in the selected section of Jebb's commentary as roots may be a little wider than a binary tree, every recorded metareference was a leaf in its tree. Representation of metareferences in the hypertext environment will be discussed in the following subsection.
4.2.3 Hypertext Reading Environment Design
In both groups, readers viewed and interacted with their test environments as HTML pages via the Chromium browser. As seen in Fig. 3, the source text on the left pane would contain bolded words associated with notes in Jebb's commentary. When users clicked on a bolded word, every note associated with that word would appear in the commentary pane on the right. Across all versions, references with inline reference resolution implemented would be printed in dark blue text. Users interacting with Experiment 1 would see statically presented inline reference resolution, whereas users interacting with Experiment 2 would see that hovering over references revealed the text excerpt those references were actually pointing to.
Experiment 1 represented its inline reference resolution statically, where each reference had a light blue box under it containing the relevant excerpt of text the reference represented. Metareferences were shown similarly inside these light blue boxes as brighter, turquoise boxes as seen in Fig. 3. Meanwhile, Experiment 2 users only saw inline reference resolution when hovering over the dark blue references, as seen in Fig. 4. References would appear as a hover balloon, or hoverable, when hovered over. When a hoverable contained metareferences, those were bolded, and users could scroll to the bottom of the page to hover over those metareferences and read their content. Our reasoning here was that while we did not want the benefits of the commentary and its references to cause too much cognitive load, too much information can cause cognitive load while reading hypertext 39, and we wanted to explore an alternative approach to minimizing extraneous load.
Across all experiments, we wanted to track interactions that specifically revealed inline reference resolution instances to the reader. Thus, in Experiment 1 we tracked all clicks that revealed these instances, and in Experiment 2 all hovers that revealed these instances were also tracked. All left-side bolded words in Experiment 1 and all hoverables in Experiment 2 would print their unique identifiers to the console when respectively clicked on and hovered over, which we would use to extract raw data on what order and frequency student interactions occurred. Participants were strongly advised not to reload or otherwise exit the page, and a full record of interaction was captured where possible after each session.
4.3 Procedure
First, participants were given a link to the pre-experiment survey along with the number of the experimental group they were assigned to and instructed to complete the survey. When the survey was completed, they took a seat at a laptop computer with a 13" monitor open to a Chromium browser with two tabs: the reading environment and that reading environment's instructions. Fig. 5 shows one of the two instruction screens participants saw as a visual reference. Participants were urged to take a moment to read all the instructions before moving onto the second tab with the reading environment in it. Once in the reading environment, participants spent no more than 10 minutes interacting with the hypertext, at which point participants were instructed to leave the computer and complete the feedback survey. When the feedback survey was complete, so were their contributions to this study.
4.4 Surveys
Participants completed one survey before their time in their assigned reading environment, and completed a feedback survey after reading. We wanted to know how students would find their environments, but we also wanted to know more about how students approach reading in hypertext environments, especially if it is not their first one. While we had questions about how our students found their test environments and how they approach hypertext reading, some of our questions needed to be asked before experiencing the test environment, as asking afterwards could have introduced bias into the results. For these reasons, rather than having one survey, we split questions into one pre-experiment survey and one post-experiment survey. Learning progress is not always as cut and dry as whether one is in the advanced or intermediate class, so both surveys asked when they started with ancient Greek (before college, as an undergraduate, or as a postgraduate), which experiment group they were in, and which class they were participating in. All surveys were hosted on Google Forms, and all participants were anonymized.
The pre-experiment survey consisted of two extra questions in addition to the three both surveys shared, and all questions were required. First, we wanted to know what translation aids students used most, and if students in different classes used different sets of aid. Non-exclusive choices for this question included the Perseus Hopper 10, Perseus Scaife 12, other hypertext environments (see Related Work), online reference works, print reference works, internet search, generative AI services, print commentaries, published translations, and personal notes. The last question on the survey asked how willing participants would be to click on/between hyperlinked references in a commentary in a digital reading environment, where 1 was extremely unwilling and 5 was extremely willing.
Moving on to the feedback survey, in addition to the three questions on both surveys, participants had three questions about their experience asking them for a rating on a five point scale, followed by two short answer questions meant for them to offer qualitative feedback on their experience. Our rating questions were these:
How helpful on a scale from 1 (extremely unhelpful) to 5 (extremely helpful) did you find your test environment in interacting with Richard Jebb's Antigone commentary?
Given a choice between a digital reading environment with hyperlinked references and your test environment, what would be your degree of preference over the two when translating? (Choices included strongly prefer test environment over hyperlinks, slightly prefer test environment over hyperlinks, no preference, slightly prefer hyperlinks over test environment, and strongly prefer hyperlinks over test environment.)
Given a choice between the two test environments, what would be your degree of preference over the two when translating? (Choices included strongly prefer static content over hoverables, slightly prefer static content over hoverables, no preference, slightly prefer hoverables over static content, and strongly prefer hoverables over static content. Both environments were summarized in this question for the ease of participants.)
We asked the following for our two short answer questions:
What features of/changes to digital presentation of commentaries in a digital reading environment do you think would best help your own understanding of the source text?
Do you have specific feedback on your test environment?
While it may be easier for participants to say what they did or did not like about their reading experience, not every participant may be able to identify at a glance exactly what features in a digital reading environment could help them learn better. For this reason, we did not require participants to answer the first of these two short answer questions, though all other questions were mandatory.
5 Results
5.1 Surveys
5.1.1 Quantitative Analysis
Of the survey's participants, 5 started learning Ancient Greek as postgraduates, 10 started as undergraduates, and 1 started learning Ancient Greek before college. Since there was only 1 individual who learned Greek before college, they are included in cross-class and cross-experiment analyses, but not in analyses where when students learned Greek, or experience, was an axis. Also, since the postgraduate to undergraduate ratio was 1:2, and there could be underlying differences across classes, we split those with undergraduate experience into intermediate undergraduate (IU) and advanced undergraduate (AU) divisions. We calculated our 95% confidence intervals for the five point scale questions like so: while we kept the numerical scores indicated on most five point scales as is, for the two preference-based questions, we assigned each answer an integer score from $1, 5$ inclusive, where one extreme was 1 (strongly preferring hyperlinks or hoverables) and each option's score incremented by 1 until arriving at 5 for the other end of the scale (strongly preferring the test environment or static content). There was overlap in most of the 95% confidence intervals for the survey questions where either one's end overlapped with the other's beginning, or one completely absorbed the range of the other. That said, Fig. 6 shows that more than half of participants (56.3%) preferred their test environment to one with hyperlinks replacing their version of inline reference resolution, and that exactly half of participants preferring static content over hoverables. From another perspective, it was found that 95% confidence intervals for preference between hyperlinks and the test environment showed a clear division between those with postgraduate Greek experience (3.816, 5.3840) and AU students tested (2.0666, 3.5334), indicating that the former group has a strong preference for the test environment, and the latter has a weaker preference for hyperlinks. Regarding willingness to click on hyperlinked references, 7 participants selected 4, 4 students selected 3, 3 students selected 5, and 2 students selected 2. Similarly, when asked about how helpful they found their test environment, half of students selected 4, 4 students selected 2, 3 students selected 3, and 1 student selected 5.
Figure 7 shows a breakdown of what students use as translation aids, delineated by class and experience. One of the most noticeable parts of this graph is that all participants use Perseus Hopper. Everything listed was used by at least half of one class. For the intermediate class, that included Perseus Hopper, Perseus Scaife, online reference works, internet search, print commentaries, published translations, and personal notes. As for the advanced class, that included Perseus Hopper, Perseus Scaife, online reference works, other hypertext environments, generative AI services, published translations, and print reference works. It was also noticed that while half the advanced class uses generative AI services as a translation aid, only a quarter of the intermediate section did. Meanwhile, 7 of 8 students in the intermediate section reported using personal notes as a translation aid, while for the advanced section that number was more than halved. Looking at these results from the axis of experience, we found that most students with postgraduate experience also used generative AI and internet search, and that most AU students used other hypertext environments. It makes sense that students in the intermediate class would depend more on personal notes and internet searches for aid, whereas more participants in the advanced class depended more on print reference works and other hypertext environments. It resonates with the idea of adopting these field-specific tools as one gets more familiar with using them. That said, it is interesting that those learning ancient Greek as postgraduates would be more likely to use generative AI as a translation aid.
5.1.2 Qualitative Analysis
For the two short answer questions in the feedback survey, one author ran a thematic analysis process 4 on a dataset containing all answers to these questions. Each answer was treated as a data extract in the set and coded with one or more codes, which were grouped and iteratively reworked into coherent themes. The two major overarching themes were feedback on aspects specific to the hypertext reading environment and feedback on conventions originating in print commentaries. We will discuss the former first, but since there is more to be discussed regarding its subtopics of proposed modifications, proposed additions, and aspects participants liked and disliked, each subtopic will have a paragraph before covering the second major theme.
Interestingly, feedback on what participants specifically liked or disliked about their test environment differed significantly by class. On the one hand, some students from the intermediate class commented about the advantages of having commentary and source text organized close together on the same screen. One participant remarked, "I think the ability to consult commentaries and the original work simultaneously is tremendously helpful to my understanding of the text, and the smoother and more intuitive such a connection is between text and commentary, the easier I find it to comprehend otherwise less accessible texts". Meanwhile, the more advanced class had specific feedback on use of hyperlinks and hoverables in their test environments. One student wrote, "If the dictionary or the comment points to the a specific grammar point, I think to see and click said point on a hyperlink would be very helpful," and others indicated similar interest in clicking on hyperlinks to commentary references and relevant dictionary pages. This resonates with the earlier observation of all participants using Perseus Hopper as a translation aid, one such hypertext environment that already contains that functionality 10. Similarly, AU students had strong feelings about hoverables. One participant said, "I liked the hoverables... I really like the idea and think it would help me compartmentalize the information so not everything is hitting me at once". While two of the three who mentioned hoverables underscored their potential to reduce the cognitive overload anticipated when reading hypertext, the third participant indicated a preference of static content over hoverables.
Students also identified things they wanted to change in their environment. Two participants mentioned a need for translations of lines resolved via inline reference resolution, and other participants suggested different strategies for changing the design of the environment to help with this cognitive overload. For instance, one student wrote, "One big paragraph with lots of different information is difficult to consume- making it hard to find the actual note you need. For example a section on the etymology and parsing of the word, then its context in the sentence, then connections to other works etc". Another student suggested highlighting the specific words in the excerpt provided that the root-level commentary item was referencing. Though some may view inline reference resolution as doing that in this paper, others may still benefit from seeing larger excerpts of the text referenced with the referenced text emphasized.
Three types of aid present in both Perseus Hopper and Scaife were requested by students as well: translation aid, vocabulary aid, and syntactic aid. Translations of referenced lines were requested before, but one student requested aligned translation 40, a feature part of Beyond Translation 13, Scaife's more experimental version. They have this to say: "I would also appreciate if there is a more aligned english to greek translation (now Perseus only have paragraphs and since sometimes Jebb doesn't translate very literally it's hard to find the corresponding English to the Greek)". This is in line with other observations made about Jebb's tendency to contemporize and oversimplify his translations, even if it went against the constructions implied by his commentary 16. Regarding vocabulary aid, participants wanted to click on words. One participant said that "a direct link to a dictionary page with the given word and similar words I think would be most helpful for me. Much like what Perseus currently has. This would limit time spent going through hundreds of dictionary pages". Like other tendencies that deviate from the hypothesis that working with hyperlinks on different screens would be enough of a cognitive load for readers to dissuade pursuit of further insight, this incarnation links itself to Perseus through the word study tool available in Hopper 10. Yet syntactic aid provided the most interesting insight on how students view the role of translation aid. While participants indicated interest in syntactic aid that dynamically updated depending on which word was most recently clicked, one participant had an interesting perspective: "I think being able to double-check the specifics of each term (such as case, number, etc.) are especially helpful to help me better understand the grammar, but I like having it hidden because I can attempt to answer my question on my own first". Similarly, another participant remarked that having "the ability to click on certain words and see what forms they are, what their definitions are, while also having access to commentary on the words, phrases, and sentences themselves is what works best". The need for syntactic focus that users could control and toggle illuminates another mode of dealing with the cognitive overload of such a hypertext environment: participants may not always want to depend on all of the tools available for translation. One could also conclude from these observations that having a UI that lets participants use a word study tool and the commentary to understand constructions in a way that does not offload too much of the needed work in translating is one possible way to meet this need. In summary, we found that not everyone's ideal hypertext reading experience was the same. Some students suggested a variety of ways to mitigate the cognitive overload that comes with reading a selection of Sophocles' Antigone and Jebb's commentary on the matter, yet others showed comfort with the amount of cognitive load experienced in settings like Perseus Hopper 10, which was reflected in their feature requests and other feedback on this test environment.
The second major theme pertains to observations made about the commentary's composition that are based on conventions predating digital hypertext. Four participants brought up the matter of glossed words in the commentary, and two of them associated glossed words with a sense of clarity of association. One participant reported an error associated with clicking on lemmas, and another participant wanted to use glossing words to differentiate between commentary sources. In this way, glossing words emerged as another strategy for mitigating cognitive load. Seven participants called out particularly counterintuitive or confusing aspects of Jebb's commentary that originated more from commentary conventions predating hypertext than as a result of being transformed for digital use. One participant wrote, "Some of the notes just referenced other authors with no explanation as to why. What is the importance of referencing other authors in regard to understanding the language?" As stated earlier, sometimes commentarians like Jebb would reference other authors or other sources with no explanation, expecting the reader to be curious and privileged enough to hunt that reference down and understand implicitly how it related to the commentary. Another participant mentioned jargon, and more than one mentioned a dislike of commentary items with multiple thick and long paragraphs (as is, unfortunately, Jebb's wont 16). Finally, commentarians have always had some expectation, no matter how realistic or not, of who their audience is and what they know. Despite having gone over the material in class previously, a few intermediate students wrote that they did not feel that they were prepared enough, be it in regards to vocabulary, processing Jebb's commentary or synthesizing meaning from the Greek itself. In conjunction with the predilections of both classes' preferred translation aids, this supports the general implication that the advanced students generally felt more comfortable in hypertext reading environments like this one as opposed to their intermediate counterparts. Reflecting on these sets of observations in relation to conventions predating digital hypertext, one can think of these conventions as a set of implicit expectations that students centuries in the future take reasonable issue with.
5.2 Interaction Data
We ran many analyses on our log data, but only a few results showed clear divisions in 95% confidence interval testing. Log records for interaction data were converted so that consecutive interactions on the same element were counted as one interaction. This normalized our data across clicks and hovers, as hover count was higher and would challenge one's ability to try to find a division in class performance across both experiments. One distinction found between experiments lies in coverage: 95% confidence intervals for percentage of total material interacted with for those working with static content on the range of (61.2114, 89.6814) and those working with hoverables on the range of (28.1205, 53.5122).
We also noticed certain patterns in the data, so we will define interaction paths as consecutive interactions between elements steadily incrementing by 1, i.e., note 10 followed by note 11 and then by note 4 contains an interaction path of length 2. For hovering data, not every note had hoverables, but hoverables sequentially next to each other, such as note 16a and note 20a, when hovered over in increasing and consecutive order, comprised a path. Like interaction paths, adjacent regressions are a consecutive sequence of actions, yet they generally begin with a move to an earlier reference and then move forward and backward by increments of one. In Experiment 1, this is moving between notes, but the equivalent in Experiment 2 would be the reader moving forward and backward when hovering over references in the same note, and hovering over different references in another note would end this sequence. For interaction paths and adjacent regressions, we took the mean length of each metric recorded for a specific participant, as well as how many of each kind were present in that participant's records. Some significant insights in Fig. 8 emerged from analyzing this data by experiment. First, those working with static content had significantly more interaction paths in their data than those working with hoverables, though mean path length remained the same. Second, those working with hoverables had significantly longer adjacent regression periods on average than participants working with static content, though the average number of periods per session stayed about the same across all divisions. That implies that static content is somewhat more likely to be traversed linearly than the hoverables.
For the last set of tests, one author applied the non-exclusive taxonomy for commentary items proposed by 1 to the sample of Jebb's notes. We did not include metrics on the reference pointer since every item contained a reference. Conversely, we will not report on the inconsistency alert type since only one of Jebb's notes was categorized that way. Both of these decisions were made since we wanted to see if there were any types of aid provided existing as subgroups within the notes that participants interacted with most. From there, we counted each participant's total interactions, and for each experiment, we averaged the number of interactions attached to a particular commentary note. To define our most-interacted-with items, we averaged those average number of interactions across experiments and separated out all average numbers of interactions above that average. We then replicated this averaging step to determine sets of all average numbers of interactions above the average for all students in both classes for each experimental group, as well as for all students in each experimental group partitioned by experience. Next, we took the numbers of each note a set member was associated with and incremented scores in syntactic aid, semantic aid, stylistic claim, and translation for each set based on what type of aid a note was. Finally, we normalized these counts to percentages. Note that these are non-exclusive, as often Jebb's notes would be categorized with a few types. As seen in Figures 9 and 10, certain scores mark the differences. While there are some differences between behavior displayed across experiment and class axes, the greatest differences seem to happen on the experience axis. Postgraduates in both experiments were more likely than average to have semantic aid comprise exactly half of their most-clicked content, which was above the general average across experiments. Also, the intermediate class in Experiment 1 reports a syntactic aid percentage higher than that of the total average by three points, whereas the advanced class reports a high semantic aid percentage ten points above the total average.
This section ends with an insight that emerged from this taxonomic investigation. In Experiment 1, when calculating average interactions for items and handling interaction path data, we noticed that across classes, particular hotspots in the hypertext were notes 4-11 and 18-20. Fig. 11 shows each item's average score for different classes and experience groups in a heatmap, and also reveals that these hotspots are unique to those with undergrad experience. Similar to the AI insight in the previous subsection, this is another interesting distinction between those with undergrad Greek experience and those without. Also, note 2 is highly interacted with across all groups and subdivisions represented in Fig. 11.
6 Discussion and Future Work
H1 postulated that commentary audiences would prefer to see citations to and quotations from primary sources. We assumed clicking between referenced material and commentary would be too much cognitive load and that such load would block students from gaining insights derived from directly examining primary sources. Yet some participants argued that using a commentary injected with inline reference resolution resulted in more text that would generate more cognitive overload for them. Also, we expected more definition between intermediate and advanced language learners, but instead instead of verification for H2, we found copious overlap between classes and more differences between experience levels. Coverage, interaction path and adjacent regression findings imply that students may view more of these notes linearly, and that regression coincides with low coverage. Factoring in findings from 35, it may be that while all words are clicked on in the latter case, students return to review the inline reference resolution elements they find helpful. Regression, general preference for static content, and low coverage for hoverables could alternatively be interpreted as readers finding this setting confusing, but thematic analysis revealed that many students considered hoverables to help reveal more commentary insights and mitigate cognitive load. Though path and coverage data did not reveal anything special about those with undergraduate experience who worked with static content, it is also significant that a few paths' popularity stems from that demographic's behavior in a way not emulated by those with postgraduate experience. Some students also wanted to replicate elements from the Perseus Hopper and Perseus Scaife environments they were familiar with, even to the point of requesting context-switching functionality, despite a majority vote to use this test environment over the hyperlink system environments (such as Perseus Hopper). We wonder whether this is just a bias affecting classes using Perseus or other hypertext reading environments, or perhaps something more nuanced than that. Lastly, more advanced students were more comfortable with tools more specific to the field, whereas the intermediate students' frequently used aids included more general-purpose tools such as personal notes and internet search. A major takeaway from this research is that while not all divisions happen neatly along predefined demographics, there are both multiple groups that have different needs when making sense of a commentary-as-hypertext, and also multiple suggestions for mitigating cognitive overload. This dataset contained many a unique method of dealing with cognitive load, and a more accessible hypertext environment should present different modes of doing so. This analysis, along with the more detailed, in person interactions with students from two different classes, gives us the foundation for a larger study including direct interaction, via video-conference as well as face-to-face, with students from many institutions in many countries.
7 Conclusion
Our study explored how readers interact with commentaries augmented by inline reference resolution. While we initially hypothesized that such integration would reduce cognitive load by minimizing context-switching, a few participants found the additional text itself burdensome. Contrary to our expectations, distinctions between intermediate and advanced language learners were subtle, yet there was a stronger distinction between groups that had started Ancient Greek at different phases of their education. Also, user familiarity with Perseus environments may have occasionally generated a subgroup preferring hyperlink-based navigation across windows despite general support for the test environment, though more testing will determine whether this bias persists in classes not using Perseus. More advanced learners also gravitated toward field-specific tools, while intermediates favored broader resources. Students also suggested multiple forms of mitigating cognitive load, and we suggest hypertext designers ideate on different strategies that work with different reading plans.
References
Sarah Abowitz, Alison Babeu, and Gregory Crane. 2024. Bridging the Understanding Gap: Helping Readers Engage Directly with Foreign-Language Sources More Easily. In Proceedings of the 24th ACM/IEEE Joint Conference on Digital Libraries. Association for Computing Machinery, New York, NY, USA. https://doi.org/10.1145/3677389.3702539
Nicola Barbieri, Fabrizio Silvestri, and Mounia Lalmas. 2016. Improving Post-Click User Engagement on Native Ads via Survival Analysis. In Proceedings of the 25th International Conference on World Wide Web (Montréal, Québec, Canada) (WWW '16). International World Wide Web Conferences Steering Committee, Republic and Canton of Geneva, CHE, 761–770. doi:10.1145/2872427.2883092
Helen Blom, Eliane Segers, Harry Knoors, Daan Hermans, and Ludo Verhoeven. 2019. Comprehension of networked hypertexts in students with hearing or language problems. Learning and Individual Differences 73 (2019), 124–137. doi:10.1016/j.lindif.2019.05.006
Virginia Braun and Victoria Clarke. 2006. Using thematic analysis in psychology. Qualitative research in psychology 3, 2 (2006), 77–101.
Joseph Chee Chang, Amy X. Zhang, Jonathan Bragg, Andrew Head, Kyle Lo, Doug Downey, and Daniel S. Weld. 2023. CiteSee: Augmenting Citations in Scientific Papers with Persistent and Personalized Historical Context. In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems (Hamburg, Germany) (CHI '23). Association for Computing Machinery, New York, NY, USA, Article 737, 15 pages. doi:10.1145/3544548.3580847
I-Jung Chen and Jung-Chuan Yen. 2013. Hypertext annotation: Effects of presentation formats and learner proficiency on reading comprehension and vocabulary learning in foreign languages. Computers & Education 63 (2013), 416–423. doi:10.1016/j.compedu.2013.01.005
Julie Coiro and Elizabeth Dobler. 2007. Exploring the online reading comprehension strategies used by sixth-grade skilled readers to search for and locate information on the Internet. Reading Research Quarterly 42, 2 (2007), 214–257. doi:10.1598/RRQ.42.2.2 arXiv:https://ila.onlinelibrary.wiley.com/doi/pdf/10.1598/RRQ.42.2.2
Suzanne Conklin Akbari and Amanda Goodman. 2023. Introduction: Commentary at the Crossroads. In Practices of Commentary: Medieval Traditions and Transmissions, Amanda Goodman, Suzanne Conklin Akbari, and Carol Symes (Eds.). Amsterdam University Press, 1–8. doi:10.1515/9781802701593-002
Gregory Crane. 1987. From the old to the new: intergrating hypertext into traditional scholarship. In Proceedings of the ACM Conference on Hypertext (Chapel Hill, North Carolina, USA) (HYPERTEXT '87). Association for Computing Machinery, New York, NY, USA, 51–55. doi:10.1145/317426.317432
Gregory Crane. 2004-. Perseus Digital Library Project. Retrieved April 21, 2025 from http://www.perseus.tufts.edu/hopper/
Gregory Crane. 2004-. Sir Richard C. Jebb, Commentary on Sophocles: Antigone. Retrieved April 25, 2025 from https://www.perseus.tufts.edu/hopper/text?doc=Perseus:text:1999.04.0023
Gregory Crane. 2019-. Scaife Viewer. Retrieved April 21, 2025 from https://scaife.perseus.org/
Gregory Crane, Alison Babeu, Lisa M Cerrato, Amelia Parrish, Carolina Penagos, Faroosh Shamsian, James Tauber, and Jake Wegner. 2023. Beyond translation: engaging with foreign languages in a digital library. International Journal on Digital Libraries 24, 3 (2023), 163–176.
Pablo Delgado, Elisabeth Stang Lund, Ladislao Salmerón, and Ivar Bråten. 2020. To click or not to click: investigating conflict detection and sourcing in a multiple document hypertext environment. Reading and Writing 33, 8 (Oct. 2020), 2049–2072. doi:10.1007/s11145-020-10030-8
Eleanor Dickey. 2007. Ancient Greek scholarship: a guide to finding, reading, and understanding scholia, commentaries, lexica, and grammatical treatises, from their beginnings to the Byzantine period. Oxford University Press, Oxford.
Patrick J. Finglass. 2016. Jebb's Sophocles. Classical Commentaries: Explorations in a Scholarly Genre (2016), 21–38.
Center for Hellenic Studies. 2023. New Alexandria Open Commentary Platform. Retrieved July 1, 2025 from https://oc.newalexandria.info/
Christopher Francese. 2011-. Dickinson College Commentaries. Retrieved July 2, 2025 from https://dcc.dickinson.edu/about-dcc
Edward Gibbon. 1925. The history of the decline and fall of the Roman Empire / by Edward Gibbon; edited with introduction, notes, appendices [and index] by J.B. Bury: v. 1. Vol. 1. Methuen & Co.; Macmillan & Co., [1909-1925], England.
Roy K. Gibson and Christina S. Kraus. 2002. The Classical Commentary: Histories, Practices, Theory. Brill. https://brill.com/edcollbook/title/7009
Sander M. Goldberg. 2015. The Future of Antiquity: An Afterword. In Classical Commentaries: Explorations in a Scholarly Genre, Christina S. Kraus and Christopher Stray (Eds.). Oxford University Press. doi:10.1093/acprof:oso/9780199688982.003.0026
Carolin Hahnel, Dara Ramalingam, Ulf Kroehne, and Frank Goldhammer. 2023. Patterns of reading behaviour in digital hypertext environments. Journal of Computer Assisted Learning 39, 3 (2023), 737–750. doi:10.1111/jcal.12709 arXiv:https://onlinelibrary.wiley.com/doi/pdf/10.1111/jcal.12709
Peter Heslin. 2016. The dream of a universal variorum: digitizing the commentary tradition. In Classical Commentaries: Explorations in a Scholarly Genre, Christina S. Kraus and Christopher Stray (Eds.). Oxford University Press, 494–511. https://doi.org/10.1093/acprof:oso/9780199688982.003.0025
Joseph C Hilleary. 2024. Facilitating Language Hacking With Digital Tools: A Study of Translation Alignment. Master's thesis. Tufts University. https://www.proquest.com/docview/3058409134
Emma Ianni. 2021. More of a Comment than a Question: Inclusive Pedagogy and the Role of Commentaries in the Classics Classroom. Teaching Citational Practice: Critical Feminist Approaches 1 (Sept. 2021). https://journals.library.columbia.edu/index.php/citationalpractice/article/view/8654
Peter J. Anderson. 2015. Heracles' Choice: Thoughts on the Virtues of Print and Digital Commentary. In Classical Commentaries: Explorations in a Scholarly Genre, Christina S. Kraus and Christopher Stray (Eds.). Oxford University Press, 483–493. doi:10.1093/acprof:oso/9780199688982.003.0024
Richard Claverhouse Jebb. 1891. Plays and fragments: with critical notes, commentary, and translation in English prose. Part 3: the Antigone. Vol. 3. University Press; Putnam, Cambridge [Cambridgeshire]: New York. https://catalog.hathitrust.org/Record/100523964 HOLLIS number: 990015031010203941.
Yushi Jing, David Liu, Dmitry Kislyuk, Andrew Zhai, Jiajing Xu, Jeff Donahue, and Sarah Tavel. 2015. Visual Search at Pinterest. In Proceedings of the 21th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining (Sydney, NSW, Australia) (KDD '15). Association for Computing Machinery, New York, NY, USA, 1889–1898. doi:10.1145/2783258.2788621
Sarah E. Kersh and Chelsea Skalak. 2018. From Distracted to Recursive Reading: Facilitating Knowledge Transfer through Annotation Software. Digital Humanities Quarterly 12, 2 (2018). https://www.digitalhumanities.org/dhq/vol/12/2/000387/000387.html
Hanhwe Kim and Stephen C. Hirtle. 1995. Spatial metaphors and disorientation in hypertext browsing. Behaviour & Information Technology 14, 4 (1995), 239–250. doi:10.1080/01449299508914637 arXiv:https://doi.org/10.1080/01449299508914637
Christina S. Kraus and Christopher Stray. 2016. Classical commentaries: explorations in a scholarly genre. Oxford University Press. https://academic.oup.com/book/3128
Cornelis van Lit and Dirk Roorda. 2024. Neither Corpus Nor Edition: Building a Pipeline to Make Data Analysis Possible on Medieval Arabic Commentary Traditions. Journal of Cultural Analytics 9, 3 (June 2024). doi:10.22148/001c.116372
Anne Mangen, Bente R. Walgermo, and Kolbjørn Brønnick. 2013. Reading linear texts on paper versus computer screen: Effects on reading comprehension. International Journal of Educational Research 58 (2013), 61–68. doi:10.1016/j.ijer.2012.12.002
Gary Marchionini and Gregory Crane. 1994. Evaluating hypermedia and learning: methods and results from the Perseus Project. ACM Trans. Inf. Syst. 12, 1 (Jan. 1994), 5–34. doi:10.1145/174608.174609
Andrea Mazzei, Tabea Koll, Frédéric Kaplan, and Pierre Dillenbourg. 2014. Attentional processes in natural reading: the effect of margin annotations on reading behaviour and comprehension. In Proceedings of the Symposium on Eye Tracking Research and Applications (Safety Harbor, Florida) (ETRA '14). Association for Computing Machinery, New York, NY, USA, 235–238. doi:10.1145/2578153.2578195
Willard McCarty. 2002. A Network with a Thousand Entrances: Commentary in An Electronic Age. In The Classical Commentary. Brill, 359–402. doi:10.1163/9789047400943016 Section: The Classical Commentary.
Emily Norton. 2023. Beyond Hypertexting the Hypertext: Annotated and GIS Adaptations of Joyce's Ulysses as Case Studies for User Experience and Engagement. In Proceedings of the 34th ACM Conference on Hypertext and Social Media (Rome, Italy) (HT '23). Association for Computing Machinery, New York, NY, USA, Article 15, 6 pages. doi:10.1145/3603163.3609051
Terhi Nurmikko-Fuller and Paul Pickering. 2021. Reductio ad absurdum?: From Analogue Hypertext to Digital Humanities. In Proceedings of the 32nd ACM Conference on Hypertext and Social Media (Virtual Event, USA) (HT '21). Association for Computing Machinery, New York, NY, USA, 245–250. doi:10.1145/3465336.3475107
Omar Taky-eddine and Redouane Madaoui. 2024. Cognitive Overload in the Hypertext Reading Environment. International Journal of English Language Studies 6, 2 (May 2024), 94–100. doi:10.32996/ijels.2024.6.2.13
Chiara Palladino, Maryam Foradi, and Tariq Yousef. 2021. Translation alignment for historical language learning: a case study. Digital Humanities Quarterly 15, 3 (2021).
Tiziano Piccardi, Miriam Redi, Giovanni Colavizza, and Robert West. 2020. Quantifying Engagement with Citations on Wikipedia. In Proceedings of The Web Conference 2020 (Taipei, Taiwan) (WWW '20). Association for Computing Machinery, New York, NY, USA, 2365–2376. doi:10.1145/3366423.3380300
Matteo Romanello. 2020-. Ajax Multi-Commentary Project. Retrieved August 2, 2024 from https://mromanello.github.io/ajax-multi-commentary/
M. Sanchiz, F. Amadieu, J. Lemarié, and A. Tricot. 2022. Do graphic and textual interactive content organizers have the same impact on hypertext processing and learning outcome? Journal of Computing in Higher Education 35 (June 2022). doi:10.1007/s12528-022-09328-z
Teresa Schurer, Bertram Opitz, and Torsten Schubert. 2023. Mind wandering during hypertext reading: The impact of hyperlink structure on reading comprehension and attention. Acta Psychologica 233 (2023), 103836. doi:10.1016/j.actpsy.2023.103836
Philipp Singer, Florian Lemmerich, Robert West, Leila Zia, Ellery Wulczyn, Markus Strohmaier, and Jure Leskovec. 2017. Why We Read Wikipedia. In Proceedings of the 26th International Conference on World Wide Web (Perth, Australia) (WWW '17). International World Wide Web Conferences Steering Committee, Republic and Canton of Geneva, CHE, 1591–1600. doi:10.1145/3038912.3052716
Yang Song, Xiaolin Shi, and Xin Fu. 2013. Evaluating and predicting user engagement change with degraded search relevance. In Proceedings of the 22nd International Conference on World Wide Web (Rio de Janeiro, Brazil) (WWW '13). Association for Computing Machinery, New York, NY, USA, 1213–1224. doi:10.1145/2488388.2488494
Roger Travis. 2019. Using Annotations in Google Docs to Foster Authentic Classics Learning. In Teaching Classics With Technology. Bloomsbury Publishing Plc, 207–216. doi:10.5040/9781350086289.ch-017
Zsofia Vörös, Jean-François Rouet, and Csaba Pléh. 2011. Effect of high-level content organizers on hypertext learning. Computers in Human Behavior 27, 5 (2011), 2047–2055. doi:10.1016/j.chb.2011.04.005 2009 Fifth International Conference on Intelligent Computing.
Do you like what you are reading? Subscribe to receive updates.
Unsubscribe anytime