PubLink: Editorial Workflow for Digital Scholarly Publications in the Humanities
Authors: Elisa Bastianello, Christopher Tomlinson, Alessandro Adamou
Source attribution: Publisher HTML for DOI: 10.1145/3648188.3677051, cross-checked against the supplied ACM PDF.
Abstract
This paper presents a customized and sustainable workflow for digital scholarly publications in the humanities that aims to overcome the difficulties of transitioning from traditional fixed formats to those that require fluid layout and structure, such as ePub, (X)HTML or XML. While the process fits well within the humanities tradition and takes into account the specificities of the humanities that distinguish them from the hard sciences, such as extended footnotes and non-DOI references, it can be easily adapted to publications in different fields. The paper outlines the various challenges encountered in the digital publication process, including the transformation of text documents into JATS XML and the process of uploading to OJS (Open Journal System). It describes our attempts to streamline the steps and provide solutions that fit within the capacity of limited budgets without compromising the content and structure of the publications. The paper concludes with an evaluation of the proposed workflow and outlines future improvements to increase its efficiency.
CCS Concepts
• Information systems → Digital libraries and archives ; • Software and its engineering~Reusability; • Applied computing → Extensible Markup Language (XML) ; • Applied computing → Hypertext languages ; Annotation; • Applied computing → Hypertext / hypermedia creation ;
Keywords
Digital publishing , Digital Humanities , OJS , JATS XML , Sustainable Workflow
ACM Reference Format
Elisa Bastianello, Christopher Tomlinson, and Alessandro Adamou. 2024. PubLink: Editorial Workflow for Digital Scholarly Publications in the Humanities. In 35th ACM Conference on Hypertext and Social Media (HT '24), September 10--13, 2024, Poznan, Poland. ACM, New York, NY, USA 5 Pages. https://doi.org/10.1145/3648188.3677051
1 INTRODUCTION
Digital scholarly publications in the humanities are finally catching on not only as fixed-layout PDFs but also as flexible hypertexts. The latter are available either in downloadable formats, such as ePub, or as online content, directly accessible in HTML, either native or generated on the fly from XML transformations. The editorial tradition in the humanities typically involves small, content-specialized teams. These teams rely on publishers primarily for the layout, printing, and distribution of printed volumes. For many years, this combination favored the move to digital in the form of PDFs generated by the same pipeline that produces print layouts. However, transitioning to formats that require fluid layout and structure demands content transformations that extend beyond traditional expertise.
Humanities publishing, which often requires the use of complex bibliographic references and footnotes, as well as an integrated iconographic apparatus in some fields, does not easily lend itself to editorial practices of the hard sciences. Humanities journals, despite their centuries-old tradition, often can't compete with the budgets of the large scientific publishers that centralize the majority of hard science publications. In order to ensure Diamond/Platinum Open Access publications for authors, who traditionally do not expect APC costs in the case of print articles, it is crucial to find solutions that can be accommodated within limited budgets. This paper presents a sustainable process consistent with the humanistic tradition, aimed at simplifying the publication of articles and papers in JATS XML format, without neglecting more traditional outputs like PDF.
2 State of the Art
In the humanistic academic tradition, authors typically draft articles using word processing software, such as Microsoft Word (.docx) or OpenOffice Write (.odt). Nevertheless, the he use of paragraph styles to structure the content and reference managers to format the bibliographic citations (such as Zotero,1) is not yet prevalent. This is due in part to the intricate nature of traditional editorial conventions, which render the optimization of reference managers a challenging endeavor. Consequently, manual formatting remains a preferred approach. Additionally, the majority of editorial guidelines are designed for visual formatting. Even when document templates are provided, they are frequently ineffective. However, conversion to HTML or XML, especially, JATS XML,2 necessitates extensive document cleanup. Consequently, solutions that require advanced technical expertise have emerged, including manual encoding with specialized software like Oxygen XML Editor, and more general-purpose options, such as Visual Studio Code.
Most digital journals are hosted on the open-source platform OJS (Open Journal System),3[2, 3, 6] which is designed to oversee the entire editorial lifespan of a scientific publication, from article submission to peer review to online publication. However, our workflow entails the completion of content submission, editorial review, and peer review outside the platform. Subsequently, only presentation galleys are uploaded while manually entering the metadata. This process is lengthy and tedious, primarily due to the prolonged loading times of individual files. Additionally, it is susceptible to errors due to the repetition of metadata submission for different articles within a limited time frame.
3 SPECIFIC REQUIREMENTS
When our institute made the decision to invest in digital scientific article publication in 2019, our objective was to ensure that online articles would be saved in XML or XHTML format. This decision was made in order to optimize long-term preservation characteristics and enhance the text's richness with semantic annotations. However, the cessation of development of ScienceMatters,4Internet Archive our initial intended publication platform, led us to survey existing platforms for alternatives.
The transition from text document to JATS XML also presented challenges, as the few existing editors (accessible at a budget) that we tested5 were only compatible with texts that had minimal to no footnotes and references with DOIs. This was due to the fact that they relied on a Crossref/DOI connection to insert the data. In the humanities, it is common practice to cite sources that were produced prior to the introduction of DOIs or other identifying codes, rendering these solutions highly ineffective. Furthermore, there was a need to provide guidance to authors on how to create documents that adhere to templates with correctly applied styles, preferably on a shareable interface.
Subsequently, further issues arose when the platform (OJS) was selected and the publication process was subjected to testing. One challenge pertains to the publication of articles in JATS XML format to OJS for online viewing. The platform presents the article metadata and abstract in a landing page that allows the user to download the PDF or to read it online as an HTML galley. The primary JATS publishing plugin on the platform, eLife Lens-Viewer,6 is integrated solely in version 1, which lacks the functionality to utilize notes. Furthermore, the plugin's development had already been at a standstill for years,7 suggesting that further efforts would be futile. The integration of obsolete components into the publishing process was deemed an unsustainable approach.
Once production commenced, however, it became evident, as previously stated, that the primary bottleneck was the uploading process to OJS. Despite the availability of a dedicated plugin, the QuickSubmit8 feature, which enables direct uploading of articles by editors for immediate online publication, proved to be a significant challenge. The plugin, in fact, requires that the editor manually enter all metadata pertaining to the article, the authors, and all publication files and associated dependencies (especially images) directly into the system form. This process is particularly susceptible to errors and is exceedingly time-consuming and tedious, particularly when the article in question is complex.
4 PROGRESS SO FAR
The following text presents a comprehensive account of the process that has led to the conceptualization of PubLink, a tool designed to facilitate the publishing process. Furthermore, it delineates the features that have been developed thus far and provides insights into the prospective development plans for the tool.
4.1 Choice of Publishing Platform
In the course of evaluating potential publishing platforms, we ascertained that a considerable number of open access platforms utilized by the humanities are based on the open-source platform OJS,[2, 3, 5, 6] developed by PKP.9Although the platform can be installed independently, we elected to rely on an experienced manager to mitigate the initial development and customization burden. In our case, the platform is hosted and maintained by Ubiquity Press,10 which has also developed its own XML parsing tool to enable readers to access articles online on the OJS interface. At the same time, in the event of a change of platform manager or a switch to an internally controlled server, the content can be moved to another instance of OJS through import and export using the native plugin, with only minimal changes.
4.2 Word Processing Software
At the time, there were no accessible word processors that could directly export in JATS format and simplify the structural formatting of the content via templates. Among the developing projects that were tested, the SciFlow Authoring platform11 had many of the desired features, despite the lack of a native JATS XML export. The software demonstrated considerable potential, particularly in regard to its straightforward management of document templates and annotation of bibliographic references. The latter is achievable both through the integrated Zotero library connection and from imported BibTeX files,12 which can be consistently updated. The formatting of the bibliography is achieved through the utilization of CSL (Citation Style Language),13 which allows the editor to modify the style during the editorial process via the selection of an alternative citation style. As a program designed specifically for scientific writing, Sciflow allows users to select structured content templates and provides guidance on the appropriate usage of paragraph styles. Furthermore, the software allows for the creation of internal references to other chapters, images, and tables in an intuitive manner, resulting in a well-organized document that is automatically populated with authors’ metadata. Consequently, starting from 2020, we elected to provide financial support for the advancement of the JATS format export and the incorporation of the article's publication metadata into the editor. This resulted in a virtuous circle, in which other institutes similarly opted to integrate SciFlow into their publication workflows. This collaborative effort culminated in the creation of the SciFlow Publishing platform, in addition to the preexisting SciFlow Authoring platform, which is now an integral part of our editorial workflow.
4.3 Handling Bibliographic References
One of the more challenging aspects of the process is the management and extraction of bibliographic references directly from the articles where they are included in the form of unstructured formatted bibliographies. Among the numerous reference managers on the market, Citavi14 stands out due to its import functionality, which enables the loading of formatted bibliographies and their comparison with selected public catalogues. This feature facilitates the processes of content validation and data transformation.15 It is important to note, however, that Citavi also has certain limitations. The data conversion is conducted using a proprietary format that does not employ CSL, which is open-source. Moreover, the bibliographic data is stored in a proprietary format that is not readily convertible, similar to the proprietary and encrypted fields utilized within the text documents annotated via this reference manager. Consequently, we elected to utilize a Zotero Lab16 institutional license to oversee the libraries comprising the bibliographic references of our digital publications and to harmonize duplicates across articles. To ensure the stability of the identifiers, we employed the BetterBibTeX plugin to generate stable citation keys. 17
In addition, we have developed and published a bespoke CSL that is aligned with our requirements.18 It can be used directly by authors to generate bibliographic citations and bibliographies and is integrated into our customized SciFlow template for our articles.
4.4 Collaboration with Authors
A crucial aspect of optimizing the workflow is the collaboration with the authors. The editorial guidelines include templates for both MS Word and OO Write, which reflect the document structure required to facilitate publication. Additionally, authors are advised to utilise SciFlow from the initial drafting stage onwards. We also request a bibliographic database in a standard format, such as BibTeX or RDF, with the article references pre-inserted, thus supplementing the formatted bibliography. This serves as the basis for the editorial stage. At this time, our digital editorial program does not include open-call journals, but rather volumes of curated essay collections. Subsequently, the document is submitted via email or by sharing the SciFlow document with the editorial team. In the course of the review phase, which is conducted by the scientific committee and the peer review, a variety of strategies have been evaluated, depending on the level of curatorial engagement in managing SciFlow documents. In certain instances, reviewers were invited to provide direct commentary on online files. In other cases, papers were downloaded from SciFlow in.docx format for peer review and copyediting. Currently, the revision control, annotation, and commenting functions are not yet sufficiently developed in SciFlow to allow it to be used exclusively at a professional editorial level. Following the processing of documents and bibliographies, they can be re-imported into SciFlow, via the Import console, which is available for use with the Publishing plan. The accepted formats include.docx and LaTeX documents, with Zotero integrated references, images, captions, tables mainly preserved. During the import phase, an improved and rectified BibTeX file can also be imported, and the updated bibliographic references are annotated by simply dragging them from the entry list into the text. The pre-processed document is subsequently imported as a new SciFlow document, where author and article metadata can be added.
4.5 Bibliography Improvement
Given the potential for inconsistencies and typographical errors in bibliographies provided by authors, we sought tools for identifying and formatting bibliographies for the first volume's articles of our series "Hertziana Studies in Art History".19 These articles, initially intended for conventional print publication, featured a bibliography in note form that required recognition and extraction. In an attempt to automate the process, a procedure was devised that aimed to identify references within the notes and reconcile the extracted content with data from WorldCat and other sources, such as catalogues of specialised libraries (e.g., Kubikat for art and architectural history documents). [1]. The process’ instability, exacerbated by difficulties accessing catalog open data for comparative and normalization purposes, negatively impacted the workflow. Furthermore, citations subsequent to the initial one, presented in abbreviated form, were not marked or identified. In this specific instance, the construction of the bibliographic database was conducted manually during the editorial phase. The development of the tool is currently on hold,in part due to the fact that the articles submitted for the subsequent volumes had to adhere to the editorial standards for final formatted bibliography with author-date citation keys. This potentially facilitates the identification of references with existing tools such as Grobid,20 even in the absence of a bibliographic database.
4.6 JATS File Manipulation
JATS file export from SciFlow, particularly element-citation format that contains structured bibliographic data, features discrepancies compared to our final document versions. These include absent label with abbreviated reference citations, disorganized bibliography, incorrect document title markings, missing start and end pages for articles, exclusion of source curators, missing URIs, and lack of original date information for reprints.The manual correction of these issues is both time-consuming and susceptible to errors, particularly given the necessity of repeated copying and pasting of code snippets and manual tag insertion. Hence, with the aid of conversational AI, we've developed XSLT and Python scripts that facilitate the transformation of SciFlow's JATS file, rectifying the bibliography and integrating absent fields with the content of the BibTeX file. This ongoing refinement process hinges on the consistency of Zotero/BibTex identifiers (bar the surname-year separating dot in the BibTex version), enabling reconciliation with article content. Minor adjustments include image encoding, involving a local alternative insertion for archiving alongside the main image published from our IIIF server, and elements of the front section.21 These remain the sole manual changes performed using text editors like Visual Studio Code and Oxygen XML Editor.
A final method of manipulation of JATS XML files involves the creation of PDF files that are loaded as an alternative galley. In our case, given that we are dealing with articles that often have a preponderant iconographic content, we currently prefer to generate the PDFs in InDesign rather than as a direct transformation. InDesign allows for the loading of XML files, but specific formatting of the content is necessary to guarantee that the structure aligns with the template styles. Additionally, since the element citation field does not incorporate formatting of bibliography entries, it must be reconstructed from BibTex and CSL. Consequently, we have expanded the processing scripts to facilitate the conversion and limit the handling of complex elements to the typesetter, while still achieving a professional layout.
4.7 File Manipulation for Publication: the PubLink platform
Given that many of the procedures mentioned previously involve using existing scripts and correcting the source code, we have decided to introduce an online interface that will act as the main workbench for all the tools, which we named PubLink. This interface allows for the application of scripts, either individually or in combination, via an online graphical interface. It comes equipped with a file manager that facilitates uploading all the elements comprising the final publication, optimizing and annotating them.
At present, the demo version is online and is primarily utilized for bibliography oversight and document preparation for OJS import.22 To manage the bibliography and extract article-related metadata for populating the OJS form, an intermediate Serial Object format has been developed to separate the JATS file elements for subsequent functionalities.
4.7.1 Bibliography Oversight. In this phase, the bibliography can be compared with either a BibTeX file or directly in the JATS XML file with the CrossRef API. Although only a portion of the references is natively assigned a Digital Object Identifier (DOI), it is becoming increasingly common for publishers to provide digital reprints of texts with DOIs. Consequently, it is possible to compare the elements with DOIs with the registered metadata, thereby identifying instances of digital reprints or missing DOIs. These can then be incorporated directly into the JATS file or exported. Furthermore, the tool is capable of analysing the JATS document in order to prepare a list of relatedIdentifiers that can be inserted in a DataCite DOI XML, thereby enhancing the metadata provided by the OJS DataCite plugin.
4.7.2 Import File Creation. As previously mentioned, importing the JATS file as a galley in OJS using the QuickSubmit function is a lengthy and tedious process. It requires both metadata entry and the need to individually upload manuscript files and all dependent elements like local images with redundant procedures. However, there is the possibility of importing complete articles with metadata and galleys from another OJS installation by using a specific “Native XML” file of the platform. This file incorporates all attachments inside as embed elements in base 64 format. Relying on existing projects, which use for example an Excel sheet,23 we have developed a transformation that can insert any missing metadata directly from the interface when starting from the intermediate Serial Object, and can add additional galleys like PDF files and cover images. Everything can be combined into a Native XML for OJS 3.3 (LTS). Upon generation, the file can be uploaded swiftly to the OJS publication platform and, after minimal checking, be published. Direct import currently creates a published file without reference issue (if not specified in the additional metadata), hence it must be unpublished to enter the correct data, a process which only takes a few minutes.
4.8 Forthcoming Developments
We are currently developing the integration of IIIF image annotations into articles and transforming them into XML format for import into InDesign, where PDFs are generated. Although direct PDF conversion is feasible, the humanities field frequently requires a more intricate presentation format, hence this additional step is necessary; however, this reflects a unique choice by our institute. SciFlow itself is capable of exporting the final document in PDF format with a correct pagination adhering to a custom template, making this additional phase avoidable. Additionally, if the OJS platform has not incorporated an XML to HTML parsing system, it is also possible to directly export an HTML ready for publication. If our editorial workflow would benefit from this functionality, it will be integrated into PubLink.
4.9 Conclusions
Our workflow, while crafted specifically for our needs, can be easily adapted for other institutes. Upon the project's completion, the source code for the PubLink platform will be made open-source,24 and its implementation through Docker containers enhances its reusability potential for other institutes. Numerous resources are available for editorial workflows, and an interface simplifying access could make digital publication as JATS XML accessible for editorial teams with lesser IT capabilities.
ACKNOWLEDGMENTS
The development of the PubLink platform was made possible by a grant from the Deutsche Forschungsgemeinschaft (DFG) - Project number 501142032.25
REFERENCES
Elisa Bastianello, Alessandro Adamou, and Nikos Minadakis. 2022. Referency: Harmonizing Citations in Transdisciplinary Scholarly Literature. In Linking Theory and Practice of Digital Libraries - 26th International Conference on Theory and Practice of Digital Libraries, TPDL 2022, Padua, Italy, September 20-23, 2022, Proceedings(Lecture Notes in Computer Science, Vol. 13541), Gianmaria Silvello, Óscar Corcho, Paolo Manghi, Giorgio Maria Di Nunzio, Koraljka Golub, Nicola Ferro, and Antonella Poggi (Eds.). Springer International Publishing, Cham, 528–532. https://doi.org/10.1007/978-3-031-16802-4_58
Brian D. Edgar and John Willinsky. 2010. A Survey of Scholarly Journals Using Open Journal Systems. Scholarly and Research Communication 1, 2 (June 2010), 22 pages. https://doi.org/10.22230/src.2010v1n2a24
Saurabh Khanna, Jon Ball, Juan Pablo Alperin, and John Willinsky. 2022. Recalibrating the Scope of Scholarly Publishing: A Modest Step in a Vast Decolonization Process. Quantitative Science Studies 3, 4 (Dec. 2022), 912–930. https://doi.org/10.1162/qss_a_00228
Jiří Kratochvíl. 2017. Comparison of the Accuracy of Bibliographical References Generated for Medical Citation Styles by EndNote, Mendeley, RefWorks and Zotero. The Journal of Academic Librarianship 43, 1 (Jan. 2017), 57–66. https://doi.org/10.1016/j.acalib.2016.09.001
Miriam Wanjiku Ndungu. 2020. Publishing with Open Journal Systems (OJS): A Librarian's Perspective. Serials Review 46, 1 (Jan. 2020), 21–25. https://doi.org/10.1080/00987913.2020.1732717
John Willinsky. 2005. Open Journal Systems: An Example of Open Source Software for Journal Management and Publishing. Library Hi Tech 23, 4 (Jan. 2005), 504–519. https://doi.org/10.1108/07378830510636300
FOOTNOTE
⁎Corresponding author.
4The original domain is no longer associated with the platform, but archived copies are available at .
5We tested Manuscripts by Atypon (now defunct https://www.manuscripts.io/about/ and XeditPro https://xeditpro.com/.
7The project main development page is available at https://github.com/elifesciences/lens, where only minor commits were carried out in the last four years.
11 https://sciflow.net/en/. Another tool we tested was the aforementioned Manuscripts, that could supported JATS but could not handle non-DOI references updates and was discarded.
12BibTex files can be exported from all the common reference managers, such as Clarivate EndNote, Elsevier Mendeley or Citavi (Now Lumivero Citavi).
18The CSL, which is currently available for English language papers only, can be accessed from the Zotero Style Repository as bibliotheca-hertziana-max-planck-institute-for-art-history or directly from the development git https://github.com/biblhertz/biblhertz_csl.
23For instance https://github.com/ualbertalib/ojsxml.

This work is licensed under a Creative Commons Attribution International 4.0 License.
HT '24, September 10–13, 2024, Poznan, Poland
© 2024 Copyright held by the owner/author(s).
ACM ISBN 979-8-4007-0595-3/24/09.
Do you like what you are reading? Subscribe to receive updates.
Unsubscribe anytime