Bringing the Web of Data to the Small Web: a Multi-Protocol Contract and the Chaykin Linked Data Server
Alessandro Adamou
Published in: HT ’26 · DOI: 10.1145/3800935.3830868 · License: CC BY 4.0
Keywords: Small Web, Non-HTTP Web, Linked Data, Gemini protocol, Spartan protocol, Nex protocol
Abstract
In response to a perceived systemic information crisis being experienced by the contemporary HTTP-based Web, a growing informal counter-movement — the Small Web, or smolweb — advocates for minimalist, content-centric protocols such as Gemini, Spartan, and Nex as deliberate alternatives to the dispersiveness of the modern Web. This paper argues that the Small Web represents a fertile and largely unexplored environment for reintroducing Linked Data and the Semantic Web natively, with the potential to circumvent the architectural flaws that perpetuate the prevailing information crisis. To this end, we derive a core contract establishing the minimal conventions a Linked Data system must observe to operate over smolweb transports, alongside additional capabilities exposed by individual protocol bindings. We further introduce Gemtext-LD, a formally specified serialisation of RDF triples into Gemtext — the sole document format common to all the protocols considered. Finally, we present a reference implementation comprising a Gemtext-LD library and Chaykin, a multi-protocol Linked Data server and proxy. This work constitutes a first step toward a Small Web of Data, and opens research directions including SPARQL integration and a principled comparison with the Linked Data Platform.
CCS Concepts: • Networks → Application layer protocols ; Public Internet; • Information systems~Data encoding and canonicalization; • Information systems → Hypertext languages ; • Information systems → Resource Description Framework (RDF) ; Entity resolution; Database web servers; • Networks → Application layer protocols ; • Networks~Public Internet; • Information systems → Data encoding and canonicalization ; • Information systems → Semantic web description languages ; • Software and its engineering → Hypertext languages ;
Full text
1 Introduction
It is widely accepted that the HTTP-based Web has undergone a comprehensive transition from a Web of Information to its current form, traversing phases that included participation and the Web of Data [10]. Upon closer examination, however, it becomes evident to the scholar focusing on any of these stages, that the disparity in community engagement with them is causing other stages to be hampered by the very architecture of the Web it sought to enhance. The HTTP-based Web is experiencing a profound “information crisis” at the expense of some of the most forward-thinking visions of it, such as the Semantic Web, which however has done little to transcend its status as a research domain.
This crisis is as much about the volume of information as it is about its navigability. The dominant Web architecture, optimised for content distribution and engagement, amplifies these problems by prioritising virality over veracity. Systemic amplifiers of misinformation, incentives to the attention economy, filter bubbles and echo chambers, and the lack of data provenance collectively impede a constructive consumption of information [15, 19]. These are not incidental defects but structural consequences of an architecture built on algorithmic ranking, third-party tracking, and engagement-optimised presentation: content is surfaced by its capacity to capture attention rather than by its veracity, and provenance is routinely severed by aggregation, embedding, and re-posting layers that this architecture rewards rather than merely tolerates. A protocol stack that structurally forecloses algorithmic feeds, client-side scripting, and third-party requests removes several of these mechanisms by design rather than by policy: there is no ranking layer left to optimise for engagement, no tracking substrate for behavioural advertising, and no visual layer through which attention-grabbing presentation could be rewarded over content itself. HTTP is technically capable of serving minimal, tracking-free, script-free content: the distinction that matters is between what a protocol permits by omission and what it prevents by design. HTTP's extensibility means that minimalism is always a voluntary, reversible convention, and the same commercial incentives already described above make it likely that, eventually, it will indeed be reversed. By contrast, minimalism by protocol-level constraint, rather than by editorial policy, mandates no scripting facility, no cookie mechanism, and no arbitrary header space in which tracking metadata could travel. This distinction is not a nostalgic preference for austerity, but a structural precondition for the trust and provenance properties we are ultimately concerned with: a client can only rely on a looked-up URI yielding unmanipulated information if the transport carrying it cannot itself be silently augmented, after the fact, with the very mechanisms responsible for the crisis described above.
The problem perceived whilst carrying out this work is therefore that HTTP cannot guarantee that it will keep serving minimal content–something only a rigidly inextensible protocol can offer, and precisely the property for which the need for new protocols, rather than disciplined old ones, is felt. This is why a counter-movement gaining momentum in response to these very symptoms takes the specific shape it does: the “Small Web”, or “smolweb”. This informal movement champions minimalist protocols designed for speed, simplicity, inextensibility, and a focus on content over presentation [5]. Newly-conceived protocols like Gemini, drawing from those of the primordial, pre-graphical Web like Gopher [1]–whose site count is also on the rise1–and Finger [20] represent a deliberate rejection of the bloat and complexity of the modern Web, harking back to vernacular, simpler eras of online interaction and re-appropriating traditional means of exploring the Web [6].
These protocols are largely the work of enthusiasts and members of specialised communities, such as computing history, IndieWeb, retrogaming, and others that have witnessed the nascent Web [2]. These individuals profoundly understand the merits of the Web of Information, yet their exposure to direct experience with, or tangible benefits of, the advancements in the Web of Data remains limited. Lacking this perspective, there is a concrete risk of their efforts, and the Small Web phenomenon in general, being perceived as a by-product of the nostalgia effect on a community of hobbyists. This raises a more precise question worth confronting: does Geminispace–the best-documented of these ecosystems–amount to a movement, or merely a self-referential community sharing an aesthetic preference? Scale alone does not resolve this, but it is not indifferent to it either: Geminispace has grown roughly tenfold in five years [17], confirming a sustained multi-year trajectory rather than a static hobbyist plateau. What distinguishes a movement from a community more decisively, however, is a shared normative stance rather than a shared taste: independent first-hand accounts of the scene converge on an explicit rejection of algorithmic curation, behavioural tracking, and engagement-optimised design as a matter of principle: not merely a preference for older aesthetics, but for independence on presentation [5]. The architectural relevance of the Small Web to a renewed Web of Data does not hinge on the size or self-organisation of its current user base. Yet the Web of Data, as realised in the concepts of Linked Data and the Semantic Web, represents one of the most significant advancements to emerge from the HTTP-based Web since it was merely publishing documents, not least because it stems from the forefathers of the Web itself [11]. As the proponents of Small Web protocols advocate for its co-existence with the HTTP one, this offers an opportunity to move beyond a mere return to tradition, by leveraging the positive aspects of the evolution of the Web to create a leaner, information-focused space [13]. This paper posits that the Small Web provides a conducive environment to re-affirm the Linked Data principles, with the potential to circumvent the architectural flaws that contribute to the prevailing crisis. This shall not prescind from acknowledging the existence of Linked Data in the HTTP realm, nor from allowing them to coexist across spaces.
This work initiates the process of reclaiming information integrity through a data-aware Web, by establishing the fundamental conventions that a Linked Data system should observe in order to make its wealth available to such a minimalist, multi-protocol, document-centric underground. We therefore use the most recent Small Web protocols for reading content: Gemini, Spartan, and Nex, as well as their sister protocols, Titan and NPS, for sending them. The reduced complexity of these protocols opens the floor to implementing Linked Data support to some degree, and the decentralised nature of the Small Web fosters a more resilient information ecosystem. There are, however, caveats that pose challenges not only to the creation of a Small Web of Data, but to the re-validation of Linked Data principles themselves once extrapolated from their original context. Addressing them with awareness is within the scope of the present work.
This paper presents a manifold contribution. Firstly, it extracts a lowest common denominator from the Small Web protocols and attempts to set a core contract on top of it for publishing Linked Data natively over them, proxying HTTP ones, and connecting the two. Then, it proposes additional features made feasible by each protocol binding. Because many of these protocols share a common document format, called Gemtext, a serialization of RDF triples into this format is introduced and formally documented. Lastly, we make a reference implementation available, in the form of a library for handling data in Gemtext, and a Linked Data server and proxy application for all these protocols, Chaykin.
1.1 New Fertile Ground for Renewing the Unfulfilled Promise of the Semantic Web
The seminal literature on Linked Data sets the following founding principles [4]:
1. Use URIs as names for things.\
2. Use HTTP URIs so that people can look up those names.\
3. When someone looks up a URI, provide useful information.\
4. Include links to other URIs, so that they can discover more things.\
In the multi-protocol context at hand, where the second principle no longer holds, a re-formulation of it could be to "Use Web URIs [...]", intending "Web" in a broad sense. We then want to ascertain that, in such an environment, the other principles still hold true. A core use case that shall be adopted to that end is the one of serving Linked Data as sites on the Small Web–which are sometimes called capsules or stations–either natively hosted or retrieved by proxy from the HTTP-based Web. This will not be achieved by extending the Small Web protocols: their characteristics, used to extract an irreducible core contract from their lowest common denominator, are regarded not as constraints, but as capabilities superimposed on said contract. Consequently, whereas the HTTP-dependent Linked Data Platform [12] is acknowledged, it cannot be taken as reference, in contrast to other approaches in the Internet of Things.
The remainder of this paper is structured as follows: Section 2 provides an overview both on the landscape of current-generation Small Web protocols and on prior approaches to minimalist management of Linked Data. Section 3 summarises the specification, covering the salient characteristics of both the core contract and wire format. Sections 4 and 5 delve into the specifics of the format and transport protocols at hand, respectively. Section 6 illustrates the reference implementation, which is validated in Section 7 against the specification and available client tools. Section 8 revisits the stated objectives and discusses them in terms of their implementation, whilst setting out medium-term prospects.
2 Background
This survey of the state of the art concisely illustrates the protocols and document format under consideration, before a glimpse into what comparable literature exists for the present work.
2.1 Contemporary Small Web protocols
Although none of these protocols have been submitted to the IETF for standardisation, they are each at various stages of maturity and enjoy diverse adoption. The present work aims at treating them equally, regardless of maturity.
2.1.1 Gemini. Gemini2 was introduced in 2019 as an evolution of traditional Gopher, with modernisations such as mandating TLS-encrypted connections and natively using a hypertextual content type, Gemtext (§2.2). Although Gemini is read-only and rigidly inextensible by design, sister protocol Titan3, which should listen to the same TCP port as Gemini (1965 by default), was designed for uploading data. In Gemini jargon, a site is called capsule, and the set of all known capsules is known as Geminispace. As of March 2026, Geminispace is known to comprise about 4850 capsules.4
2.1.2 Spartan. Admittedly a hobbyist-oriented protocol inspired by Gopher and Gemini, Spartan5 strives for extreme simplicity, opting for unencrypted TCP, reducing its status codes to 4, and dropping request parameters. It also accepts Gemtext as content type but, unlike Gemini, it supports data upload natively. Spartan uses TCP port 300.
2.1.3 Nex. Nightfall Express, or Nex,6 is likely the least developed of these protocols, yet it still enjoys third-party support. It is a minimalist, stateless protocol whose request payload (on port 1900) consists of an optional path, and that has no response headers. It supports a line-oriented plaintext content type with hyperlinks. Nex is read-only, but can be paired with sister protocol NPS7 for uploading data to port 1950. A Nex online resource is called a station.
The focus of this work is on the protocols that have seen the most extensive development and adoption with third-party client and/or server implementations (§7.2). A plethora of further internet protocols that can be likened to the Small Web are being proposed, including Mercury, Scorpion, Guppy, and Molerat.8 For any of these to be considered in the future, even in the event that they should challenge our current core contract and wire format, they will require substantial adoption beyond the reference implementation offered by the proponents of the protocols themselves.
Table 1 consolidates what distinguishes the five protocols from one another, and, not coincidentally, what distinguishes all of them from HTTP as it is exploited by the Linked Data Platform: no verb space, no general header mechanism, and no content negotiation, the three pillars LDP's semantics are built on (§2.3).
Table 1: Feature comparison of the five Small Web protocols, focused on properties that bear on the feasibility of an LDP-like binding. Titan and NPS are mostly wire-identical to their read counterpart, differing only in the operation admitted.
Feature | Gemini | Titan | Spartan | Nex | NPS |
|---|---|---|---|---|---|
Default ports | 1965 | as Gemini | 300 | 1900 | 1915 |
Operation(s) admitted | read | submit | read + submit | read | submit |
Verb/method space (as in HTTP) | none | none | none | none | none |
Operation(s) admitted | read | submit | read + submit | read | submit |
Response metadata beyond status | single meta string | as Gemini | single meta string | none | as Nex |
Status codes | two-digit | as Gemini | single-digit | none | as Nex |
Transport encryption | TLS 1.2+ | as Gemini | none | none | as Nex |
Default document format | Gemtext (native) | as Gemini | Gemtext | plaintext, =>-links only | as Nex |
Content negotiation | none | as Gemini | none | none | as Nex |
2.2 Gemtext format
Gemtext (unofficial MIME type text/gemini; filename extension .gmi) is the default information interchange format for Gemini, also supported by Spartan. It is a lightweight markup format with support for hyperlinks, headings, preformatted text, bulleted lists, and blockquotes.9
A key characteristic of Gemtext is that lines are strictly delimited, therefore a single line can have one purpose and one only; this implies, for example, that a list item cannot contain a link. Unlike HTML or most of Markdown, multiple lines are not meant to be joined together by a rendering engine.
2.3 Related Work
Standards-wise, the reference W3C recommendation for a Web protocol binding of Linked Data is the Linked Data Platform (LDP) [12, 18]. It delineates regulations for using HTTP to read and write RDF resources organized into typed containers, giving clients a uniform protocol for create-read-update-delete (CRUD) operations and membership management over Linked Data on the web. By contrast, the SPARQL query language specification has HTTP bindings wired into it [8]. As LDP is strongly aimed at leveraging the full potential of HTTP semantics, this work does not seek to implement or replicate LDP in the Small Web, but rather to encode the strict interactions assumed by the latter in terms of data consumption.
The academic literature on the aforementioned Small Web protocols of the current generation is virtually non-existent. However, that of bridging Linked Data over to non-HTTP protocols is not an entirely novel idea, as it has been explored in the Internet of Things, under constraints that are largely orthogonal to the ones addressed here. Hasemann et al. devised a method for constrained sensor nodes to publish RDF descriptions of themselves over 6LoWPAN/CoAP [9]; their concern, however, is the compactness of the RDF encoding for devices with kilobytes of RAM and no filesystem, while the transport exposed to clients remains a document-level, CoAP-based interface with explicit GET/POST/DELETE/SUBSCRIBE operations, i.e. a request/response model with a method set and a subscription mechanism that plays the role headers and content negotiation play in HTTP. The closest work to our intent is that of Loseto et al., who map the Linked Data Platform's container and membership model onto CoAP [14]; their own reference implementation describes its proxy component as translating “LDP-HTTP request methods and headers into the corresponding LDP-CoAP ones,” aimed squarely at machine-to-machine Web of Things scenarios. Both lines of work therefore translate yet retain, at the transport level, a verb-and-header apparatus and target machine clients that never need a document to double as human-navigable hypertext, unlike the wire format of most of the Small Web (§3.2). They therefore cannot be transposed wholesale onto a core contract that defines only a request URL and a binary OK/Error outcome (§3.1); they nonetheless corroborate the premise that a leaner transfer protocol need not preclude Linked Data compliance.
Data decentralisation as proposed by the Solid framework is also worthy of observation, not least because it shares this work's ambition of returning agency over Linked Data to its producers. Solid Pods are LDP containers exposing the standard HTTP method set, with access control implemented as a separate ACL resource whose association is advertised through a Link header, with resource/storage discovery similarly conducted through Link headers rather than through the request path or body [16]. None of these mechanisms has a counterpart in the present core contract, which defines no headers and no verb beyond an implicit request (and, via sister protocols, an equally implicit upload). Solid adoption has seen prudent yet steady progression. Notably, deployments such as SOLIOT do not merely graft Solid's access-control philosophy onto an unchanged HTTP substrate: Bader and Maleshkova explore adapting it to CoAP as well as another lightweight IoT protocol, MQTT [3], showing that Solid's data-control model can migrate away from HTTP. Yet CoAP retains a REST-like method set and an options mechanism, while MQTT retains topics, quality-of-service levels, and a broker-mediated publish/subscribe contract—neither dispenses with the request/response or headers-equivalent apparatus altogether. The applicability of the Solid data model, and in particular its container-based resource organisation and access-control paradigm, to header-less, single-verb Small Web transports such as Gemini and Spartan therefore remains an open question, one only partially anticipated by SOLIOT's move away from HTTP.
3 Specification
With one of the main goals of this project being to effectively operationalise the publishing of and access to Linked Open Data using existing Smolweb protocols, the specification was derived bottom-up for the protocols themselves, with a core contract serving as a lowest common denominator among those that support a reading operation and those that support uploading, and optional capabilities that are unique to some protocols.
This section illustrates the contract for transport and for the wire format—there are no server-side storage bindings.
In what follows, the requirement semantics assumes to adhere to RFC 2119 insofar as adopting the terminology: must / should / may / must not.
Figure 1 previews the reasoning behind what follows: five heterogeneous protocols are reduced to a single, minimal Core Contract (§3.1)–the same one that, via proxying, also admits HTTP(S) resources (§3.2)–which is in turn realised through one shared wire format (§3.2–4). Protocol-specific capabilities beyond this floor are reintroduced, but only as optional bindings (§5), never as requirements on the contract itself.
Five labelled boxes, one each for the protocols Gemini, Titan, Spartan, Nex, and NPS, each annotated with its read/submit capability, transport (TLS or plaintext), and status-code support. Arrows from all five boxes point down into one wide box labelled Core Contract: one IRI in, OK or Error out. A further arrow points down from the Core Contract box into a box labelled Wire format, PDL-HL, leading to Gemtext-LD serialisation.
Figure 1: Core Contract as the lowest common denominator of the five Small Web protocols. Protocol-specific capabilities (TLS, status codes, read/submit etc.) are not requirements of the contract, but optional bindings layered back on top of it (§5).
3.1 Core Contract
The atom of information is the RDF triple: (subject, predicate, object). This specification retains the assumptions of the RDF 1.1 Concepts and Abstract Syntax specification, whereby subject is a non-literal node, predicate is an IRI, and object is any RDF node [7]. A resource description may be the set of all triples sharing a node as subject or object.
3.1.1 Identity of RDF Resources. Every resource that can be requested by a client is identified by an IRI. IRIs may use the <http[s]://>
scheme, or a Small Web scheme (e.g., gemini://, spartan://) if natively served in that protocol.
3.1.2 Read Request. The Linked Data resource being requested is defined by the request URL alone, as it is the only omnipresent and consistent element across protocols. The IRI of said resource is provided in one or two ways:
as the full request URL, e.g. gemini://example.smol/0123456789abcdef; or\
encoded in the URL path, in the form <protocol>//<host>[:<port>]/<rsrcIri>, where <rsrcIri> is URL-encoded. Example: nex://example.smol/http%3A%2F%2Fwww.wikidata.org%2Fentity%2FQ257469.\
A client sends exactly one IRI per request using either method. Whether the resource data are to be locally retrieved by the server, or proxied from a third-party host, is determined solely by the first or second request form being used, respectively, i.e. the part of the request path that follows the system one is an URL-encoded IRI. HTTP(S) resources will typically be proxied. Method 2 should not be used to proxy to the same host that received the request.
Lastly, there is no assumption on returning a “not found” response when there are no data describing a possible entity URI: this is in line with the underlying assumption of the Semantic Web or the LDP specification, where a HTTP 404 or an empty body are both possible, but neither is mandated. Hence, only “OK” and “Error” outcomes are defined.
3.1.3 Hosting vs. Retrieval. A provider that hosts Linked Data natively on a Small Web site, i.e. a capsule or station, may commit to giving each data entity a cross-protocol identity, by minting URIs for each supported protocol. One key consequence is that, due to its repercussions on the logical interpretation of entity statements, the authority host and domain of resource IRIs never change across protocols. This no-rewrite policy implies, for instance, that there shall be no gemini://www.wikidata.org/... resource to match the original HTTP one—barring the unlikely event that Wikidata starts serving to the Small Web. For that reason, it is also safe to assume ontological equivalence, such as with implied owl:sameAs or owl:equivalentClass statements, across entity IRIs that differ by protocol. Because domain names are not rewritten, the server supporting multiple protocols, be they HTTP or Small Web, shall be one and the same, therefore it should retrieve the same statements regardless of the protocol: it may therefore add equivalence statements such as (nex://example.smol/entity1 owl:sameAs gemini://example.smol/entity1).
Figure 2 operationalises the Core Contract as described: whether an IRI is resolved locally or proxied from HTTP(S) is decided once, by scheme alone, and both paths converge on one binary outcome.
A flowchart beginning at a client request carrying one IRI. A decision on the IRI scheme branches left to a proxy fetch for http or https resources, using Turtle with an RDF/XML fallback and a scheme-swap retry, or right to a local lookup for a native Small Web scheme. Both branches converge on a decision of whether a resource description was found. If not, the flow ends in an Error outcome. If found, the resource is serialised as Gemtext-LD, in expanded or condensed mode, ending in an OK outcome.
Figure 2: Resolution flow for a Core Contract request.
3.2 Wire format and Hyperlinking
For the so-called wire format, i.e. the markup format assumed for data transfer by default, we postulate the following:
that it must be a prefix-dispatched, line-oriented markup format;\
that it supports hyperlinking by means of URL parsing.\
The acronym PDL-HL will be use to refer to the above. In (1), we call a markup format line-oriented when whitespace is not collapsed across newlines, unlike in HTML; and prefix-dispatched when the syntactic construct into which a line is parsed is determined solely by its leading k characters, i.e. a fixed-length prefix discriminator, or sigil. It follows that each line constitutes an independent syntactic unit. No assumption is made, however, as to how many consecutive lines constitute an atomic semantic unit: that determination is delegated to the hypertext layer, which is why (2) is required.
The assumed semantics of hyperlinks warrant a commitment at the core contract level regarding whether or not wire formats must reflect the identity of Linked Data resources in their rendering. Because there is no guarantee that a line representing a link has an anchor text, the identity of the referenced RDF resource must be reconstructable from the rendered hyperlink alone. Concretely, while a hyperlink should resolve to a representation of the resource under any supported protocol, it must be rendered so as to display the protocol of its identity (typically HTTP for non-Small-Web resources). This separation of concerns permits URL (but not entity IRI) rewriting in formats that support anchor texts, preserves isomorphism between wire-format Linked Data and RDF, and ensures that Linked Data principles 3 and 4 continue to hold: it remains possible to discover entities and resolve their identities, even when this should need requests to branch off from the Small Web to HTTP—a circumstance that several protocol designs contemplate.
4 Gemtext-LD
Gemtext (§2.2) is a PDL-HL format, and one that all five protocols can express.10 It is human-readable and navigable, with hyperlinks as first-class entities. To implement the wire format specified in §3.2, this specification proposes a Gemtext rendering for Linked Data that is round-trip-parseable to RDF. No alternative content negotiation is needed as the MIME type text/gemini is assumed.
In Gemtext, each line carries exactly one syntactic role, determined by its prefix, and no inline markup is permitted. This implies that an RDF triple cannot be expressed as a single line, if at least the subject and, if an IRI, the object of that triple are to be rendered as hyperlinks. The Gemtext-LD format therefore seeks a trade-off between hyperlink density and length of context per semantic unit. The rule of thumb is that every hyperlink should use the same protocol as the one of the request that caused the document to be returned: if the RDF resource is third-party (i.e. its data are not stored on the receiving site), it shall be rewritten in proxy form according to the respective protocol binding (§5 below).
A summary of the Gemtext-LD specification, available in the project's source code repository, follows.
4.1 Document Structure
A response body is a sequence of resource descriptions. A resource description takes the following form:
Level-1 headings therefore delimit groups of statements. For property blocks, two serialisation modes are defined:
Expanded (default): one Gemtext line per triple.\
Condensed (optional): triples grouped under level-2 headings of the form "## <predicate>".\
4.2 Predicate Encoding — Expanded Mode
In this encoding, described in Table 2, the creation of hyperlinks for objects–their URIs if resources, or datatype IRIs if literals–is prioritised. Each triple (subject, P, object) produces exactly one Gemtext line.
Table 2: Expanded mode object encoding.
Object type | Output line |
|---|---|
IRI (HTTP or Small Web scheme) | => <uri> <P> : <shortUri> |
IRI (other scheme e.g. URN) | <P>: <uri> |
Blank node | <P>: :<id> |
Simple literal | <P>: "<value>" |
Language-tagged literal | <P>: "<value>"@<lang> |
Datatyped literal |
|
The example below renders the Wikidata statements that entity Q257469 is a video game with two English titles. It assumes a Gemini proxy is running on localhost (port 1965 implicit). Line 2 is wrapped for display purposes.
4.3 Predicate Encoding — Condensed Mode
Condensed mode supports QNames, whose prefixes are declared in a preamble with level-1 heading # Prefixes. Prefix URIs are not hyperlinks. Triples are grouped by predicate: groupings may be sorted. Each predicate group produces:
A level-2 heading: ## <shortP> \
If P is an HTTP or Small Web URI, a property navigation link: => <P> ↗ <shortP> 11 \
One line per object value, as defined in Table 3.\
A trailing blank line.\
Table 3: Condensed mode object encoding (within a predicate group).
Object type | Output line |
|---|---|
IRI (HTTP or Small Web scheme) | => <uri> <shortUri> |
IRI (other scheme e.g. URN) | <uri> |
Blank node | :<id> |
Simple literal | "<value>" |
Language-tagged literal | "<value>"@<lang> |
Datatyped literal |
|
The listing below, sans the trailing blank line, shows the same statements in condensed mode. Note that, due to the increased context length and the more verbose syntax for the preamble and predicate headers, this mode is only recommended for RDF descriptions with non-human-readable IRIs and many statements sharing the same predicate.
Because of the significant differences between expanded and condensed encoding, it should be the client application's responsibility to request the latter if needed. To that end, HTTP-like query parameters may be supported.
5 Protocol Bindings
What was described in Sections 3.1 and 3.2 can be summarised into a single statement: that “IRI-in → PDL-HL-out” is the irreducible contract that every Small Web protocol can and shall satisfy; even Nex, which has no headers at all.
Let us now delve into the specifics for each Small Web protocol supported by this work.
5.1 Gemini (Read)
Gemini supports read requests with no parameters. Of the status codes it returns, 10-19 is for requesting user input.
Transport: TLS 1.2+, port 1965\
Request line: gemini://<host>[:<port>]/<path>rn \
Response: <status> <meta>rn<body> \
<meta> may include any of lang and charset, i.e. the only response parameters that a Gemini client should process: these are assumed to be the union of languages and charsets of all the RDF literals present in the Gemtext-LD.
IRI derivation: the full request URL is the lookup IRI iff the path is not an absolute URL; otherwise, the path.
Interactive entry point: a request to / with no query string returns a Gemini input prompt (status 10). The user's input is the IRI to resolve.
Responses: status 40 (temporary failure) for validation errors; status 20 for success, MIME text/gemini.
5.2 Titan (Submit, depends on Gemini)
Titan defaults to the same port as Gemini. A Titan request is distinguished by the titan:// scheme in the request line.
Request line: titan://<host>/<path>;size=<n>;rn \
Payload: exactly <n> bytes; the first line of the payload is the IRI to resolve.\
Response: same as Gemini (20 text/gemini or 40 on error).\
5.3 Spartan (Read and Submit)
Being the only protocol to natively support submit, it needs dedicated path semantics to denote either operation.
Transport: plaintext TCP, port 300\
Request line: <host> <path> <length>rn \
Payload (if length > 0): exactly length bytes\
Path semantics: /submit reads payload, first line as IRI; otherwise /<anything> is treated as the IRI to resolve.
Response line: <status> <meta>rn<body>. A successful response status is “2 text/gemini”, whereas temporary and permanent failures are “4 <message>” and “5 <message>”, respectively
5.4 Nex (Read)
Transport: plaintext TCP, port 1900\
Request: a single line /<percent-encoded-IRI>n \
Response: raw plaintext body, no status line\
The IRI is URL-decoded before lookup. Since Nex has no status header, outcomes (success, error) are communicated entirely within the body. A Gemtext body is assumed to be returned, of which Nex clients treat only => lines specially.
5.5 NPS (Submit, depends on Nex)
NPS is the write companion to Nex, as Titan is to Gemini. Unlike Titan, it defaults to a separate port.
Transport: plaintext TCP, port 1915\
Request: multiple text lines terminated by a line containing only the "." character\
Response: raw plaintext body, no status line\
The first non-empty line of the payload is the IRI to resolve. Multi-line payloads allow editor-friendly submission: the user composes text, appends a terminator line, and pipes to the server.
6 Implementation
The core contract, bindings for all the protocols discussed, and a parser and serializer for the Gemtext-LD format have all received a multi-platform reference implementation. Rust was chosen as the implementation language due to its resource efficiency and the significant amount of Rust software already available for Gemini.
Two separate units of Rust code, or crates, are available: a library for the Gemtext-LD format, and a binary Linked Data server and proxy, named Chaykin, that uses it. The code is portable and can be compiled for any operating system and architecture for which there is a Rust toolchain, or built as a Docker image. A Chaykin binary consists of a small executable (6-8.5MB on disk depending on the target platform), which bootstraps in less than 10MB of system memory on Unix systems, while the corresponding Docker image takes less than 35MB of storage space. At present, the Gemtext-LD specification only supports statements as plain triples, with no named graphs. A syntax that supports named graphs (i.e. quads) and an isomorphism to RDF for statement annotation are both being considered for future work.
The implementation is open source and dual-licensed MIT and Apache 2.0, as is customary in many FOSS projects in Rust. A single Git repository12 holds the source code of both Chaykin and Gemtext-LD, as well as releases (both of source code and as multi-platform executable binaries), specifications, and usage documentation. Prepared Docker images to host on DockerHub are underway.
7 Validation
Conformance cannot be inferred merely from co-developing the specification and the implementation, as that argument would be circular. We instead treat every normative (must/should) requirement of the Core Contract (§3.1) and the five protocol bindings (§5) as an individually falsifiable claim, discharged either by an automated test in the round-trip suite below, or by a live protocol probe against the running server. A full requirement-by-requirement checklist mapping each claim to the test or probe that exercises it is maintained alongside the source code (§6) and referenced, rather than reproduced, here for space. Other salient validity factors follow.
7.1 Round-trip fidelity of Gemtext-LD
To ensure a lossless conversion of RDF data into Gemtext-LD and back to RDF, the gemtext-ld crate carries 42 unit tests, covering five node types (IRIs, blank nodes, and three forms of literals) in both expanded and condensed serialisation modes. A substantial subset targets edge cases: misclassifying newlines in literals, non-registered namespace prefixes, subjects and predicates within registered namespaces, non-HTTP datatype URIs, complex BCP 47 language tags, embedded double-quotes, non-HTTP IRI objects in condensed mode, declaration and round-tripping of document-local prefixes for namespaces outside the registered set (§4), and disambiguation of nested, overlapping namespaces (e.g. Wikidata's prop/ versus prop/direct/). The implementation was refined until the suite passed and all edge cases round-tripped correctly without modification. As this suite grows with the format, the count above is a snapshot as of this writing; the authoritative, current figure is reproducible by running the test suite in the repository (§6).
7.2 Client-side Data Retrieval
The Chaykin server returns data per the client requests in all the five supported protocols. This has been verified through dedicated Small Web clients and general-purpose command line tools.
For the former, the browsers Lagrange13 and Alhena14 were used. Both are actively developed, cross-platform, and aiming to support all the protocols contemplated here, though neither supports NPS at the moment. For Spartan upload and Titan, these browsers open dialogues for inputting the entity URI in plaintext, which they URLencode in the request. Following the standard practice of IETF/W3C working groups, it was verified that: IRIs are resolved, links are navigable, condensed/expanded modes display correctly, and the HTTP proxy path fetches third-party RDF/XML or Turtle data.
Both clients successfully retrieved Gemtext-LD data requested through all supported protocols, though Alhena's behaviour with data retrieved via Nex was to save the payload to disk, due to strictly depending on the file extension. A revision of the protocol bindings to support lightweight, protocol-specific content negotiations and redirecting to URLs ending in .gmi could be considered. Both browsers seamlessly follow protocol switches in hyperlinks, e.g. spartan:// links in a document retrieved via Gemini, though this is an edge case assumed across different capsules/stations.
The same characteristics, sans navigability and rendering, were verified using command line internet clients netcat and openssl. The following is a minimal test suite in Bash scripting on a local server accessed in loopback.
Two store backends are available: a naive store performing a linear scan per request—O(n) for expanded Gemtext-LD and O(nlog n) for condensed, with n = number of triples—and an indexed store trading a costlier, tree-structured load for an O(k) resource lookup, with k = number of statements about the requested resource, obviating the re-grouping and re-sorting condensed mode otherwise requires. Preliminary benchmarks confirm the expected trade-off–faster lookup, slower load–but also that its magnitude is sensitive to the shape of the data, not merely its size; a systematic characterisation across representative Linked Data distributions is left to the extended validation report mentioned below.
7.3 Scope and Threats to Validity
The above establishes conformance and cross-client interoperability for a single reference implementation; it is not a broader empirical assessment of adoption, usability, or comparative advantage over incumbent approaches such as the Linked Data Platform. The capsules exercised throughout this section–a synthetic dataset for round-trip testing and a proxied Wikidata entity for client retrieval–are rather short of a representative sample of Linked Open Data in the wild. Establishing that the contract is not merely satisfiable but useful requires deploying it over real capsules and stations at scale, gathering reports from a wider set of clients, and the SPARQL-over-Small-Web and LDP-comparison studies already outlined as future work (§8). An extended validation report–the full conformance checklist, client logs, and performance benchmarks across representative Linked Data distributions–is maintained alongside the codebase (§6) and updated independently of this paper.
8 Conclusions
The aim of this work is to initiate the process of determining the extent of scholarly recognition that can be granted to the hitherto sparse Small Web movement. Recognising that it is not the result of an organised campaign, but rather an expression of individual reflections on contemporary interaction with the Web, this contribution organises these reflections and builds upon them by offering a purely academic perspective: that of Linked Data.
By providing an initial set of proposals for an irreducible contract that should hold for protocols present and future, bindings for contemporary protocols, and a document-oriented serialisation format, we have made significant progress in terms of serving and proxying RDF entity descriptions. Whilst we debate whether this work should attempt a full re-specification of the Linked Data Platform, there emerge immediate research questions to be addressed in the imminent future, the most pressing one being possible support for SPARQL querying within the contract. Rather than seeking to allow cross-protocol SPARQL querying over the Small Web at all costs, we wish to frame our research intent as finding a common ground for querying the HTTP web, the Geminispace, and other sites and stations, without compromising the minimalism imposed upon the Small Web protocols.
Additional work concerns the evolution of Gemtext-LD to support named graphs and RDF, and an evaluation of the expressiveness, scope efficiency, and information density of this format. Only then will it make sense to compare the contract against the LDP on a comparison axis that is not "which is better" but "what is lost and what is gained by stripping HTTP", possibly helping bestow upon the Small Web the scientific dignity we believe it merits.
Notes
Veronica-2 stats, Internet Archive captures: https://web.archive.org/web/20260000000000/https://gopher.floodgap.com/gopher/gw?gopher/0/v2/vstat.
Gemini Protocol: https://geminiprotocol.net/.
Geminispace statistics (through Gemini): gemini://gemini.bortzmeyer.org/software/lupa/stats.gmi.
Spartan protocol (HTTP-proxied page): https://portal.mozz.us/spartan/spartan.mozz.us/.
Nex protocol: https://portal.mozz.us/nex/nightfall.city/nex/.
Nightfall's Postal Service (NPS): https://nightfall.city/nps/info/.
An online document that provides a round-up of such protocols is being maintained by the proponent of Scorpion: https://dbohdan.com/archive/scorpion/zzo38computer.org/smallweb.txt/ (archived from the original being served via Scorpion).
Gemini hypertext format, aka “gemtext”, specification: https://geminiprotocol.net/docs/gemtext-specification.gmi.
The Nex protocol only mandates a plaintext format in which lines beginning with => are treated as link lines. A Nex client will follow => links but treat all other lines as plain text. Gemtext is therefore a valid Nex payload, but Nex does not assume Gemtext semantics beyond link lines.
This is a legal syntax as Gemtext and all the protocols being treated either support UTF-8 or mandate no encoding.
GitHub repository: https://github.com/alexdma/chaykin.
Lagrange browser: https://gmi.skyjake.fi/lagrange/.
Alhena browser: https://www.metaloupe.com/alhena/alhena.html.
References
[1] Farhad Anklesaria and Mark McCahill. 1993. The Internet Gopher. In Intelligent Information Retrieval: The Case of Astronomy and Related Space Sciences, A. Heck and F. Murtagh (Eds.). Springer Netherlands, Dordrecht, 119–125. https://doi.org/10.1007/978-0-585-33110-29
[2] Alessio Antonini, Megan Bushnell, Christopher Ohge, Francesca Benatti, Alessandro Adamou, and Sam Brooker. 2023. Hypertext as Method: Reflections on Hypertext as Design Logic. In Proceedings of the 34th ACM Conference on Hypertext and Social Media (Rome, Italy) (HT ’23). Association for Computing Machinery, New York, NY, USA, Article 45, 4 pages. https://doi.org/10.1145/3603163.3609074
[3] Sebastian R. Bader and Maria Maleshkova. 2020. SOLIOT: Decentralized Data Control and Interactions for IoT. Future Internet 12, 6 (2020), 105. https://doi.org/10.3390/fi12060105
[4] Tim Berners-Lee. 2006. Linked Data. W3C Design Note. Accessed: 2026-04-27. https://www.w3.org/DesignIssues/LinkedData.html
[5] Kevin Boone. 2026. Small web, IndieWeb, Gemini… A guide to the retro-web. Accessed: 2026-04-27. https://kevinboone.me/web-adjacent.html
[6] Mike Caulfield. 2015. The Garden and the Stream: A Technopastoral. Keynote presented at the Digital Pedagogy Lab Institute, Madison, WI. Accessed: 2026-04-27. https://hapgood.us/2015/10/17/the-garden-and-the-stream-a-technopastoral/
[7] Richard Cyganiak, David Wood, and Markus Lanthaler. 2014. RDF 1.1 Concepts and Abstract Syntax. W3C Recommendation. World Wide Web Consortium. https://www.w3.org/TR/rdf11-concepts/
[8] Lee Feigenbaum, Gregory Todd Williams, Kendall Grant Clark, and Elias Torres. 2013. SPARQL 1.1 Protocol. W3C Recommendation. World Wide Web Consortium. https://www.w3.org/TR/sparql11-protocol/
[9] Henning Hasemann, Alexander Kröller, and Max Pagel. 2012. RDF Provisioning for the Internet of Things. In 2012 3rd International Conference on the Internet of Things (IoT). IEEE, 143–150. https://doi.org/10.1109/IOT.2012.6402316
[10] James Hendler and Tim Berners-Lee. 2010. From the Semantic Web to social machines: A research challenge for AI on the World Wide Web. Artificial Intelligence 174, 2 (2010), 156–161. https://doi.org/10.1016/j.artint.2009.11.010
[11] Pascal Hitzler. 2021. A review of the Semantic Web field. Commun. ACM 64, 2 (2021), 76–83. https://doi.org/10.1145/3397512
[12] Arnaud Le Hors and Steve Speicher. 2013. The linked data platform (LDP). In 22nd International World Wide Web Conference, WWW ’13, Rio de Janeiro, Brazil, May 13-17, 2013, Companion Volume, Leslie Carr, Alberto H. F. Laender, Bernadette Farias Lóscio, Irwin King, Marcus Fontoura, Denny Vrandecic, Lora Aroyo, José Palazzo M. de Oliveira, Fernanda Lima, and Erik Wilde (Eds.). International World Wide Web Conferences Steering Committee / ACM, 1–2. https://doi.org/10.1145/2487788.2487790
[13] Martine S. Lenders, Carsten Bormann, Thomas C. Schmidt, and Matthias Wählisch. 2025. A Leaner and Faster Web: How CBOR Can Improve Dynamic Content Encoding in JSON and DNS over HTTPS. CoRR abs/2512.12067 (2025), 20 pages. arXiv:2512.12067 https://doi.org/10.48550/ARXIV.2512.12067
[14] Giuseppe Loseto, Saverio Ieva, Filippo Gramegna, Michele Ruta, Floriano Scioscia, and Eugenio Di Sciascio. 2016. Linking the Web of Things: LDP-CoAP Mapping. In The 7th International Conference on Ambient Systems, Networks and Technologies (ANT 2016) / The 6th International Conference on Sustainable Energy Information Technology (SEIT-2016) / Affiliated Workshops, May 23-26, 2016, Madrid, Spain(Procedia Computer Science), Elhadi M. Shakshuki (Ed.). Elsevier, 1182–1187. https://doi.org/10.1016/J.PROCS.2016.04.244
[15] Eli Pariser. 2011. The Filter Bubble: What the Internet Is Hiding from You. Penguin Press, New York.
[16] Andrei Vlad Sambra, Essam Mansour, Sandro Hawke, Maged Zereba, Nicola Greco, Abdurrahman Ghanem, Dmitri Zagidulin, Ashraf Aboulnaga, and Tim Berners-Lee. 2016. Solid: A Platform for Decentralized Social Applications Based on Linked Data. Technical Report. MIT CSAIL & Qatar Computing Research Institute. http://emansour.com/research/lusail/solidprotocols.pdf
[17] Roy Schestowitz. 2021. 2021: The Year of Gemini on the Internet. Techrights. Accessed: 2026-07-23. https://techrights.org/o/2021/09/29/gemini-phenomenal-growth/
[18] Steve Speicher, John Arwe, and Ashok Malhotra. 2015. Linked Data Platform 1.0. W3C Recommendation. World Wide Web Consortium. https://www.w3.org/TR/ldp/
[19] Tim Wu. 2016. The Attention Merchants: The Epic Scramble to Get Inside Our Heads. Alfred A. Knopf, New York.
[20] David Zimmerman. 1991. The Finger User Information Protocol. RFC 1288. RFC Editor. https://www.rfc-editor.org/rfc/rfc1288
Source
Imported from ACM’s structured HTML source. ACM Reference Format: Alessandro Adamou. 2026. Bringing the Web of Data to the Small Web: a Multi-Protocol Contract and the Chaykin Linked Data Server. In 37th ACM Conference on Hypertext (HT '26), September 14--18, 2026, London, United Kingdom. ACM, New York, NY, USA 9 Pages. https://doi.org/10.1145/3800935.3830868
Do you like what you are reading? Subscribe to receive updates.
Unsubscribe anytime