How to Assess the Exhaustiveness of Longitudinal Web Archives: A Case Study of the German Academic Web
Longitudinal web archives can be a foundation for investigating structural and content-based research questions. One prerequisite is that they contain a faithful representation of the relevant subset of the web. Therefore, an assessment of the authority of a given dataset with respect to a research question should precede the actual investigation. Next to proper creation and curation, this requires measures for estimating the potential of a longitudinal web archive to yield information about the
doi
10.1145/3372923.3404836
name
How to Assess the Exhaustiveness of Longitudinal Web Archives: A Case Study of the German Academic Web
pages
5
acm_url
https://dl.acm.org/doi/10.1145/3372923.3404836
authors
Michael Paris, Robert Jäschke
doi_url
https://doi.org/10.1145/3372923.3404836
license
restricted
summary
Longitudinal web archives can be a foundation for investigating structural and content-based research questions. One prerequisite is that they contain a faithful representation of the relevant subset of the web. Therefore, an assessment of the authority of a given dataset with respect to a research question should precede the actual investigation. Next to proper creation and curation, this requires measures for estimating the potential of a longitudinal web archive to yield information about the
keywords
dataset; longitudinal; web archive; focused web crawl; exhaustive
source_pdf
HT-2020_51-35_3372923/3372923.3404836.pdf
import_kind
full_text
open_access
false
ccs_concepts
• Information systems →Data extraction and integration;
displayAuthor
Michael Paris, Robert Jäschke