Printed Gateways to the Internet: Internet Directory Books and Their Use in Web Archive Research
Date
relationships.isAuthorOf
Journal Title
Journal ISSN
Volume Title
Publisher
Abstract
In the 1990s and early 2000s, Internet directory books — print books featuring categorized listings of Internet resources — were among the bestselling computer and Internet reference books in the United States. Just as these books once guided their readers through the Internet and the web, they can help researchers navigate, curate, and evaluate web archive collections today. Internet directory books are largely absent from existing historical accounts of Internet search and navigation. In this dissertation, I first offer a historical overview that traces the genre's emergence in the early 1990s through its eventual decline in the early 2000s. Specifically, by examining publishers' marketing materials and different types of book reviews, I explain the books' continued popularity through the 1990s and early 2000s, despite the availability of online search engines and web directories. I argue that the books fulfilled a genuine navigational and pedagogical need of many Internet users in the 1990s and early 2000s by offering human-curated catalogs of Internet resources that could be understood and browsed in a familiar format. The rest of the dissertation focuses on how these books could be used by researchers to navigate and curate archived web materials, and to evaluate and detect the archival gaps and deficits in existing web archive collections. To do this, I built a dataset of 95,380 unique historical URLs collected from fifteen directory books published in five languages between 1998 and 2001. I detail the technical and curatorial decisions I took to create the dataset, and I demonstrate how the dataset can be used to help researchers access and curate archived web materials through three public-facing digital humanities projects. I then use a subset of the dataset to compare the Wayback Machine's archival coverage and quality for Chinese- and English-language URLs. While the Wayback Machine has high archival coverage rates for URLs in both languages, the median Chinese-language URL received 36 captures between 1996 and 2005, against 110 captures received by its English-language counterpart. The median archived Chinese-language web page examined in the dissertation also has a higher proportion of missing embedded resources. These findings — made possible by the historical URL listings in Internet directory books — highlight the unevenness of the Wayback Machine's record of the early web across different languages. In summary, the dissertation traces the history of Internet directory books, and shows how their listings can be used by researchers of the early web today.
Description
Thesis (Ph.D.)--University of Washington, 2026
