Link Rot
A reader's summary of why the footnote you clicked last year may lead nowhere today, and what that does to an argument that depends on it.
The term at a glance
Link rot is the gradual failure of hyperlinks as the pages they point to move, change or disappear. The familiar symptom is the “404 Not Found” error, but a link can also rot by redirecting to a home page, to a parked domain or to an unrelated page. The phrase is a metaphor borrowed from decay in the physical world, and it quietly makes the process sound natural, when most dead links are the result of somebody's decision or neglect.
Origin
The problem is as old as the web. In 1998, Tim Berners-Lee, who invented it, published a short note on the W3C site titled Cool URIs don't change. His argument was that when a page disappears, it is almost always because its owner chose to reorganize or stopped caring, and that good practice is to design addresses that outlast the software currently serving them. The advice was widely admired and widely ignored. Within a few years “link rot” was a standard complaint among web writers, librarians and academics.
History and context
Archiving began almost at the same time as the decay. The Internet Archive was founded by Brewster Kahle in 1996 and opened its Wayback Machine to the public in 2001, and it has since stored far more than a trillion page captures. In 2014 Jonathan Zittrain, Kendra Albert and Lawrence Lessig reported that more than 70% of the URLs in several Harvard law journals, and about half of those in US Supreme Court opinions, no longer led to the information originally cited, which led to Perma.cc. The same year, a study of millions of scholarly articles found that roughly one in five suffered from reference rot. In May 2024 the Pew Research Center put numbers on the whole web: 38% of pages that existed in 2013 were no longer accessible in 2023, about a quarter of all pages from the decade were gone, 23% of news pages and 21% of government pages contained at least one broken link, and 54% of Wikipedia articles had at least one reference pointing to a dead page. Social platforms are a special case, because a deleted post or an account that goes private takes its replies and quote-posts with it.
Main ideas
Link rot and content drift are different failures
Link rot is the plain case: the address returns an error or nothing at all. Content drift is subtler, because the address still works but the page now says something else. Researchers who studied scholarly citations (Klein and colleagues, 2014) used the umbrella term reference rot for both, and found that drift can leave a citation pointing at a page that no longer supports the claim it was cited for.
A link is a promise nobody is obliged to keep
A URL names a place on someone's server, not a document. Publishers redesign sites, change content-management systems, shut down, get acquired or simply stop paying for hosting, and nothing in the web's design forces them to keep old addresses alive. Berners-Lee's 1998 note argued that this is a choice, not a law of nature: a site owner who plans well can keep an address working for decades.
The web is not forgetful in a uniform way
Old pages from big institutions, universities and governments often survive, while personal blogs, small publishers, startups and social-media posts vanish fastest. What disappears is therefore skewed toward the ordinary and the informal, which is a kind of survivorship bias for historians of the period.
Archiving is the main countermeasure
The Internet Archive's Wayback Machine, opened to the public in 2001, stores dated snapshots of pages, and Perma.cc, launched in 2013 by the Harvard Law School Library, lets authors freeze a cited page and receive a permanent link. Wikipedia runs bots that look for dead references and point them at archived copies. All of these depend on someone capturing the page while it still exists.
Preservation is a public good with no owner
Keeping a page online costs its publisher a little every year and brings that publisher little in return, while the benefit of a durable record goes to everyone who cites it. That mismatch is why the job has fallen to libraries and nonprofits, and why a fix that relies on every publisher behaving well has never worked.
Critique
- Not every dead link is a loss. Much of what disappears was ephemeral to begin with: product pages for discontinued goods, event listings, press releases past their date. Counting every error as damage overstates what a future reader would actually want.
- Archives are partial and have their own gatekeepers. A snapshot exists only if a crawler or a person saved it, pages behind logins and heavily scripted sites are captured badly, and a copy kept by one nonprofit is only as safe as that nonprofit. Rights holders have also challenged archiving in court.
- Permanence has costs too. People sometimes have good reasons to take things down, and the European right to erasure reflects that. A web where nothing can rot would clash with privacy, and the same snapshot that rescues a source can resurface a post its author deleted, much as in the Streisand effect.
Impact
The practical effect is on anyone who argues from sources: a footnote that cannot be checked weakens an essay, a court opinion or a news story, and readers rarely notice until they try to follow it. Link rot also raises a quieter worry about the shape of the record. If the old web thins out while a flood of machine-written pages replaces it, as the dead internet theory and the research on model collapse suggest could happen, what remains easy to find may be less trustworthy than what is gone. The usual advice is simple: when citing an online source, save a copy to an archive and link to both, and treat a page that cannot be preserved as a claim that cannot be fully checked.