Skip to main content
Should Your Web Archive Serve Researchers or Just Store URLs?Archival Management
5 min readFor Archivists and Digital Preservation Specialists

Should Your Web Archive Serve Researchers or Just Store URLs?

Your web archive isn't just a vault; it's a dynamic research tool that thrives on active use. Many institutions treat archived websites as static collections, but this approach overlooks their true value: capturing change over time and evolving as collections themselves.

If you manage a web archive or oversee digital preservation, you're responsible for a resource that researchers want to explore, not just browse. This checklist will help you determine if your archive supports actual research or merely preserves content that can't be effectively queried.

Prerequisites

Before using this checklist, ensure:

  • Your web archive has at least 12 months of crawl data.
  • You have documentation of your archive's contents (crawl scope, frequency, format).
  • You can identify at least three potential user groups (computational researchers, historians, policy analysts).
  • You have the authority to propose changes to search infrastructure or access methods.

Checklist Items

1. Search infrastructure reflects research patterns, not just keyword matching

Your archive's search function should support how researchers work: tracking changes over time, comparing versions, and filtering by date ranges and domains. If your only search option is a site-level A-Z list or a single keyword box, you're forcing researchers to work around your system.

Done looks like: Researchers can filter by time period, domain, and content type. They can compare snapshots of the same page across dates. Search results show context, such as when the page was captured and how many versions exist.

2. You've observed real users searching your archive in the past six months

The National Archives conducted observed search exercises where one participant searched while another recorded their approach to dead-ends, filtering, and navigation. You can't fix usability problems you haven't watched happen.

Done looks like: You have notes from at least five observed search sessions showing where users get stuck, what workarounds they invent, and which features they ignore.

3. Sensitive content policies are documented and visible to users

Web archives face unique takedown challenges; content may need removal after preservation due to privacy concerns, legal requirements, or policy changes. Researchers need to know that gaps exist and why.

Done looks like: Your takedown policy is published. Removed content leaves a tombstone record explaining the removal category (not the specific reason). Researchers understand the archive is incomplete by design.

4. You offer multiple access methods for different research scales

Some researchers want to read specific pages closely, while others need full-text exports or derived datasets. Your archive should support both without forcing everyone through the same interface.

Done looks like: You provide at least three access paths: browser-based replay for close reading, bulk data access (even if limited), and documented APIs or file formats (WARC, CDX) for computational work.

5. Special collections or curated entry points exist

An archive containing thousands of government websites can overwhelm new users. Curated collections, themed groupings that showcase research possibilities, give researchers a starting point.

Done looks like: You maintain at least two special collections organized by theme, event, or policy area. Each collection includes context about why these sites matter and what research questions they might answer.

6. Researchers contribute to collection priorities

Your users know what's missing. The National Archives invited workshop participants to propose guest-curated special collections, leading to student placements focused on building new research pathways.

Done looks like: You have a documented process for researchers to suggest crawl targets, collection themes, or access improvements. You've acted on at least one user suggestion in the past year.

7. Training materials address web archive-specific concepts

Researchers unfamiliar with web archives need to understand crawl frequency, snapshot versioning, and format quirks (WARCs, CDX indexes). Don't assume they'll figure it out.

Done looks like: You offer documentation or workshops explaining how web archives differ from live websites, what researchers can and can't expect to find, and how to interpret capture dates and missing content.

8. You track what researchers want to study, not just what they access

Usage logs show what people clicked. Research proposals show what they wanted to accomplish but couldn't. The National Archives asked workshop participants to design imaginary research projects, revealing common themes: change over time, close reading of specific pages, and full-text analysis.

Done looks like: You've collected research proposals or project descriptions from at least ten potential users. You can identify three recurring research patterns your current system doesn't support well.

Common Mistakes

Treating web archives like paper collections. Paper archives are closed collections; everything arrived, got processed, and sits in boxes. Web archives grow continuously while researchers use them. You're managing a living collection that must balance ongoing capture with current access.

Assuming computational researchers are your only audience. Yes, some users want bulk data and API access. But historians, policy analysts, and students want to read specific pages, compare versions, and understand context. Build for both.

Hiding collection gaps. Removed content, failed crawls, and scope limitations aren't failures, they're documentation opportunities. Researchers trust archives that acknowledge what's missing more than archives that pretend to be complete.

Building in isolation. Your users will invent workarounds, suggest features, and propose research directions you haven't considered. The National Archives learned that researchers wanted to collaborate as guest curators, a solution the institution hadn't imagined.

Next Steps

Start with item 2: observe actual users searching your archive. You'll learn more from watching five people struggle with your interface than from reviewing analytics for six months.

Then address item 5 by creating one pilot special collection. Choose a theme where you already have strong coverage, a policy area, a government transition, or a specific time period. Document why these sites matter and what questions they help answer.

Finally, tackle item 6 by establishing a lightweight feedback mechanism. You don't need formal advisory boards. A quarterly email asking "what research would you do if you could?" generates actionable priorities.

Your web archive captures how government communicated, how policies evolved, and how public information changed over time. Whether researchers can actually study those patterns depends on whether you've designed for research use or just URL preservation.

You Might Also Like