Skip to main content
Can We Archive the Web Without Wrecking the Planet?Archival Management
4 min readFor Archivists and Digital Preservation Specialists

Can We Archive the Web Without Wrecking the Planet?

The Growing Concern

Researchers recently gathered at Durham University for the NetDRIVE 2026 Summer School, focusing on digital infrastructure and sustainability. These discussions have sparked urgent questions for anyone managing web Accessioning programs. Your web archive is expanding, storage costs are rising, and leadership is concerned about the organization's carbon footprint. You're caught between preserving digital records and meeting environmental commitments.

Here's what practitioners are asking.

Q1: Does Web Accessioning Have a Significant Carbon Footprint?

Yes, and it's larger than most teams realize.

Every crawl uses electricity. Each stored WARC file requires power and cooling. Accessing an archived page involves server processing. Multiply this by terabytes of data, redundant copies, and 24/7 infrastructure, and you're looking at significant energy consumption.

The question isn't whether web Accessioning impacts the environment, but whether your program can justify that impact with clear value. Capturing everything indiscriminately because "storage is cheap" creates an environmental liability alongside your preservation liability.

Q2: How Do We Measure Our Web Archive's Environmental Impact?

Start with measurable factors.

First, assess your storage footprint. How many terabytes are you storing? How many copies? Where are those servers located, and what's the energy mix of that region? A data center in Iceland using geothermal power has a different carbon profile than one in a coal-dependent grid.

Next, examine your crawl frequency and scope. Are you recrawling static sites weekly when monthly would suffice? Are you capturing entire domains when only specific records are needed? Every unnecessary crawl wastes energy.

Don't get bogged down in precise carbon emission calculations. Focus on variables you control: crawl frequency, retention periods, storage redundancy, and access patterns. These decisions drive your environmental impact.

Q3: Balancing Legal Compliance with Sustainability

"Comprehensive" doesn't mean "everything, forever."

Legal obligations require preserving records relevant to anticipated litigation, not crawling every page daily. This is a Records Control Schedule issue, not a technology one.

Work with legal to define what constitutes a record in your web content. Marketing campaign pages? Likely records with defined retention. Temporary event announcements? Maybe not. The 47th revision of a product spec page? Depends on your industry and regulations.

Once you've established criteria, design targeted crawls that meet compliance without capturing unnecessary data. If legal pushes back, ask for specific regulations or case law requiring indiscriminate Accessioning. They likely won't find it.

Q4: Is AI a Sustainable Solution for Accessioning?

It depends on what you're replacing.

If you're capturing everything and considering AI to filter intelligently, you might reduce your storage footprint. However, training and running AI models consumes energy. Compare the energy cost of AI processing against storing everything it would filter out.

For most, a well-designed Business Classification Scheme applied to web content outperforms AI in accuracy and energy efficiency. You don't need machine learning to determine that investor relations pages are records and cafeteria menus aren't.

AI might help identify substantive content changes, avoiding multiple identical captures of static pages. But simple hash comparison (fixity checking) can achieve this with less computational overhead.

Q5: Making the Case for Sustainability to Leadership

Focus on cost and risk, not just environmental concerns.

Sustainable web Accessioning is about running a defensible, cost-effective program. Eliminating unnecessary captures and optimizing storage reduces infrastructure costs. Implementing rational retention periods instead of "keep everything forever" reduces eDiscovery exposure and storage costs.

Frame it like this: "Our current approach captures 40TB annually with no defined retention limits. By aligning appraisal criteria with our Records Control Schedule, we can reduce capture volume by X percent, cut storage costs, and meet all compliance obligations."

If your organization has Net Zero commitments, connect your program to those goals. But make the business case first. Sustainability that saves money while reducing legal risk is a winning strategy.

Q6: A Practical Change to Implement This Quarter

Adopt event-based retention for your web archives.

Currently, you might keep everything until budget constraints force a decision. That's not a strategy; it's deferred decision-making.

Choose one record series in your web archive. Define the business trigger that starts the retention clock. For product documentation, it might be product discontinuation. For campaign sites, it might be the campaign end date. Then set a retention period based on your Records Control Schedule.

Proving you can apply defensible disposition to web archives establishes a framework for sustainable practice. You're not just reducing storage costs and environmental impact; you're showing that your program makes conscious, documented decisions about what to preserve and why.

That's the foundation of both sustainability and defensibility.

Further Resources

The International Internet Preservation Consortium maintains working groups on web Accessioning practices, including sustainability. ISO 14721 (the OAIS reference model) provides a framework for long-term preservation decisions, including appraisal and disposition.

Your best resource is your own Records Control Schedule. If you don't have retention periods defined for web content, that's the real issue. Solve that, and sustainability follows.

You Might Also Like