AI crawlers from Meta and Alibaba almost destroyed a volunteer-run LGBT history archive

| Source: Fast Company AI

Tags: web crawlers, Meta, Alibaba, AI training data, scraping, robots.txt, digital preservation

Aggressive AI web crawlers attributed to Meta and Alibaba nearly shut down a volunteer-run LGBT history archive by driving up hosting costs enough that the site founder spent nights building bot defenses — highlighting collateral damage of AI training data collection on small public-interest sites.

Details

A volunteer-operated LGBT history archive faced a near-existential crisis when AI crawlers linked to Meta and Alibaba hammered the site with traffic that spiked hosting costs beyond what the volunteer operator could sustain. The site's founder was forced to spend considerable time learning web infrastructure defenses to keep the crawlers out.\n\nThe incident represents an emerging pattern of AI training data collection imposing real costs on small, non-commercial sites that lack the leverage to fight back or negotiate licensing deals the way large publishers can. Questions arise about whether robots.txt opt-outs are being respected by all crawlers, and whether AI companies have obligations to smaller publishers beyond what major media agreements cover.\n\nNote: The Fast Company source article is sparse on technical specifics — exact traffic volumes, crawl rates, and specific cost figures are not reported in the available material.