There’s an Amazon warehouse where books end up in pieces for the AI: employee revelations

Written by Jason Miller

404 Media’s crusade against book-destroying AI continues. After the investigation published in mid-August, an Amazon employee revealed further details about the company’s modus operandi. The company, as previously revealed by the investigation and now by the man interviewed, who remained anonymous, scans and then destroys thousands of books inside a Las Vegas warehouseturning them into training material for their artificial intelligence systems.

The structure is called VGT3 and shares the complex with LAS8the center where Amazon manages the printing on demand of books. Here the volumes arrive on pallets, are stripped of their covers by automatic machinery and the loose pages end up on rows of high-speed scanners.

Among the volumes that ended up in the chain were new and used books, library copies, British government documents and publications in German, Russian and Japanese. The discovery came after a bookseller, suspecting the sale to an anonymous AI company, hid a tracker in a shipment of rare volumes that ended up right inside the Las Vegas warehouse.

Why rare books are tempting for AI, Amazon is not the only one

The warehouse workers, according to the testimony collected, were initially told that the operation was to prepare content for the devices Kindle. The employee does not hide the discomfort of a job that eliminates copies that are difficult to replace: “I don’t like it, because these books can’t be reused”he declared, underlining how the operation reduces the overall availability of certain editions.

The employee described an industrial operation dedicated to massive destruction and scanning of books intended for the creation of dataset for training artificial intelligence systems. The interview, given anonymously because staff are not authorized to speak to the press, offers a first-hand look at a process that workers themselves struggle to understand.

According to the employee, they arrive every day boxes and large containers of books of all kinds: new texts still sealed, volumes discontinued from libraries, academic materials from the University of London and even government documents presented to the British Parliament. Books come first scanned for barcodes to eliminate duplicates, then sent to stations where a machine cuts the back with a single stroke of the blade. The loose pages are then scanned by approximately 20–25 machines similar to banknote counters, capable of quickly digitizing each sheet.

The employee describes a work environment chaotic and unorganizedwith procedures that change from day to day. Rejected or duplicate books are thrown in bulk in large containers, officially destined for suppliers, although the staff doubts whether they are actually recovered.

The workers were initially told that the operation was for Kindles, but the employee claims that the explanation was not credible due to publishing rights. Only later did he understand that the volumes they were being destroyed to power AI systems. A practice that judges negatively, as written above: many books would be rare or difficult to findand their destruction would reduce the number of copies available in the world.

Non-catalogue texts have particular value for training linguistic models: they are not available onlinethey don’t have never crossed automatic generation systems And reduce the risk of model collapsethe phenomenon whereby an AI trained on content already generated by another AI gets progressively worse. Amazon isn’t the only company hunting for this type of material.

The story is, in fact, part of a broader rush by artificial intelligence laboratories towards quality texts, which are increasingly scarce online. Publishers, writers and bibliophiles have reacted harshly, defining the physical destruction of the volumes as an excessive price to fuel models which, paradoxically, could one day return precisely those contents in the form of generated text.

Amazon, when asked about the matter, said it purchases the books “through commercial channels to improve the products and services that customers use.” A generic formula, which neither confirms nor denies in detail the specific use for the training of one’s artificial intelligence models, nor explains why the scan must pass through the physical destruction of each copy. The phenomenon is certainly not new, and we have addressed it several times in the past months:

  • Buying old books to destroy them: how AIs search for uncontaminated text
  • AI companies buy thousands of academic books and then destroy them: here’s why they do it

Jason Miller

I'm Jason Miller, and I've been passionate about technology and storytelling for over a decade. As a lead writer at Herald Editorials, I strive to bring clarity and creativity to complex tech topics. When I'm not writing, you'll find me exploring the latest gadgets or hiking in the great outdoors.