Every author group I hang out in runs the same ritual every week.
Someone drops a screenshot of their copyright page into the chat that’s some variation of: “No part of this book may be used to train artificial intelligence.” Forty people pile in with “Stealing this!” Everyone exhales and goes back to writing, feeling like they just hung a clove of garlic on the windowsill and the vampires will have to find another house.
I get it. But you won’t find that line in my books, and it isn’t because I’m reckless. We’re standing guard over a battlefield that was cleared months ago while the real action moved down the street.
The fight over AI training is mostly over. Your new book doesn’t need to be in a model’s training data to show up inside an AI answer today.
While authors were pasting anti-training disclaimers into Vellum, modern AI moved on to retrieval. That isn’t a stealth attack or a legal loophole. It’s a permission slip that has had your signature on it since the day you clicked “Agree” on your retailer dashboard. We all did.
I covered the authorship side of this back in June for Indie Author Magazine: who owns AI-assisted drafts, what the Copyright Office actually wants disclosed, and how to register your work properly. This is the other half of the story, which is where training law actually landed, how retrieval works when a reader asks a chatbot about your plot, and why the publishing contracts we already signed cover it.
The Training War Has a Scoreboard (and Models Don’t Hold Your Manuscript)
The copyright-page disclaimer rests on the idea that an AI model swallows your book whole, keeps the PDF on a hard drive somewhere, and regurgitates it on request like a well-read cat.
The courts spent two years taking that idea apart.
Training leaves patterns, not files. In Bartz v. Anthropic (2025), the court ruled that training models on lawfully acquired books is transformative fair use. The machine reads your text once alongside millions of other books, nudges a few billion mathematical weights by a hair, and discards the original. In Kadrey v. Meta (2025), top technical experts couldn’t force Llama to cough up more than 50 consecutive words of the plaintiffs’ books. Your villain’s monologue is not sitting in a database waiting to be lifted.
The $1.5B Anthropic settlement was about piracy, not training. When final approval landed in July 2026, authors got checks of roughly $3,000 per registered title because Anthropic downloaded books from torrent sites like LibGen. In that same case, the judge said plainly that buying physical paperbacks, shredding them, scanning them, and training on them was legal fair use. The pirated files were the crime, and the training wasn’t.
New books aren’t in training sets anyway. Unless your novel is quoted word-for-word on thousands of websites the way Harry Potter is (which hit roughly 42% memorization in stress tests), a trained model holds zero recoverable manuscript text from your work. Your debut cozy mystery is not, statistically speaking, Harry Potter.
So the training fight is scored, and your brand-new book can still show up in a detailed AI answer tomorrow.
The Real Engine: From Training to Retrieval
While everyone was arguing about training datasets, the tech world quietly shifted to Retrieval-Augmented Generation, or RAG, because nobody in tech has ever named anything well.
Training is a massive, static event done up front. Retrieval happens live, the moment a user types a prompt.
When a reader asks an AI tool a specific question about your book, the system doesn’t rummage through the model’s memory. It fetches the relevant pages directly from a live file or database, hands those paragraphs to the AI as a temporary cheat sheet, writes a summary from that cheat sheet, and then forgets it ever saw the thing.
Nothing was trained and no model learned your plot points, but your text was retrieved, processed, and served on demand like room service for chatbots.
You see this on the open web every day with Google AI Overviews and Perplexity. For your ebooks, though, the database being searched isn’t scraped off the open web by some random crawler. It lives on the retailer server where you uploaded it, on purpose, with a cover and a blurb attached.
You Already Authorized the Index
Your manuscript isn’t floating around for web bots to trip over. It sits in retailer databases, and retailers don’t need to crawl your site because they already hold the pristine, high-res file you handed them under your distribution contract.
Go open Section 5.5 of the KDP Terms and Conditions, which I promise is riveting. In exchange for putting your book in the store, you granted Amazon an irrevocable, nonexclusive license to:
“reproduce, index and store Books... and reformat, convert and encode Books.”
Index. Store. Convert. Encode.
That clause wasn’t written to pull a fast one on anyone. It was written so Amazon could host your file, run standard search, build “Look Inside” previews, and format your book for different devices. But under the hood, interactive tools like Ask This Book on Kindle aren’t relying on AI training or fair use at all. They’re RAG retrieval with a conversational interface bolted on.
Amazon’s own technical breakdown says as much. The feature uses your manuscript purely as temporary prompt context, it’s classified as an extension of standard search functionality, and it runs as an always-on feature across the store with no author opt-out.
Amazon didn’t need to ask us for new AI training permissions because retrieval is indexing, and we all said yes to indexing the day we published our first title.
Fighting the Right Problem
Back in 2023, the Authors Guild drafted a model clause reserving the right to “reproduce and/or use the Work for purposes of training artificial intelligence technologies.” Thousands of indies pasted it onto page two, and traditional publishers slotted it in beside their standard copyright notices.
That clause was built for 2023. It addresses whether a model can ingest text to learn up front, and it says nothing about real-time search, indexing, or retrieval prompts, because RAG wasn’t what anyone was staring at when it was drafted. It’s a good lock on a door nobody uses anymore.
Natural-language disclaimers don’t override fair use in the US anyway. Over in Europe, the Higher Regional Court of Hamburg (OLG Hamburg, Dec 2025) ruled that human-readable copyright notes don’t satisfy the legal requirement for machine-readable opt-outs. So in at least one major jurisdiction, the sentence isn’t even legible to the thing it’s aimed at.
Where That Leaves Us
This isn’t a call to panic or yank your books off retail shelves. It’s a call to see the board clearly so we stop spending energy on disclaimers whose main job is making us feel better.
Training law has settled boundaries. Training on legally acquired copies leans hard toward transformative use, while sourcing from pirate archives remains flatly illegal.
Retrieval is modern search. Tools that answer questions about your book in real time run on database permissions and search indexing rather than model training.
Standard retail terms already cover it. Distribution agreements like KDP §5.5 grant retailers the right to index files for interactive search, and you agreed to that long before anyone said “chatbot” out loud.
Direct sales are the one place you hold the pen. If you don’t want a platform indexing your files or running interactive search against your text, selling direct on your own site is the only space where you write the license terms yourself. Everywhere else you’re a guest.
We don’t need to keep re-fighting the 2023 war over whether a model “swallowed” our words. We just need to understand that when we hit publish, we agreed to let retailers index our files, and in 2026, search indexes talk back.
My original piece on copyright registration and ownership is over at Indie Author Magazine: AI Copyright for Indie Authors: Where We Stand Now. One update to that article: the $1.5B Anthropic settlement it references received final judicial sign-off on July 20, 2026.


