If it happened and you get to court, it's probably fairly simple to prove in court (subpoena, discovery, etc). It's not like training will happen only once and then they'll delete the data and all references to it. When training models, you want to have a full lineage of how you obtained it.
Not really, that’s what discovery is for. You just ask Anthropic and OpenAI to hand over every internal document and message that contains your companies name, or matches any reasonable query about where training data comes from.
It very hard for a company to do anything without leaving some kind of paper trail behind that can be discovered in court. Not without crippling their own operations by simply refusing to digitise or write down anything.
Lots of companies filled with people working very hard to obscure their shady practices have been hoisted by their own internal docs. Just look at any major Uber, Google, Apple, Microsoft lawsuit. Do really think Anthropic and OpenAI are gonna be better at destroying their paper trail before the lawsuit starts?
If they did not use my data for training, they should permit me to run the first 1-3 layers of the model locally and send them dense hidden state vectors. In my experience these compress very nicely without much effort.
It can be very hard to prove in court that your data was used for training