If they could declare with certainty that Buckminster's and Alpoge's usage data had been totally excluded from training, would that set a worse precedent and reflect poorly on their de-identification process (and data access safeguards moreover)?
This may sound like a charitable interpretation of OpenAI's remark, but consider that the lie would be (I think) impossible to falsify from the outside. They could easily just say "no sir we didn't peek" unless:
1. The conspiracy to peek at codex sessions involved enough people that the risk of one snitching is non-negligible
2. Lawyers advised it would be a bad idea to make such a remark, whether true or false
> If they could declare with certainty that Buckminster's and Alpoge's usage data had been totally excluded from training, would that set a worse precedent and reflect poorly on their de-identification process (and data access safeguards moreover)?
No; if they said "we can see that Tristan opted out of model improvement, therefore we are confident his work and ideas did not improve our model," that would be an excellent and reassuring precedent.
This requires keeping history if, at the time, Buckmaster's account had a certain flag set, because just because the account has the flag now doesn't mean it had the flag at a certain moment in the past. And even if they had such history, it's not obvious whether they just load all data as-is into training.
A totally reasonable pipeline may be unauditable for this purpose.
This may sound like a charitable interpretation of OpenAI's remark, but consider that the lie would be (I think) impossible to falsify from the outside. They could easily just say "no sir we didn't peek" unless:
1. The conspiracy to peek at codex sessions involved enough people that the risk of one snitching is non-negligible
2. Lawyers advised it would be a bad idea to make such a remark, whether true or false