Only if you interpret statements as being binary logic.
"seem to be" carries semantic meaning here: I'm stating my interpretation of the situation based on data we have available (which is limited) and my prior.
Put another way: "We can't say that for sure, but my money is on it not being a simple case of intellectual property theft"
Imagine if they broke it down to each distinct source, that'd be several billion cases of copyright infringement (though it's going to be determined by what courts think and that often comes down to "who can afford the best lawyers" in practice if not intent).
Apparently if I use lib-gen, that's copyright infringement and I'm exposed to legal risk but it seems fine to download all of it if your intent is "train an AI" so far.
Other people are allowed to have their priors, too. Even under a non-informative prior, the weight of the evidence (Tristan's account, OpenAI's announcement, and Bubeck's "denial", if you want to call it that, plus multiple other mathematicians coming forward with similar experiences) pushes the probability mass toward some degree of impropriety.
What exact evidence are you incorporating into your prior to come out with this posterior?
We are giving you an opportunity to correct yourself. You are instead trying to make your nonsensical statement make sense. Not only does the first part of your sentence literally contradict the second part:
> We don't have enough accurate knowledge to say [one way or the other], and it doesn't seem to be the case at all [based on our incomplete knowledge].
But it is in no way equivalent to this:
> We can't say that for sure, but my money is on it not being a simple case of intellectual property theft
Based on: my experience working on AI for 32 years, including a decade at Google including working on large-scale model training systems that used user data and complied with various user policies around data retention, along with a few decades working in science/tech making decisions around ambiguous data.
In short, I have a well-tuned intuition and a huge set of priors, and applied them to the limited knowledge we have about this situation.
The human mathematicians didn't solve the Navier-Stokes problem, they solved the Euler problem. And they were extensively using LLMs to drive the work, as described in the Buckmaster statement.
Any way you cut it, this is a major achievement for AI, besotted with human drama over whose prompt should be recognized by the history books.
I don't think we should assume a millenium puzzle has been solved, yet. Astra showed impressive capacity for cheating when it was faced with impossible cybersecurity challenges. It seems equally plausible at this stage that it's found a bug in Lean.
Xerox is incidentally a really good example, because precisely nobody ended up using the desktop experience Xerox made. They ended up using the desktop experience that Microsoft and Apple made and shipped while Xerox the actual company faded and memory of those original parc research teams faded into obscurity.
Incentives are one thing, even adjusting for them it's huge, and I don't understand this incentive play for only openai, academics have perverse incentives too, to overreport, overclaim, publication bias etc why are we scrutinizing AI industry to such high degree when they have demonstrated capability and often times are off by a model release at worst.
I swear there's nobody blinder than those who won't see.