ARC-AGI takes it very seriously: they've never tested Fable, because they won't run on the eval set without ZDR.
Another possible way to look at this is that any benchmark reporting Fable scores is potentially contaminated.
ARC-AGI takes it very seriously: they've never tested Fable, because they won't run on the eval set without ZDR.
Another possible way to look at this is that any benchmark reporting Fable scores is potentially contaminated.