> We wanted to see why Uzbekistan didn’t jump out, so we reproduced it in our comment (Extended Data Fig 1). It turns out that Uzbekistan wasn’t even the biggest outlier, but that the version they had published had the axes cropped so you couldn’t see the outliers (see red boxes in our version). This seemed indicative of a different issue, which is why we documented it in the comment.
Cropping the chart to hide the outliers is so bad that I can't tell if they're incompetent or malicious. I wouldn't be surprised if this is the kind of thing an LLM would produce in the hands of an operator not paying too much attention, but the paper was published in the time period before LLMs were everywhere in publishing.
We need something akin to the international geophysical year, but for data integrity. Make it an interdisciplinary priority to clean house and root out papers that are hanging by a thread of included / excluded outliers, biased samples, and outright fraud. It would be humbling, but we'd be in much better shape afterwards.
The article suggests it's unreasonable numbers in the original Uzbekistan data source and that other datapoints may have been worse, the authors just didn't correctly execute their basic checks.
"It turns out that Uzbekistan wasn’t even the biggest outlier, but that the version they had published had the axes cropped so you couldn’t see the outliers..."
> We wanted to see why Uzbekistan didn’t jump out, so we reproduced it in our comment (Extended Data Fig 1). It turns out that Uzbekistan wasn’t even the biggest outlier, but that the version they had published had the axes cropped so you couldn’t see the outliers (see red boxes in our version). This seemed indicative of a different issue, which is why we documented it in the comment.
Cropping the chart to hide the outliers is so bad that I can't tell if they're incompetent or malicious. I wouldn't be surprised if this is the kind of thing an LLM would produce in the hands of an operator not paying too much attention, but the paper was published in the time period before LLMs were everywhere in publishing.
Makes me wonder how many more papers out there have hard-to-pin-down errors like that
And how useful potentially AI could be to spot those (even if retrospectively)
Good on them for the retraction. It's good to see science at work.
We need something akin to the international geophysical year, but for data integrity. Make it an interdisciplinary priority to clean house and root out papers that are hanging by a thread of included / excluded outliers, biased samples, and outright fraud. It would be humbling, but we'd be in much better shape afterwards.
I'm confused as to what the actual issue was. What was the data for which Uzbekistan was the outlier and why?
The article suggests it's unreasonable numbers in the original Uzbekistan data source and that other datapoints may have been worse, the authors just didn't correctly execute their basic checks.
"It turns out that Uzbekistan wasn’t even the biggest outlier, but that the version they had published had the axes cropped so you couldn’t see the outliers..."
> had the axes cropped so you couldn’t see
Almost sounds intentional...
> This seemed indicative of a different issue, which is why we documented it in the comment.
Yea, that different issue is fraud.
My read is that the model had too much variance; more regularization was needed.