| |
'Comically bad' datasets used to train clinical models for stroke and diabetes
Researchers discovered that datasets hosted on Kaggle and used to train clinical prediction models for stroke and diabetes contain severe quality issues, including celebrity images, duplicates, and unverified patient data unsuitable for medical research. A stroke detection paper published in Scientific Reports, for example, used a dataset featuring Sylvester Stallone and other celebrities alongside images of Bell's palsy and children, leading to an editor's warning and potential retraction. The findings highlight a broader problem affecting potentially thousands of papers across online data repositories that use unreliable datasets without proper verification.
Read Full Article →
← More Science news