The Claim
A high feature-to-sample ratio in machine learning models is strongly associated with an increased risk of overfitting, particularly when non-robust validation methods such as K-fold cross-validation are used, because higher ratios increase the likelihood of detecting spurious correlations in random noise, leading to inflated performance estimates.
What the research says
Not yet evaluated
We are still looking at what the research says.
These are independent scores, not a percentage. Higher-grade studies count more, so a single strong opposing study can outweigh several weaker ones.
If a machine learning model has more features (like traits or measurements) than data points (like people or samples), it can 'cheat' by finding fake patterns in random noise, especially when tested the wrong way — making it look better than it really is.
See the scientific wording
The feature-to-sample ratio is a strong indicator of overfitting risk in machine learning models, with higher ratios (e.g., more features than samples) leading to increasingly inflated performance estimates when using non-robust validation methods like K-fold cross-validation, due to the increased probability of finding spurious correlations in noise.
What the research says
1 studyStudy: Machine learning algorithm validation with a limited sample size
The study shows that when you have more features than samples, especially in small studies, regular cross-validation can trick you into thinking your model works better than it really does. Using better testing methods fixes this problem.
Score breakdown, mechanism chain, raw evidence, ideal studies needed & 1 supporting studies
Not medical advice. For informational purposes only. Always consult a qualified healthcare professional before making health decisions.