Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

Its possible. More generally, its also possible the human graders were doing a bad job; the ML system can only learn 'essay quality' to the extent that the training data reflects it.

However, the kaggle supplied 'straw-man' benchmark, which worked solely based on the count of characters and words in the essay, had an score of .647 with the training data. (The score metric used isnt trivial to interpret - it was 'Weighted Mean Quadratic Weighted Kappa' - but for reference the best entries had a score of ~.8 at the end)

The score of .647, just using length, is quite high. For length to have this powerful a causal predictive effect, the human graders would have to be weighting for length, as a feature, very heavily.

I can't rule that out; but I think its highly likely a major component of the predictive effect of length was correlative, rather than causal.



Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: