the nature of the test was to see if the models can effectively compose an image of a novel concept outside the training set. If they are trained on it, it ceases to be an interesting test to some extent.
I would urge you to re-read the blog post you are commenting on. It pretty clearly explains how it is an interesting test independently of "see[ing] if the models can effectively compose an image of a novel concept outside the training set".
it's still interesting because there's no pelican-on-bike model, and if you're training a model well enough, then it should be obvious when a model has reached "AGI" or whatever.
The author primarily talks about the compression–prediction equivalence and also provides some working code linked in Github https://github.com/nathan-barry/gzipt
Every prediction model is inherently a compressor, and all compression algorithms are prediction models.
Reference:
Language Modeling Is Compression — Delétang et al., DeepMind, 2023. The prediction-compression equivalence, with the Chinchilla-beats-PNG result.
It will always be running my first local model and seeing its responses. A close second is watching the full thought traces of DeepSeek as this was and is still censored by major closed labs.
reply