Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

As much as our brain contents are unlicensed copies to the extent we can reproduce copyrighted work: If the model can recite copyrighted portions of text used in training, the model weights are a derivative work. Because the weights obviously must encode the original work. Just because lossy compression was applied the original work should still be considered present as long as it's recognizable. So the weights may not be published without license. Seems rather straightforward to me and I do wonder how Meta thinks they get around this.

Now if the likes of Openai and Google keep the model weights private and just provide generated text, they can try to filter for derivative works, but I don't see a solution that doesn't leak. If a model can be coaxed into producing a derivative work that escapes the filter, then boom, unlicensed copy was provided. If I tell the model to mix two texts word by word, what filter could catch this? What if I tell the model to use a numerical encoding scheme? Or to translate into another language? For example assuming the model knows a bunch of NYT articles by heart, as was already demonstrated: If have it translate one of those articles to French for me, that's still an unlicensed copy!

I can see how they will try to get these violations legalized like the DMCA safe-harbored things, but at the moment they are the ones generating the unlicensed versions and publishing them when prompted to do so.



Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: