M1inki
It has long been no secret to my readers that the copyright law appears to me as an entity that is hopelessly obsolete and extremely expensive in terms of enforcement, and therefore will immediately wither away as soon as the state ceases to carry out this very enforcement. However, with the appearance of large language models, all grounds have emerged to believe that copyright will wither away even and significantly sooner than the state.
There are various arguments in defense of copyright. There is the utilitarian one: if everyone takes someone else’s content for free, the incentive to compose this content will vanish, and that is the end of progress. There is the objectivist one: a person has a right to the product of their mind, and damn it, because that’s what Ayn Rand said. There is the ethical one: labor must be paid for.
And then it turns out that to train another, more advanced version of a large language model, it needs to be fed the entire corpus of texts ever created by humanity, and it will devour them and ask for seconds. If this is not done, then a promising direction of information technology development will hit a dead end. Those training the model are not concerned with the artistic merits of the text. For them, whether it is the Bible or a Harry Potter fanfic—it is simply text possessing a key quality: written by a human. Knowledge of the Bible is more important for a language model, but only for the reason that, firstly, it is more voluminous, and secondly, references to it are far more frequent in other texts and in user queries.
And that’s it, we have immediately lost the utilitarian argument, because it is precisely for the sake of progress that texts need to be fed to the neural network. In exchange, we have acquired very wealthy lobbyists for the abolition of copyright, since companies developing large language models can quite well compete in terms of capital size with digital publishing houses. Moreover, even if for particularly widely circulated texts the rightsholders manage to defend the right not to acquaint LLMs with them, who wins from the fact that in response to “summarize such-and-such story by Stephen King” the chatbot will answer “I don’t know a single work by this author, let me instead summarize such-and-such story by David Friedman”? Surely it won’t be Stephen King’s publishers nor Stephen King himself who will benefit from such a thing.
How will content for training neural networks be obtained? Very simple. Content is usually already physically present on the web in free access; it just has a “pirated” label, and therefore law-abiding companies have to pretend it doesn’t exist. However, this is too flimsy a little fence, and it won’t stop those who consider themselves entitled to step over it. The only person a neural network won’t get is some Rodion Belkovich, because he in principle does not post his books on the web, even for money))) This is a huge loss for the future interlocutors of intellectual chatbots, but what can be done, humanity has lost many valuable texts throughout its history, shit happens.
