2026-09-11 | Data Protection

Copyright Limitations of Text and Data Mining
The basis of training artificial intelligence is so-called text and data mining, during which texts and data in digital form are analysed using automated analytical methods. Although the technology is a driving force of innovation, it encounters serious copyright obstacles. EU legislation distinguishes between the purposes of use: the Directive provides a mandatory exception for research organisations and cultural heritage institutions for text and data mining carried out for the purposes of scientific research. General, commercial-purpose data mining (such as the training of a commercial AI model), however, is subject to stricter conditions. This general exception may only be applied if the right holders have not expressly excluded the use. In the case of content made publicly available on the internet, authors, publishers and right holders may also reserve their rights by machine-readable means – for example, through metadata or the terms of use of the website. If the right holders exercise this possibility of prohibition, AI laboratories may not use the works for training their models without permission.
Processing of Personal Data and the Right to Be Forgotten
Language models, however, process not only copyright-protected works but also enormous amounts of personal data. During the training process and user interactions, the strict rules of the General Data Protection Regulation (GDPR) also apply. The principles require that the processing of personal data must in all cases be carried out lawfully, fairly and in a fully transparent manner in relation to the data subject. The processing of personal data is considered lawful only if the data subject has given his or her freely given, specific, informed and unambiguous consent to it, or if the processing has another established legal basis (for example, the legitimate interests of the controller). For developers, the greatest technological nightmare is the “right to be forgotten”, that is, the right to erasure: the data subject is entitled to request the deletion of data concerning him or her and the termination of further data processing. This right applies in particular where the data subject withdraws the consent given to the processing, or where the processing of personal data otherwise does not comply with the requirements of the Regulation. From an engineering perspective, however, retrospectively making a trained neural network “forget” a specific item of personal data is an almost impossible task.
Corporate Risks: What Should Companies Pay Attention to During Everyday Use?
ChatGPT contains legal pitfalls not only for developers but also for companies using it. During everyday business operations, employees often upload trade secrets, client data or draft contracts to AI platforms so that the program can assist with drafting, translation or analysis. By doing so, however, the company – as controller – may easily breach the GDPR principles relating to purpose limitation and data security, since some platforms may also use the data entered for the system’s own learning. In order to avoid personal data breaches and violations of confidentiality obligations (NDAs), it is essential to create an internal corporate AI policy that precisely specifies what data may be uploaded to open artificial intelligence systems.

