As the first court in Germany, the Hamburg Regional Court (‘Landgericht Hamburg‘) ruled on Artificial Intelligence whether datasets used for AI training activities may infringe German copyright law (Judgment as of 27 September 2024 – file no. 310 O 227/23).

Background

The plaintiff is a photographer who made one of his photos freely available to the public via a photo agency’s website, subject to the following restrictions:

“ […] you may not: […]

  1. Use automated programs, applets, bots or the like to access the…com website or any content thereon for any purpose, including, by way of example only, downloading Content, indexing, scraping or caching any content on the website.

The defendant is a non-profit organization dedicated to promoting research activities in AI. The defendant creates and provides open datasets consisting of text-image pairs ready for use to train generative AI. By way of matching a large number of texts and images, AI is able to learn how people, animals or objects look like, based on which AI can be empowered to distinguish between and artificially create people, animals or objects by itself.

For the purpose of AI training activities, defendant downloaded and stored a copy of a variety of pictures from publicly available resources. The defendant runs a software over these copies which analyses as to whether the content of each photo matches its particular description. In case of a mismatch, the software eliminates the particular picture from the dataset as the latter needs to be reliable and consistent to ensure proper AI training. The defendant’s dataset comprises nearly six billion text-image pairs including one photo created by the plaintiff.

The plaintiff requested the defendant to stop any reproducing activities as to plaintiff’s photo in question.

The court’s decision

At a glance

The court dismissed the claim, essentially because of the following reasons:

  • According to Section 15 para. 1 no. 1 of the German Copyright Act (“Urheberrechtsgesetz” – UrhG), it is the plaintiff in its position as the author of the photo in question who holds exclusive rights of reproduction. ʼRight of reproductionʼ means the right to produce copies of that particular photo, whether on a temporary or on a permanent basis (Section 16 para. 1 UrhG) which can only be waived by way of (i) plaintiff’s consent or (ii) by statutory exception. In absence of plaintiff’s explicit consent, the court had to deal with the question whether Section 44b para. 2 sent. 1- and/or Section 60d para. 1 UrhG, as a statutory exception for text and data mining for scientific purposes, applies to the benefit of the defendant. The court answered this question in the affirmative.
  • Section 60d para. 1 UrhG applies to ʼtext and data miningʼ in terms of Section 44b para. 1 UrhG. Section 44b para. 1 UrhG specifies ʼtext and data miningʼ as an automatic analysis of datasets by automated means for the purpose of gathering information, in particular patterns, trends and/or correlations. The court holds the opinion that this definition also encompasses text and data mining activities for the purpose of AI-training activities.
  • Section 60d para. 2 sent. 1 and 2 UrhG sets out that, inter alia, research organizations shall be entitled to make reproductions for text and data mining activities for the purpose of scientific research. Research organizations comprise universities, research institutes, but also other institutions being active in scientific research if they (i) pursue non-commercial purposes, (ii) reinvest all their profits in scientific research or (iii) act in the public interest based on a state-approved mandate. The court decided that defendant can invoke on this statutory waiver as, according to the court, defendant pursues a scientific purpose in terms of Section 60d para. 2 sent. 1 and 2 UrhG. Furthermore, the court assumed a non-commercial purpose to be given as the plaintiff was not in the position to reasonably demonstrate particular influence of a commercial third party which may prevent defendant to invoke the statutory exception (Section 60d para. 2 sent. 3 UrhG).
Further background and detail

AI-enthusiasts have a particular interest in collecting as many material suitable to train AI as possible. Opulent databases of text and image pairs provide large-scaled opportunities for effective AI trainings and promise better and more reliable outputs. In order to pull together huge amounts of training data, web scraping seems to be a pragmatic option. It means that a software automatically scans through websites available in the public domain, reproduces protected material such as photos, videos, sounds or alike and compile them in a dataset for AI training purposes. These activities are at risk to potentially conflict existing copyright laws because, according to the basic principle set out in Section 15 para. 1 no. 1 UrhG, it is the exclusive right of the particular author to reproduce his/her work.

The court had to decide whether the defendant can successfully invoke Sections 44b para. 2 and/or 60d para. 1 UrhG as a statutory exception for systematic reproduction in the context of creating datasets suitable and ready to run AI training sessions. The court came to the conclusion that this is the case:

Key Take-aways
  • The case is not dealing with AI-trainings as such, but preparation of datasets ready and suitable to run AI training sessions.
  • The court holds the opinion that “scientific research” in terms of Section 60d para. 1 UrhG means methodical and systematic pursuit of new knowledge in general without direct production of such being required.
  • The burden of proof for commercial influence according to Section 60d para. 2 sent. 3 UrhG preventing a party to invoke the statutory waiver for scientific research is with the rights holder.
  • The court’s decision is open for appeal to the Hanseatic Higher Regional Court of Hamburg (ʼOberlandesgericht Hamburgʼ). Since Section 60d UrhG is based on Art. 3 and 2 lit. 1 of the DSM-Directive the case or at least similar disputes are at risk also to be referred to the European Court of Justice for request for preliminary ruling. In the meantime, the first serve in the match between copyright and AI has been made and rights holders should definitely pay attention. EU-copyright law does, with the good intention of fostering innovation, provide exceptions and limitations for text and data mining. These exceptions can allow for free reproduction of protected material in datasets that are subsequently used for AI-training. Certainly, this only represents one of many scenarios concerning use of copyright protected material by intelligent systems.