4 L’IA (1) : La protection par la PI des inputs et outputs 4 L’IA (1) : La protection par la PI des inputs et outputs
4.1 US District Court, 18 Aout 2023, Thaler v. Perlmutter, No. 22-CV-384-1564-BAH 4.1 US District Court, 18 Aout 2023, Thaler v. Perlmutter, No. 22-CV-384-1564-BAH
UNITED STATES DISTRICT COURT FOR THE DISTRICT OF COLUMBIA
STEPHEN THALER,
Plaintiff,
v.
SHIRA PERLMUTTER, Register of Copyrights and Director of the United States Copyright Office, et al.
Defendants.
Civil Action No. 22-1564 (BAH) Judge Beryl A. Howell
MEMORANDUM OPINION
Plaintiff Stephen Thaler owns a computer system he calls the “Creativity Machine,” which he claims generated a piece of visual art of its own accord. He sought to register the work for a copyright, listing the computer system as the author and explaining that the copyright should transfer to him as the owner of the machine. The Copyright Office denied the application on the grounds that the work lacked human authorship, a prerequisite for a valid copyright to issue, in the view of the Register of Copyrights. Plaintiff challenged that denial, culminating in this lawsuit against the United States Copyright Office and Shira Perlmutter, in her official capacity as the Register of Copyrights and the Director of the United States Copyright Office
(“defendants”). Both parties have now moved for summary judgment, which motions present the sole issue of whether a work generated entirely by an artificial system absent human involvement should be eligible for copyright. See Pl.’s Mot. Summ. J. (Pl.’s Mot.”), ECF No. 16; Defs.’ Cross-Mot. Summ. J. (“Defs.’ Mot.”), ECF No. 17. For the reasons explained below, defendants are correct that human authorship is an essential part of a valid copyright claim, and therefore plaintiff’s pending motion for summary judgment is denied and defendants’ pending cross-motion for summary judgment is granted.
I. BACKGROUND
Plaintiff develops and owns computer programs he describes as having “artificial intelligence” (“AI”) capable of generating original pieces of visual art, akin to the output of a human artist. See Pl.’s Mem. Supp. Mot. Summ. J. (“Pl.’s Mem.”) at 13, ECF No. 16. One such AI system—the so-called “Creativity Machine”—produced the work at issue here, titled “A Recent Entrance to Paradise:”
Admin. Record (“AR”), Ex. H, Copyright Review Board Refusal Letter Dated February 14, 2022
“(Final Refusal Letter”) at 1, ECF No. 13-8.
After its creation, plaintiff attempted to register this work with the Copyright Office. In his application, he identified the author as the Creativity Machine, and explained the work had been “autonomously created by a computer algorithm running on a machine,” but that plaintiff sought to claim the copyright of the “computer-generated work” himself “as a work-for-hire to the owner of the Creativity Machine.” Id., Ex. B, Copyright Application (“Application”) at 1,
ECF No. 13-2; see also id. at 2 (listing “Author” as “Creativity Machine,” the work as “[c]reated autonomously by machine,” and the “Copyright Claimant” as “Steven [sic] Thaler” with the
transfer statement, “Ownership of the machine”). The Copyright Office denied the application on the basis that the work “lack[ed] the human authorship necessary to support a copyright claim,” noting that copyright law only extends to works created by human beings. Id., Ex. D, Copyright Office Refusal Letter Dated August 12, 2019 (“First Refusal Letter”) at 1, ECF No. 13-4.
Plaintiff requested reconsideration of his application, confirming that the work “was autonomously generated by an AI” and “lack[ed] traditional human authorship,” but contesting the Copyright Office’s human authorship requirement and urging that AI should be “acknowledge[d] . . . as an author where it otherwise meets authorship criteria, with any copyright ownership vesting in the AI’s owner.” Id., Ex. E, First Request for Reconsideration at 2, ECF No. 13-5. Again, the Copyright Office refused to register the work, reiterating its original rationale that “[b]ecause copyright law is limited to ‘original intellectual conceptions of the author,’ the Office will refuse to register a claim if it determines that a human being did not create the work.” Id., Ex. F, Copyright Office Refusal Letter Dated March 30, 2020 (“Second Refusal Letter”) at 1, ECF No. 13-6 (quoting Burrow-Giles Lithographic Co. v. Sarony, 111 U.S. 53, 58 (1884) and citing 17 U.S.C. § 102(a); U.S. Copyright Office, Compendium of U.S. Copyright Office Practices § 306 (3d ed. 2017)). Plaintiff made a second request for reconsideration along the same lines as his first, see id., Ex. G, Second Request for Reconsideration at 2, ECF No. 13-7, and the Copyright Office Review Board affirmed the denial of registration, agreeing that copyright protection does not extend to the creations of non-human entities, Final Refusal Letter at 4, 7.
Plaintiff timely challenged that decision in this Court, claiming that defendants’ denial of copyright registration to the work titled “A Recent Entrance to Paradise,” was “arbitrary, capricious, an abuse of discretion and not in accordance with the law, unsupported by substantial evidence, and in excess of Defendants’ statutory authority,” in violation of the Administrative
Procedure Act (“APA”), 5 U.S.C. § 706(2). See Compl. ¶¶ 62–66, ECF No. 1. The parties agree upon the key facts narrated above to focus, in the pending cross-motions for summary judgment, on the sole legal issue of whether a work autonomously generated by an AI system is copyrightable. See Pl.’s Mem. at 13; Defs.’ Mem. Supp. Cross-Mot. Summ. J. & Opp’n Pl.’s Mot. Summ. J. (“Defs.’ Opp’n”) at 7, ECF No. 17. Those motions are now ripe for resolution. See Defs.’ Reply Supp. Cross-Mot. Summ. J. (“Defs.’ Reply”), ECF No. 21.
II. LEGAL STANDARD
- Administrative Procedure Act
The APA provides for judicial review of any “final agency action for which there is no other adequate remedy in a court,” 5 U.S.C. § 704, and “instructs a reviewing court to set aside agency action found to be ‘arbitrary, capricious, an abuse of discretion, or otherwise not in accordance with law,’” Cigar Ass’n of Am. v. FDA, 964 F.3d 56, 61 (D.C. Cir. 2020) (quoting 5 U.S.C. § 706(2)(A)). This standard “‘requires agencies to engage in reasoned decisionmaking,’ and . . . to reasonably explain to reviewing courts the bases for the actions they take and the conclusions they reach.” Brotherhood of Locomotive Eng’rs & Trainmen v. Fed. R.R. Admin., 972 F.3d 83, 115 (D.C. Cir. 2020) (quoting Dep’t of Homeland Sec. v. Regents of Univ. of Cal. (“Regents”), 140 S. Ct. 1891, 1905 (2020)). Judicial review of agency action is limited to “the grounds that the agency invoked when it took the action,” Regents, 140 S. Ct. at 1907 (quoting Michigan v. EPA, 576 U.S. 743, 758 (2015)), and the agency, too, “must defend its actions based on the reasons it gave when it acted,” id. at 1909.
B. Summary Judgment
Pursuant to Federal Rule of Civil Procedure 56, “[a] party is entitled to summary judgment only if there is no genuine issue of material fact and judgment in the movant’s favor is proper as a matter of law.” Soundboard Ass’n v. FTC, 888 F.3d 1261, 1267 (D.C. Cir. 2018) (quoting Ctr. for Auto Safety v. Nat'l Highway Traffic Safety Admin., 452 F.3d 798, 805 (D.C. Cir. 2006)); see also Fed. R. Civ. P. 56(a). In APA cases such as this one, involving cross-motions for summary judgment, “the district judge sits as an appellate tribunal. The ‘entire case’ on review is a question of law.” Am. Bioscience, Inc. v. Thompson, 269 F.3d 1077, 1083–84 (D.C. Cir. 2001) (footnote omitted) (collecting cases). Thus, a court need not and ought not engage in fact finding, since “[g]enerally speaking, district courts reviewing agency action under the APA’s arbitrary and capricious standard do not resolve factual issues, but operate instead as appellate courts resolving legal questions.” James Madison Ltd. by Hecht v. Ludwig, 82 F.3d 1085, 1096 (D.C. Cir. 1996); see also Lacson v. U.S. Dep’t of Homeland Sec., 726 F.3d 170, 171 (D.C. Cir. 2013) (noting, in an APA case, that “determining the facts is generally the agency’s responsibility, not [the court’s]”). Judicial review, when available, is typically limited to the administrative record, since “[i]t is black-letter administrative law that in an [APA] case, a reviewing court should have before it neither more nor less information than did the agency when it made its decision.” CTS Corp. v. EPA, 759 F.3d 52, 64 (D.C. Cir. 2014) (internal quotation marks and citation omitted).
III. DISCUSSION
Under the Copyright Act of 1976, copyright protection attaches “immediately” upon the creation of “original works of authorship fixed in any tangible medium of expression,” provided those works meet certain requirements. Fourth Estate v. Public Benefit Corporation v. Wall- Street.com, LLC, 139 S. Ct. 881, 887 (2019); 17 U.S.C. § 102(a). A copyright claimant can also register the work with the Register of Copyrights. Upon concluding that the work is indeed copyrightable, the Register will issue a certificate of registration, which, among other advantages, allows the claimant to pursue infringement claims in court. 17 U.S.C. §§ 410(a), 411(a); Unicolors v. H&M Hennes & Mauritz, L.P., 142 S. Ct. 941, 944–45 (2022). A valid copyright exists upon a qualifying work’s creation and “apart” from registration, however; a certificate of registration merely confirms that the copyright has existed all along. See Fourth Estate, 139 S. Ct. at 887. Conversely, if the Register denies an application for registration for lack of copyrightable subject matter—and did not err in doing so—then the work at issue was never subject to copyright protection at all.
In considering plaintiff’s copyright registration application as to “A Recent Entrance to Paradise,” the Register concluded that “this particular work will not support a claim to copyright” because the work lacked human authorship and thus no copyright existed in the first instance. First Refusal Letter at 1; see also Final Refusal Letter at 3 (providing the same rationale in the final reconsideration decision). By design in plaintiff’s framing of the registration application, then, the single legal question presented here is whether a work generated autonomously by a computer falls under the protection of copyright law upon its creation.
Plaintiff attempts to complicate the issues presented by devoting a substantial portion of his briefing to the viability of various legal theories under which a copyright in the computer’s work would transfer to him, as the computer’s owner; for example, by operation of common law property principles or the work-for-hire doctrine. See Pl.’s Mem. at 31–37; Pl.’s Reply Supp. Mot. Summ. J. & Opp’n Def.’s Cross-Mot. Summ. J. (“Pl.’s Opp’n”) at 11–15, ECF No. 18. These arguments concern to whom a valid copyright should have been registered, and in so doing put the cart before the horse.1 By denying registration, the Register concluded that no valid copyright had ever existed in a work generated absent human involvement, leaving nothing at all to register and thus no question as to whom that registration belonged.
The only question properly presented, then, is whether the Register acted arbitrarily or capriciously or otherwise in violation of the APA in reaching that conclusion. The Register did not err in denying the copyright registration application presented by plaintiff. United States copyright law protects only works of human creation.
Plaintiff correctly observes that throughout its long history, copyright law has proven malleable enough to cover works created with or involving technologies developed long after traditional media of writings memorialized on paper. See, e.g., Goldstein v. California, 412 U.S. 546, 561 (1973) (explaining that the constitutional scope of Congress’s power to “protect the ‘Writings’ of ‘Authors’” is “broad,” such that “writings” is not “limited to script or printed material,” but rather encompasses “any physical rendering of the fruits of creative intellectual or aesthetic labor”); Burrow-Giles Lithographic Co. v. Sarony, 111 U.S. 53, 58 (1884) (upholding the constitutionality of an amendment to the Copyright Act to cover photographs). In fact, that malleability is explicitly baked into the modern incarnation of the Copyright Act, which provides that copyright attaches to “original works of authorship fixed in any tangible medium of expression, now known or later developed.” 17 U.S.C. § 102(a) (emphasis added). Copyright is designed to adapt with the times. Underlying that adaptability, however, has been a consistent understanding that human creativity is the sine qua non at the core of copyrightability, even as that human creativity is channeled through new tools or into new media. In Sarony, for example, the Supreme Court reasoned that photographs amounted to copyrightable creations of “authors,” despite issuing from a mechanical device that merely reproduced an image of what is in front of the device, because the photographic result nonetheless “represent[ed]” the “original intellectual conceptions of the author.” Sarony, 111 U.S. at 59. A camera may generate only a “mechanical reproduction” of a scene, but does so only after the photographer develops a “mental conception” of the photograph, which is given its final form by that photographer’s decisions like “posing the [subject] in front of the camera, selecting and arranging the costume, draperies, and other various accessories in said photograph, arranging the subject so as to present graceful outlines, arranging and disposing the light and shade, suggesting and evoking the desired expression, and from such disposition, arrangement, or representation” crafting the overall image. Id. at 59–60. Human involvement in, and ultimate creative control over, the work at issue was key to the conclusion that the new type of work fell within the bounds of copyright.
Copyright has never stretched so far, however, as to protect works generated by new forms of technology operating absent any guiding human hand, as plaintiff urges here. Human authorship is a bedrock requirement of copyright.
That principle follows from the plain text of the Copyright Act. The current incarnation of the copyright law, the Copyright Act of 1976, provides copyright protection to “original works of authorship fixed in any tangible medium of expression, now known or later developed, from which they can be perceived, reproduced, or otherwise communicated, either directly or with the aid of a machine or device.” 17 U.S.C. § 102(a). The “fixing” of the work in the tangible medium must be done “by or under the authority of the author.” Id. § 101. In order to be eligible for copyright, then, a work must have an “author.”
To be sure, as plaintiff points out, the critical word “author” is not defined in the Copyright Act. See Pl.’s Mem. at 24. “Author,” in its relevant sense, means “one that is the source of some form of intellectual or creative work,” “[t]he creator of an artistic work; a painter, photographer, filmmaker, etc.” Author, MERRIAM-WEBSTER UNABRIDGED DICTIONARY, https://unabridged.merriam-webster.com/unabridged/author (last visited Aug. 18, 2023); Author, OXFORD ENGLISH DICTIONARY, https://www.oed.com/dictionary/author_n (last visited Aug. 10, 2023). By its plain text, the 1976 Act thus requires a copyrightable work to have an originator with the capacity for intellectual, creative, or artistic labor. Must that originator be a human being to claim copyright protection? The answer is yes.2
The 1976 Act’s “authorship” requirement as presumptively being human rests on centuries of settled understanding. The Constitution enables the enactment of copyright and patent law by granting Congress the authority to “promote the progress of science and useful arts, by securing for limited times to authors and inventors the exclusive right to their respective writings and discoveries.” U.S. Const. art. 1, cl. 8. As James Madison explained, “[t]he utility of this power will scarcely be questioned,” for “[t]he public good fully coincides in both cases [of copyright and patent] with the claims of individuals.” THE FEDERALIST NO. 43 (James Madison). At the founding, both copyright and patent were conceived of as forms of property that the government was established to protect, and it was understood that recognizing exclusive rights in that property would further the public good by incentivizing individuals to create and invent. The act of human creation—and how to best encourage human individuals to engage in that creation, and thereby promote science and the useful arts—was thus central to American copyright from its very inception. Non-human actors need no incentivization with the promise of exclusive rights under United States law, and copyright was therefore not designed to reach them.
The understanding that “authorship” is synonymous with human creation has persisted even as the copyright law has otherwise evolved. The immediate precursor to the modern copyright law—the Copyright Act of 1909—explicitly provided that only a “person” could “secure copyright for his work” under the Act. Act of Mar. 4, 1909, ch. 320, §§ 9, 10, 35 Stat. 1075, 1077. Copyright under the 1909 Act was thus unambiguously limited to the works of human creators. There is absolutely no indication that Congress intended to effect any change to this longstanding requirement with the modern incarnation of the copyright law. To the contrary, the relevant congressional report indicates that in enacting the 1976 Act, Congress intended to incorporate the “original work of authorship” standard “without change” from the previous 1909 Act. See H.R. REP. NO. 94-1476, at 51 (1976).
The human authorship requirement has also been consistently recognized by the Supreme Court when called upon to interpret the copyright law. As already noted, in Sarony, the Court’s recognition of the copyrightability of a photograph rested on the fact that the human creator, not the camera, conceived of and designed the image and then used the camera to capture the image. See Sarony, 111 U.S. at 60. The photograph was “the product of [the photographer’s] intellectual invention,” and given “the nature of authorship,” was deemed “an original work of art . . . of which [the photographer] is the author.” Id. at 60–61. Similarly, in Mazer v. Stein, the Court delineated a prerequisite for copyrightability to be that a work “must be original, that is, the author’s tangible expression of his ideas.” 347 U.S. 201, 214 (1954). Goldstein v. California, too, defines “author” as “an ‘originator,’ ‘he to whom anything owes its origin,’” 412 U.S. at 561 (quoting Sarony, 111 U.S. at 58). In all these cases, authorship centers on acts of human creativity.
Accordingly, courts have uniformly declined to recognize copyright in works created absent any human involvement, even when, for example, the claimed author was divine. The Ninth Circuit, when confronted with a book “claimed to embody the words of celestial beings rather than human beings,” concluded that “some element of human creativity must have occurred in order for the Book to be copyrightable,” for “it is not creations of divine beings that the copyright laws were intended to protect.” Urantia Found. v. Kristen Maaherra, 114 F.3d 955, 958–59 (9th Cir. 1997) (finding that because the “members of the Contact Commission chose and formulated the specific questions asked” of the celestial beings, and then “select[ed] and arrange[d]” the resultant “revelations,” the Urantia Book was “at least partially the product of human creativity” and thus protected by copyright); see also Penguin Books U.S.A., Inc. v. New Christian Church of Full Endeavor, 96-cv-4126 (RWS), 2000 WL 1028634, at *2, 10–11 (S.D.N.Y. July 25, 2000) (finding a valid copyright where a woman had “filled nearly thirty stenographic notebooks with words she believed were dictated to her” by a “‘Voice’ which would speak to her whenever she was prepared to listen,” and who had worked with two human co-collaborators to revise and edit those notes into a book, a process which involved enough creativity to support human authorship); Oliver v. St. Germain Found., 41 F. Supp. 296, 297, 299 (S.D. Cal. 1941) (finding no copyright infringement where plaintiff claimed to have transcribed “letters” dictated to him by a spirit named Phylos the Thibetan, and defendant copied the same “spiritual world messages for recordation and use by the living” but was not charged with infringing plaintiff’s “style or arrangement” of those messages). Similarly, in Kelley v. Chicago Park District, the Seventh Circuit refused to “recognize[] copyright” in a cultivated garden, as doing so would “press[] too hard on the[] basic principle[]” that “[a]uthors of copyrightable works must be human.” 635 F.3d 290, 304–06 (7th Cir. 2011). The garden “ow[ed] [its] form to the forces of nature,” even if a human had originated the plan for the “initial arrangement of the plants,” and as such lay outside the bounds of copyright. Id. at 304. Finally, in Naruto v. Slater, the Ninth Circuit held that a crested macaque could not sue under the Copyright Act for the alleged infringement of photographs this monkey had taken of himself, for “all animals, since they are not human” lacked statutory standing under the Act. 888 F.3d 418, 420 (9th Cir. 2018). While resolving the case on standing grounds, rather than the copyrightability of the monkey’s work, the Naruto Court nonetheless had to consider whom the Copyright Act was designed to protect and, as with those courts confronted with the nature of authorship, concluded that only humans had standing, explaining that the terms used to describe who has rights under the Act, like “‘children,’ ‘grandchildren,’ ‘legitimate,’ ‘widow,’ and ‘widower[,]’ all imply humanity and necessarily exclude animals.” Id. at 426. Plaintiff can point to no case in which a court has recognized copyright in a work originating with a non-human.
Undoubtedly, we are approaching new frontiers in copyright as artists put AI in their toolbox to be used in the generation of new visual and other artistic works. The increased attenuation of human creativity from the actual generation of the final work will prompt challenging questions regarding how much human input is necessary to qualify the user of an AI system as an “author” of a generated work, the scope of the protection obtained over the resultant image, how to assess the originality of AI-generated works where the systems may have been trained on unknown pre-existing works, how copyright might best be used to incentivize creative works involving AI, and more. See, e.g., Letter from Senators Thom Tillis and Chris Coons to Kathi Vidal, Under Secretary of Commerce for Intellectual Property and Director of the U.S. Patent and Trademark Office, and Shira Perlmutter, Register of Copyrights and Director of the U.S. Copyright Office (Oct. 27, 2022), https://www.copyright.gov/laws/hearings/Letter-to- USPTO-USCO-on-National-Commission-on-AI-1.pdf (requesting that the United States Patent and Trademark Office and the United States Copyright Office “jointly establish a national commission on AI” to assess, among other topics, how intellectual property law may best “incentivize future AI related innovations and creations”).
This case, however, is not nearly so complex. While plaintiff attempts to transform the issue presented here, by asserting new facts that he “provided instructions and directed his AI to create the Work,” that “the AI is entirely controlled by [him],” and that “the AI only operates at [his] direction,” Pl.’s Mem. at 36–37—implying that he played a controlling role in generating the work—these statements directly contradict the administrative record. Judicial review of a final agency action under the APA is limited to the administrative record, because “[i]t is black- letter administrative law that in an [APA] case, a reviewing court should have before it neither more nor less information than did the agency when it made its decision.” CTS Corp., 759 F.3d at 64 (internal quotation marks and citation omitted). Here, plaintiff informed the Register that the work was “[c]reated autonomously by machine,” and that his claim to the copyright was only based on the fact of his “[o]wnership of the machine.” Application at 2. The Register therefore made her decision based on the fact the application presented that plaintiff played no role in using the AI to generate the work, which plaintiff never attempted to correct. See First Request for Reconsideration at 2 (“It is correct that the present submission lacks traditional human authorship—it was autonomously generated by an AI.”); Second Request for Reconsideration at 2 (same). Plaintiff’s effort to update and modify the facts for judicial review on an APA claim is too late. On the record designed by plaintiff from the outset of his application for copyright registration, this case presents only the question of whether a work generated autonomously by a computer system is eligible for copyright. In the absence of any human involvement in the creation of the work, the clear and straightforward answer is the one given by the Register: No.
Given that the work at issue did not give rise to a valid copyright upon its creation, plaintiff’s myriad theories for how ownership of such a copyright could have passed to him need not be further addressed. Common law doctrines of property transfer cannot be implicated where no property right exists to transfer in the first instance. The work-for-hire provisions of the Copyright Act, too, presuppose that an interest exists to be claimed. See 17 U.S.C § 201(b) (“In the case of a work made for hire, the employer . . . owns all of the rights comprised in the copyright.”).3 Here, the image autonomously generated by plaintiff’s computer system was never eligible for copyright, so none of the doctrines invoked by plaintiff conjure up a copyright over which ownership may be claimed.
IV. CONCLUSION
For the foregoing reasons, defendants are correct that the Copyright Office acted properly in denying copyright registration for a work created absent any human involvement. Plaintiff’s motion for summary judgment is therefore denied and defendants’ cross-motion for summary judgment is granted.
An Order consistent with this Memorandum Opinion will be entered contemporaneously.
Date: August 18, 2023
BERYL A. HOWELL
United States District Judge
Footnotes
1 In pursuing these arguments, plaintiff elaborates on his development, use, ownership, and prompting of the AI generating software in the so-called “Creativity Machine,” implying a level of human involvement in this case entirely absent in the administrative record. As detailed, supra, in Part I, plaintiff consistently represented to the Register that the AI system generated the work “autonomously” and that he played no role in its creation, see Application at 2, and judicial review of the Register’s final decision must be based on those same facts.
2 The issue of whether non-human sentient beings may be covered by “person” in the Copyright Act is only “fun conjecture for academics,” Justin Hughes, Restating Copyright Law’s Originality Requirement, 44 COLUMBIA J.L. & ARTS 383, 408–09 (2021), though useful in illuminating the purposes and limits of copyright protection as AI is increasingly employed. Nonetheless, delving into this debate is an unnecessary detour since “[t]he day sentient refugees from some intergalactic war arrive on Earth and are granted asylum in Iceland, copyright law will be the least of our problems.” Id. at 408.
3 In any event, plaintiff’s attempts to cast the work as a work-for-hire must fail as both definitions of a “work made for hire” available under the Copyright Act require that the individual who prepares the work is a human being. The first definition provides that “a ‘work made for hire’ is . . . a work prepared by an employee within the scope of his or her employment,” while the second qualifies certain eligible works “if the parties expressly agree in a written instrument signed by them that the work shall be considered a work made for hire.” 17 U.S.C. § 101 (emphasis added). The use of personal pronouns in the first definition clearly contemplates only human beings as eligible “employees,” while the second necessitates a meeting of the minds and exchange of signatures in a valid contract not possible with a non-human entity.
4.2 Hamburg District Court, 27 septembre 2024, LAION v. Robert Kneschke, 24 310 O.22723 4.2 Hamburg District Court, 27 septembre 2024, LAION v. Robert Kneschke, 24 310 O.22723
Official headnotes (translated from the German by David Wright-Policepayeh)
1. A reproduction is not deemed to be ephemeral if data is automatically deleted during an analysis process, but where the deletion does not occur ‘user-independently’ and instead as a result of a corresponding deliberate programming of the analysis process. Nor is downloading merely an undertaking process to the analysis performed if the data is to be analysed by specific software (para. 62).
2. An act of reproduction is subject to the exception provision of Sec. 44b(2) Copyright Act if it is carried out for the purpose of text and data mining. A teleological reduction of the exception provision because the European legislator was not yet aware of the ‘AI problem’ when the underlying directive provision was adopted in 2019 cannot be assumed. This is because the technical development does not concern data mining as such, but rather the performance of the artificial neural networks trained with the data. However, the exception provision does not apply where a reservation of use has been effectively declared (para. 66).
3. Reproductions for text and data mining are permitted for the purposes of scientific research by research organisations (Sec. 60d Copyright Act). Scientific research generally refers to the methodical and systematic pursuit of new knowledge. In particular, the term scientific research does not presuppose any subsequent success of the research. Whether the dataset is also used by commercial undertakings to train or further develop their AI systems is irrelevant because research by commercial undertakings is still also research. It is irrelevant whether an organisation also conducts scientific research in the form of developing its own AI models in addition to creating corresponding datasets. A non-commercial purpose already results from the fact that an organisation makes the dataset publicly available free of charge (para.103).
4. Admittedly, research organisations cannot invoke the exception provision of Sec. 60d Copyright Act if they cooperate with a private undertaking that has a decisive influence on the research organisation and preferential access to the results of scientific research. However, it is not sufficient to point out that one of the co-founders of the research organisation is employed by a commercial AI undertaking as ‘Head of Machine Learning Operations’, for example. The mere fact that members of the association work for the commercial AI undertaking does not prove that this undertaking has a decisive influence on the organisation’s research work. In addition, it would at least have to be claimed that the organisation granted the commercial undertaking preferential access to the results of its scientific research, namely to the dataset at issue. It is not sufficient for the commercial AI undertaking to have trained its service using the dataset at issue (para. 112).
Hamburg Regional Court (Landgericht Hamburg), judgment of 27 September 2024 – 310 O 227/23
Facts of the case:
[1] The defendant is an association that was founded with a constituent meeting on 7 July 2021 (minutes in Exhibit B7, statutes in Exhibit Bl). The specific purpose of the defendant’s activities is disputed between the parties.
[2] The defendant makes what is known as a dataset for image-text pairs publicly available free of charge under the name ‘L.’. This is a type of tabular document that contains hyperlinks to publicly accessible images or image files on the internet as well as further information about the corresponding images, including an image description (also known as alternative text) that provides information about the content of the image in text form. The dataset comprises 5.85 billion corresponding image-text pairs. The dataset can be used to train so-called generative artificial intelligence.
[3] The dataset was created after the defendant was founded in the second half of 2021. For this purpose, the defendant had used an existing dataset of C. C. F. from the USA (www. c..org), which contained the relevant URLs together with a textual description of the respective image content for a kind of random cross-section of the images that could be found on the internet. The defendant then extracted the URLs for the images from this dataset and downloaded the images from their respective storage locations. The defendant then used software to check the images to see whether the description of the image content already in the existing dataset actually matched the content to be seen in the image. Images where the text and image content did not match sufficiently were filtered out. For the remaining images, the metadata, in particular the URL of the image’s storage location and the image description, were extracted and included in a new dataset, the L. Whether the downloaded image files were subsequently deleted again is disputed between the parties – at least with regard to the photograph at issue.
[4] The image at issue was, as part of the aforementioned process, also captured, downloaded, analysed and included with its metadata in the L. dataset. Specifically, an image file posted on the website of the photo agency B. (https://www. b..com) with a watermark of the photo agency B. was downloaded.
[5] On the website of photo agency B., the following text was to be found on the subpage https://www. b..com/de/usage.html since at least 13 January 2021:
[6]‘RESTRICTIONS
[7]YOU MAY NOT:
[8](…)
[9]18. use automated programs, applets, bots or the like to access the B..com website or any content thereon for any purpose, including, by way of example only, downloading Content, indexing, scraping or caching any content on the website.’
[10] The plaintiff alleges an infringement of the copyright in the photograph at issue in the form of an unauthorised reproduction by the defendant as part of the analysis process.
[11] The plaintiff claims that he is the author of the photograph identified in the operative part of the judgment. The undertaking B. was entitled to offer for sale and to display the photo at issue on its website b..com and to offer licences for the photo; in this respect, B. was the holder of simple, sublicensable rights of use.
[12] The – undisputed – reproduction that took place as part of the analysis process infringed the plaintiff’s rights under Sec. 16 Copyright Act; in particular, it was not covered by the exception provisions of Secs. 44a, 44b and 60d of the Act:
[13] The plaintiff argues that the exception provision of Sec. 44a Copyright Act is not relevant, the independent downloading of a photograph in particular does not constitute a temporary act within the meaning of this provision.
[14] Nor, according to the plaintiff, is the reproduction covered by Sec. 44b Copyright Act. The aggregation of data for the purpose of AI training is not text or data mining within the meaning of Sec. 44b. Neither the European nor the German legislator had such a use ‘in mind’ when creating the exception provision of Art. 4 of the EU Directive on Copyright in the Digital Single Market (DSM Directive) and Sec. 44b Copyright Act respectively. In the case of text and data mining within the meaning of Sec. 44b Copyright Act, only ‘information hidden in the data should be made accessible’, ‘but the content of the intellectual creation should not be used’. However, the so-called ‘AI-web scraping’ at issue here is precisely about the intellectual content of the works used for training purposes ‘and ultimately about the creation of identical or similar competing products’. Moreover, according to the defendant’s own disclaimer (printed on Sheet 47), the dataset is ‘uncurated’. Finally, the collection and storage for the creation of parallel archives is excluded from the exception provision of Sec. 44b Copyright Act according to the legislator’s expressly declared intent.
[15] In addition, ‘the mass incorporation of copyrighted works for training purposes in the context of generative AI’ impairs the normal exploitation of copyrighted works, because it creates the conditions for replacing authors in many cases, or in any event, by offering a free competing product, makes it considerably more difficult to exploit the work. However, according to Art. 7(2) DSM Directive in conjunction with Art. 5(5) of the EC Directive on harmonisation of certain aspects of copyright and related rights in the information society (InfoSoc Directive), this precludes the application of the exception provision.
[16] In the plaintiff’s view, the reproduction is in any event inadmissible due to the reservation of use declared on the website www. b..com pursuant to Sec. 44b(3) Copyright Act. The picture agency’s declaration to this effect is attributable to the plaintiff, since it distributes the photo at issue for him. Contrary to the defendant’s view, the reservation is also machine-readable within the meaning of Sec. 44b(3) second sentence Copyright Act. The requirements in this regard are no higher than for machine readability by humans; however, the reservation is written in printed letters. Moreover, the text is also recognisable to a computer program as being a reservation. The ChatGPT service was able to recognise the corresponding reservation, and specific tools such as WebOpt-Out could recognise reservations such as those on b..com.
[17] Nor can the defendant invoke the exception provision of Sec. 60d Copyright Act. The plaintiff disputes that the defendant fulfils the requirements of Sec. 60d Copyright Act in terms of facts, namely:
[18] – that it was ‘registered in the register of associations’ at the time of the act of reproduction at issue here;
[19] – that document Bl submitted by the defendant represented the defendant’s valid statutes or had represented them at the time of the act of reproduction at issue;
[20] – that the members of the association and the board of directors are volunteers or were volunteers at the time of the act of reproduction at issue;
[21] – that the defendant is or was at the time of the act of reproduction at issue exclusively active in research or that the defendant conducts scientific research, pursues non-commercial purposes and reinvests all profits in scientific research or is active in the public interest within the framework of a state-recognised mandate; furthermore, according to the statutes submitted by the defendant, the purpose of the defendant is only the ‘promotion of research’ and not ‘research’ as such, and it is also unclear what research is supposed to be in the collection created by the defendant, which is (undisputedly) made available to other undertakings;
[22] – that the defendant creates and tests its own client models on the basis of the training data in order to further research the possibilities of AI technology;
[23] – that by making the training dataset publicly available, other researchers and interested parties should be offered the opportunity to train their own AI models; according to the defendant’s own statements, the dataset at issue was also used to train the services ‘DALL-E 2’, ‘Midjourney’ and ‘Stable Diffusion’ of the provider S. AI; however, these were operated by (purely) commercial undertakings; insofar as the defendant denies training the first two services, it would in any case have been possible for them to use the dataset.
[24] Moreover, the defendant cannot pursuant to Sec. 60d(2) No. 3 Copyright Act invoke the privilege of Sec. 60d of the Act. The defendant obviously cooperates intensively with commercial providers:
[25] – Thus, there is apparently a cooperation with the private undertaking S. AI, which has a direct influence on the defendant through the financing of the dataset in question and the filling of relevant positions at the defendant by its own employees. According to an interview with its founder and managing director, S. AI financed the L. dataset.
[26] – Members of the defendant’s ‘team’ are also commercially active ‘in many places’ in the same field for large tech undertakings, including as employees of S. AI.
[27] – Furthermore, in a chat on the platform ‘D.’ on September 28, 2021, the defendant’s co-founder, R. V. pressed for a rapid completion of the ‘1B dataset’, as they had received funding of USD 5,000 from a certain ‘J.’ or his undertaking; the data should (already) be made available to the latter even if the dataset could not yet be made available to the public. The ‘J.’ in question is an employee of the commercial AI provider M.
[28] Finally, the defendant cannot rely on so-called ‘implied consent’. He – the plaintiff – did not make the photograph at issue freely accessible, but arranged for it to be offered by the B. agency for the granting of paid licences.
[29] Alongside the petition for a cease-and-desist order with respect to the reproduction, the plaintiff had initially requested information about the extent to which the photograph had been used. At the hearing on 11 July 2024, the parties agreed that this request for information was settled.
[30] The plaintiff now still requests
[31] that the defendant be ordered to refrain from reproducing and/or permitting the reproduction of the following photograph on pain of a fine of up to EUR 250,000, or alternatively imprisonment for up to six months for each individual case of non-compliance,
[32] Image removed
[33] for the purpose of creating AI training datasets, as was done in the context of the creation of the L. dataset (…).
[34] The defendant requests
[35] that the action be dismissed.
[36] It disputes that the plaintiff created the image at issue himself or that he is otherwise entitled to assert infringements of rights in his own name with regard to the image, and that at the time the image was captured by the defendant, the plaintiff was not entitled to assert infringements of rights in his own name with regard to the image at issue.
[37] Above all, however, the (one-time) download of the image at issue in the context of the creation of the L. dataset admittedly constitutes a copyright-relevant reproduction, but this is covered by the exception provisions of Secs. 44a, 44b and 60d Copyright Act as well as the plaintiff’s implied consent:
[38] Firstly, the reproduction that took place was covered by the exception provision of Sec. 44a Copyright Act. The defendant does not permanently store the images; rather, the images are only used for analysis for a short time and then immediately automatically and irrevocably deleted. The reproduction does not have any independent economic significance.
[39] The exception provision of Sec. 44b Copyright Act also applies. The analysis of image data and the extraction of metadata for the training of artificial intelligence is a main application of text and data mining according to the intention of the legislator. Nor are digital parallel archives created, as the downloaded images are not permanently stored, and all that is recorded is hyperlinks. The exception under Sec. 44b(3) Copyright Act does not apply:
[40] – According to the plaintiff’s submission, it was not he himself as the rightholder, but the operator of the website www. b..com as a third party who declared this reservation of use; the plaintiff himself even expressly stated in his e-mail of 13 February 2023 (Exhibit B5) that he had neither the qualifications nor the economic means to declare a reservation of use.
[41] – Moreover, the reservation is not explicit, as the passage on the website www. b..com is general and lists various unauthorised actions. There is no explicit mention of text and data mining or reproductions.
[42] – Furthermore, the characteristic of machine readability is missing. A clause written in natural language is generally not machine-readable within the meaning of Sec. 44b(3) Copyright Act. A prerequisite for machine readability in this sense is that the clause can be automatically processed by software. This requires the corresponding information to be encoded. In any case, it is necessary that the text at least contains specific keywords such as ‘data mining’.
[43] – The reservation was also clearly not intended as one in accordance with Sec. 44b(3) Copyright Act. The fact that, according to the plaintiff’s submission, the clause was already present on the website on 13 January 2021 makes it clear that the clause could not have been drawn up ‘with a view to the provision in Sec. 44b(3) Copyright Act’, since the statutory provision had not yet come into force at that time. Moreover, ‘nor was it credible’ that a US provider would invoke a reservation of use based on German law.
[44] In any case, it – the defendant – can invoke the exception provision of Sec. 60d Copyright Act.
[45] – It – the defendant – is a non-profit association consisting of researchers and, according to the association’s statutes (Exhibit Bl), is dedicated to research; in particular, it has set itself the task of further developing self-learning algorithms in the sense of artificial intelligence and making them available to the general public. To this end, it provides datasets and models free of charge, as well as creating and testing its own AI models based on the training data.
[46] – Its activities also constitute ‘research’. Simply by making transparent on the internet how the training datasets are created, it contributes to the acquisition of knowledge about the training of artificial intelligence. In this way, other researchers could follow the steps for creating the datasets and build on them. In addition, it published a scientific paper entitled ‘L.: An open large-scale dataset for training next generation image-text models’ on the L. dataset at issue for the first time on 17 September 2022 (Exhibit B8). By 5 April 2024, this paper had been cited a total of 1,403 times in other scientific works and had also received further awards.
[47] – In addition, the defendant also trains its own AI models on the basis of the datasets it has created in order to gain insights into how AI can be improved through appropriate training.
[48] – The natural persons that comprise the defendant, i.e., the board members and other members of the association, are also ‘researchers’. The plaintiff’s reference to the fact that individual team members work in the same field for large tech undertakings has no relevance to the question of whether the defendant itself is commercially active. Moreover, the persons concerned work for the defendant on a voluntary basis and therefore are obliged to earn their living elsewhere.
[49] – The fact that the datasets made publicly available by the defendant are also used by commercial providers is irrelevant for the application of the exception provision of Sec. 60b Copyright Act. Moreover, the DALL-E 2 and Midjourney services were not actually trained with the defendant’s dataset.
[50] – Nor does the reverse exception provided for in Sec. 60d(2) No. 3 Copyright Act apply in the present case. The undertaking S. AI had indeed made computing resources available to the defendant during the start-up phase. However, the same had also been provided by J. S. C. (JSC). The defendant did not receive any financial support in the form of money from S. AI. Nor had there been any further cooperation with this undertaking. In any case, S. AI did not receive preferential access to the research results. Nor did S. AI have any decisive influence. Neither S. AI itself nor any of its legal representatives were members of the defendant.
[51] Finally, the plaintiff has given so-called implied consent in favour of the uses by the defendant. The reproduction at issue is to be classified as a customary act of use.
[52] It should also be noted that the plaintiff himself earns money with AI-generated images. This calls into question his ‘need for legal protection’.
[53] With regard to the further details of the parties’ submissions, reference is made to the written submissions and exhibits on the file, insofar as they were made the subject of the in-person hearing, as well as to the minutes of the hearing of 11 July 2024 (Sheets 120-123).
[54] The parties submitted disallowed written submissions dated 29 August and 11 September 2024 (plaintiff) and 20 September 2024 (defendant) to the file.
Grounds:
I.
[55] The action is admissible but is unsuccessful on the merits. It is true that the defendant has encroached on the plaintiff’s exploitation rights by reproducing the photograph at issue. However, this encroachment is covered by the exception provision of Sec. 60d Copyright Act. Against this background, there is no need for a conclusive finding as to whether the defendant can additionally invoke the exception provision of Sec. 44b Copyright Act.
[56] In any event, the photograph at issue is protected as a photographic work pursuant to Sec. 72(1) Copyright Act. After inspecting the raw data found on the plaintiff’s laptop, this court also has no doubt as to the plaintiff’s status as photographer, Sec. 72(2) Copyright Act. The plaintiff is also entitled to assert claims for infringement pursuant to Sec. 97 Copyright Act, including the right to require cessation of the infringement pursuant to para. 1 of the provision; the defendant has not demonstrated that the plaintiff granted the picture agency B. more extensive rights than (sublicensable) simple rights of use. The picture agency B. provided the photo with a watermark; this was a transformation requiring consent within the meaning of Sec. 23(1) first sentence Copyright Act, so that the consent of the plaintiff as the author was also required in principle for its exploitation. As part of the downloading process, the defendant reproduced this version within the meaning of Sec. 16(1) Copyright Act without having obtained the plaintiff’s consent.
[57] However, the defendant was entitled to do so on the basis of a statutory authorisation. It is true that the reproduction was not covered by the exception provision of Sec. 44a Copyright Act (below, 1.), and whether the defendant can rely on the exception provision of Sec. 44b Copyright Act appears doubtful (below, 2.). However, there is no need for a final decision on the latter in the present case, as the act of reproduction was in any case covered by the exception provision of Sec. 60d Copyright Act (below, 3.).
1.
[58] The act of reproduction is not covered by the exception provision of Sec. 44a Copyright Act.
[59] According to this provision, temporary acts of reproduction are permitted which are transient or incidental and constitute an integral and essential part of a technological process and whose sole purpose is to enable transmission in a network between third parties by an intermediary or a lawful use of a work or other protected subject matter and which have no independent economic significance.
[60] The reproduction in the present case was neither transient nor incidental.
a)
[61] A reproduction is transient within the meaning of Sec. 44a Copyright Act if its duration is limited to what is necessary for the proper completion of the technological process in question, it being understood that that process must be automated so that it deletes that act automatically, without human intervention, once its function of enabling the completion of such a process has come to an end (judgment of the CJEU, 16 July 2009, C-5/08 – Infopaq v. Danske Dagblades Forening, para. 64 (juris) on Art. 5(1) DSM Directive).
[62] The defendant’s reliance on the fact that the files were ‘automatically’ deleted as part of the analysis procedure it carried out cannot establish the transience of the copying in the aforementioned sense. Apart from the fact that the defendant has not stated anything about the concrete duration of the storage, the deletion was not ‘user-independent’, but rather due to a corresponding deliberate programming of the analysis process by the defendant.
b)
[63] A reproduction is incidental within the meaning of Sec. 44a Copyright Act if it neither exists independently of, nor has a purpose independent of, the technological process of which it forms part (judgment of the CJEU, 5 June 2014, C-360/13, para. 43 (juris)).
[64] In the present case, the image files were downloaded specifically in order to analyse them using specific software. Thus, the downloading is not merely a process incidental to the analysis carried out, but a conscious and actively controlled procurement process preceding the analysis.
2.
[65] Whether the defendant can invoke the exception provision of Sec. 44b Copyright Act appears doubtful in the present case. It is true that the download carried out by the defendant is in principle subject to the exception provision of Sec. 44b(2) Copyright Act, in particular it was carried out for the purpose of text and data mining within the meaning of Sec. 44b(1) Copyright Act (below, a). However, without this requiring a final decision in the present case, there are some indications that the act of reproduction was not permissible under Sec. 44b(2) Copyright Act due to an effectively declared reservation of use within the meaning of Sec. 44b(3) Copyright Act (below, b).
a)
[66] The act of reproduction at issue is as a matter of principle subject to the exception provision of Sec. 44b(2) Copyright Act.
[67] (1) The download at issue was made for the purpose of text and data mining within the meaning of Sec. 44b(1) Copyright Act. According to this provision, text and data mining is the automated analysis of individual or several digital or digitised works for the purpose of gathering information, in particular regarding patterns, trends and correlations. In any event, this is to be affirmed for the act of reproduction at issue here (below, (a)); a teleological reduction of the exception cannot be considered in this respect (below, (b)).
[68] In the present case, there is therefore no need to decide the further question, which has been discussed in detail in the literature, as to whether or not the training of artificial intelligence in its entirety is subject to the exception provision of Sec. 44b Copyright Act (for a detailed opinion, see BeckOK UrhR/Bomhard, 42nd edn. 15 February 2024, Copyright Act Sec. 44b paras. 1la-1lb with further references; see also in detail the study ‘Urheberrecht & Training generativer KI – technologische und rechtliche Grundlagen’, commissioned by the Initiative Urheberrecht and submitted as Exhibit K11).
[69] (a) The defendant carried out the act of reproduction for the purpose of obtaining information on ‘correlations’ in the literal sense of Sec. 44b(1) Copyright Act. The defendant downloaded the photograph at issue from its original storage location in order to compare the image content with the image description already stored for the text using software that was already available – apparently the OpenAI CLIP application. This analysis of the image file to compare it with a pre-existing image description constitutes an analysis for the purpose of obtaining information about ‘correlations’ (namely the question of non-matching/matching images and image descriptions). The fact that the defendant analysed the images included in the L. dataset in this way was not disputed as such by the plaintiff.
[70] Nor is the applicability of Sec. 44b(1) Copyright Act – contrary to the opinion of the plaintiff (reply pp. 13 et seq., Sheet 47 et seq. of the files) – excluded on the grounds that the defendant did not ‘curate’ the L. B5 dataset created by him according to a ‘disclaimer’ issued for this dataset. The disclaimer reproduced by the plaintiff refers solely to a warning that the dataset had not been searched for ‘disturbing content’ or the like. However, such an – additional – filtering of the dataset to be created is not a prerequisite for the application of Sec. 44b(1) Copyright Act and does not contradict the assumption that the downloaded images – as stated – were analysed for the correlation between image content and image description.
[71] (b) Nor is the act of reproduction at issue to be excluded from the exception provision of Sec. 44b Copyright Act by way of a teleological reduction.
[72] Although the exclusion of the reproduction of data for the purpose of AI training by way of teleological reduction is occasionally advocated in the literature on the grounds that Sec. 44b Copyright Act only covers the development of ‘information hidden in the data’, but not the use of ‘the content of the intellectual creation’ (Schack, NJW 2024, 113; in this direction also Dormis and Stober, Urheberrecht und Training generativer KI-Modelle, Annex K1l, pp. 67 et seq., differentiating between semantics and syntax), there are doubts as to whether this is convincing, as it is not sufficiently clear what the difference is between ‘information hidden in the data’ and ‘the content of the intellectual creation’ in the case of digitised works.
[73] As to the further argument that ‘AI web scraping’ is about the intellectual content of the works used for training purposes and ‘ultimately’ about the creation of content of identical or similar competing products (Schack, loc. cit.), this court is of the opinion that this argumentation does not distinguish strictly enough between:
[74] – firstly, the creation of a dataset (which is the sole subject of the dispute here) that can – also – be used for AI training;
[75] – secondly, the subsequent training of the artificial neural network with this dataset; and
[76] – thirdly, the subsequent use of the trained AI for the purpose of creating new image content.
[77] Admittedly, this latter functionality may already be the aim when the training dataset is created. However, at the time of compiling the training dataset, it is neither foreseeable in what way the second step (training) will be successful, nor what specific content can be generated by the trained AI in the third step (in the application of the AI). The concrete application possibilities for a rapidly developing technology such as AI are therefore not conclusively foreseeable at the time of the creation of the training dataset and therefore cannot be determined with legal certainty. Due to this legal uncertainty, the mere general – and only – intention at the time of the creation of the training dataset to obtain future AI-generated content is not a suitable criterion for assessing the legal admissibility of the creation of the training dataset as such.
[78] Finally, to argue, in support of a teleological reduction of the exception provision of Sec. 44b Copyright Act, that the European legislator ‘simply did not yet have the AI problem’ ‘on its radar’ when creating the underlying provision of the Directive (Art. 4 DSM Directive) in 2019 (Schack, loc. cit.; likewise for the training of AI models Dormis and Stober, loc. cit., pp. 71 et seq., 87 et seq.), is of itself clearly not sufficient for a teleological reduction. In particular, account must be taken of the fact that technical developments in the field of artificial intelligence since 2019 concern less the type and scope of data mining for the procurement of training data (at issue here), but rather the performance of the artificial neural networks trained with the data (accordingly, Dormis and Stober, loc. cit., p. 95, also assume that the mere creation of training datasets ‘in advance of the actual training’ may well fall below the TDM threshold). It should also be noted that the C. C. F. database accessed by the defendant has been in creation since 2008 (!), cf. https:// c.. o./ o.
[79] Apart from this, the current European legislator of the AI Act (Regulation (EU) 2024/1689 of 13 June 2024, OJ L 12 July 2024 p. 1) has unequivocally expressed that the creation of datasets intended for the training of artificial neural networks is also subject to the restriction of Art. 4 of the DSM Directive. According to Art. 53(1)(c) of the AI Act, providers of AI models with a general purpose are obliged to put in place a policy in particular to identify and comply with a reservation of rights invoked pursuant to Art. 4(3) of the DSM Directive.
[80] The fact that the creation of datasets intended for the training of artificial neural networks is also subject to the exception provision of Art. 4 of the DSM Directive is, moreover, generally in accordance with the assessment of the German legislator in the context of its implementation of the aforementioned exception provision in 2021 (Explanatory Memorandum to the draft bill BT-Drs. 19/27426, p. 60).
[81] (c) Nor, ultimately, does the so-called three-step test enshrined in Art. 5(5) InfoSoc Directive (in conjunction with Art. 7(2)(1) DSM Directive) justify a different assessment. According to this test, the exceptions provided for may only be applied in certain special cases in which the normal exploitation of the work or other protected subject matter is not impaired and the legitimate interests of the rightholder are not unreasonably prejudiced. These requirements are met in the present case.
[82] The reproduction relevant to copyright law in the present case is limited to the purpose of analysing the image files for their conformity with a pre-existing image description and subsequent entry into a dataset. It is not apparent, and is not claimed by the plaintiff, that this use would impair the exploitation possibilities of the works concerned.
[83] It is true that the dataset created in this way may subsequently be used to train artificial neural networks and the resulting AI-generated content may compete with the works of (human) authors. However, this alone does not justify considering the creation of the training datasets as prejudicing the exploitation rights to works within the meaning of Art. 5(5) InfoSoc Directive. This must be the case simply because the consideration of merely future technical developments, which are not yet foreseeable in detail, does not allow for a legally certain distinction between permissible and impermissible uses (see similarly (b) above).
[84] Since, in the event of doubt, it can never be ruled out on the basis of current technological developments that the knowledge gained by means of text and data mining will be used to train artificial neural networks which can then compete with authors, the contrary position would ultimately even require text and data mining within the meaning of Sec. 44b Copyright Act to be prohibited in its entirety; however, such a complete invalidation of the exception provision would obviously run counter to the legislative intention and therefore cannot represent a tenable interpretation.
[85] (2) The image file downloaded by the defendant was also, as undisputed by the plaintiff, lawfully accessible within the meaning of Sec. 44b(2) first sentence Copyright Act.
[86] A work is ‘lawfully accessible’ in this sense in particular if it is freely accessible on the internet (Explanatory Memorandum to the draft bill BT-Drucks. 19/27426, p. 88).
[87] This can be assumed for the image downloaded by the defendant. Contrary to the plaintiff’s initial submission, the defendant did not download the ‘original image’ reproduced in the application for cessation of the infringement initially formulated in the statement of claim – which would only have been made available by the picture agency B. if a license had been purchased – but rather a version of the image bearing a watermark of the picture agency. This was obviously the preview image posted on the agency’s website for advertising purposes. However, precisely this watermarked preview image had been made freely accessible on the internet by the agency.
b)
[88] However, there are some indications that the exception provision of Sec. 44b(2) Copyright Act does not apply in the present case – without this requiring a final decision – since there was an effectively declared reservation of use within the meaning of para. 3 of the provision; in particular, the reservation of use indisputably declared on the B..com website is likely to meet the requirements for machine readability within the meaning of Sec. 44b(3) second sentence Copyright Act.
[89] (1) There is much to suggest that the reservation of use declared on the agency’s website was declared by a person authorised to do so and that the plaintiff can also rely on this to protect his own rights.
[90] According to the wording of Sec. 44b(3) Copyright Act, ‘the rightholder’ can declare the reservation of use. This means that account must be taken not only of declarations of reservation by the author himself, but also by subsequent rightholders, whether they are successors in title or holders of rights derived from the author. According to the plaintiff’s conclusive submission (minutes of 11 July 2024 p. 3, Sheet 122 of the files), he had granted the B. picture agency simple rights of use to the original picture that could be sublicensed. The picture agency was then itself the rightholder of the pictures posted on its website and was therefore able to declare a reservation of use in accordance with Sec. 44b(3) Copyright Act without further ado; it is neither apparent nor has it been asserted that this was prevented by agreements with an in rem effect in the contractual relationship between the plaintiff and the picture agency.
[91] The plaintiff is probably also entitled to invoke this declaration of reservation by his licensee. From a commercial point of view, the original photo at issue was exploited via the agency. Thus, in practice, the specific decision as to which third party was to receive the authorisation for which use lay with the agency; the agency was not under any obligation to conclude a contract. In such a situation, this court believes that there is much to suggest that the author may invoke a reservation declared by his licensee pursuant to Sec. 44b(3) Copyright Act when asserting the prohibition rights he retained.
[92] (2) The defendant’s objection that the prohibition of use for web crawlers declared in the agency’s general terms and conditions vis-à-vis its customers cannot be formulated ‘in relation to Sec. 44b(3) Copyright Act’ if only because of the time situation, is also irrelevant. It is not a prerequisite for the legal effects of the declaration that it is consciously declared with regard to a specific version of the law.
[93] (3) The reservation is also sufficiently clearly formulated. Article 4(3) DSM Directive requires an explicit declaration of the reservation of use. This explicitness requirement must therefore be taken into account when interpreting Sec. 44b(3) Copyright Act in conformity with the Directive (see also the Explanatory Memorandum to the draft bill BT-Drs. 19/27426, 89). The declared reservation must therefore be both expressis verbis (not implied) and so specific (concrete and individual) that it unequivocally covers a specific content and a specific use (Hamann, ZGE 16 (2024), p. 134). The reservation of use formulated on the picture agency B.’s website easily meets these requirements.
[94] Furthermore, the argument that a reservation of use declared for all works posted on a website contradicts the explicitness requirement of Sec. 44b(3) Copyright Act (thus in extension of his own abstract conclusion Hamann, loc. cit., p. 148) is not convincing. This is because the scope and content of the reservation explicitly declared for all works posted on a website can be determined beyond doubt and is therefore explicitly declared.
[95] (4) Finally, there are also considerable aspects in support of the argument that the reservation of use satisfies the requirements of machine readability within the meaning of Sec. 4[4]b(3)1 second sentence Copyright Act.
[96] In view of the underlying legislative intention to enable automated searches by web crawlers (cf. Explanatory Memorandum to the draft bill BT-Drucks. 19/27426, p. 89), the term ‘machine readable’ should certainly be interpreted in the sense of ‘machine understandable’ (see Hamann, loc. cit., pp. 113, 128 et seq.).
[97] However, this court is inclined to regard a reservation of use written solely in ‘natural language’ as ‘machine understandable’ (unlike the probably predominant view in the literature, see Hamann, loc. cit. pp. 131 et seq., 146 et seq. with further references to the current state of opinion, including a reference to a contribution by the defendant’s counsel here, namely Akinci and Heidrich, IPRB 2023, 270, 272, who apparently also take the Court’s view; however, the paper was not directly accessible to the Court by the time the judgment was handed down). However, the question of whether and under what specific conditions a reservation declared in ‘natural language’ can also be regarded as ‘machine understandable’ will always have to be answered in the light of the technical development existing at the relevant time of the use of the work.
[98] Accordingly, the European legislator also laid down in the AI Act that providers of AI models must put in place a policy in particular to identify and comply with a reservation of rights expressed pursuant to Art. 4(3) of the DSM Directive ‘including through state-of-the-art technologies’ (Art. 53(1)(c) AI Regulation). However, these ‘state-of-the-art technologies’ beyond doubt include AI applications that are capable of capturing the content of text written in natural language (according in particular to the defendant’s counsel Akinci and Heidrich in the article IPRB 2023, 270, 272, not directly accessible to the Chamber, cited here according to Hamann, loc. cit. p. 148, who incidentally affirms this possibility in technical terms, loc. cit.). In this respect, there is every indication that the legislator of the AI Act had precisely such AI applications in mind with its reference to ‘state-of-the-art technologies’.
[99] Against such a view, some argue that it leads to a circular conclusion: If it is required that the operator of the text and data mining must use AI applications to check whether a reservation of use has been declared, then this AI-supported search in turn requires a pattern analysis, which already satisfies the actus reus of text and data mining within the meaning of Sec. 44b(1) Copyright Act; in other words, only the application of the exception decides on the permissibility of its application (according to Hamann, loc. cit., p. 148). This court does not share this assessment: Contrary to this view, the copyright-relevant act of use requiring justification is not the performance of a ‘pattern analysis’ as such, but the reproduction of the copyright work within the meaning of Sec. 16 Copyright Act. It does not appear to be mandatory that the prior finding of such works on the internet and their verification as to whether reservations within the meaning of Sec. 44b(3) second sentence Copyright Act have been declared requires quasi upstream further text and data mining within the meaning of Sec. 44b(1) Copyright Act, since it is in particular conceivable that the website content is processed through the use of web crawlers, in which only fleeting and incidental reproductions are made, which in turn are already justified under Sec. 44a Copyright Act.
[100] Furthermore, the objection is raised to this court’s broader understanding of the term ‘machine readability’ that this term is understood more narrowly by the European legislator in a different context. In this context, reference is made to Recital 35 of the PSI Directive (Directive (EU) 2019/1024), which requires, among other things, ‘simple’ recognisability for ‘machine readability’ within the meaning of that Directive (according to BeckOK UrhR/Bomhard, 42nd edn. 15 February 2024, Copyright Act Sec. 44b para. 31 with further references); it is argued that this cannot be assumed for a reservation formulated only in natural language. However, such an argument presupposes that the terms of both directives must be understood in the same way. This court has doubts as to whether such an equation of the terms is convincing, as the directives have different objectives: While the PSI Directive deals with the purely unilateral access of the public to information or the purely unilateral obligation of public authorities to publish certain information, Art. 4(3) of the DSM Directive deals with a balance between the text and data mining users’ interests (in being able to do this as easily and as legally securely as possible) and the rightholders’ interests (in securing their rights as easily and as effectively as possible). In the opinion of this court, this balance of interests cannot be resolved one-sidedly in favour of the users of text and data mining by considering only the simplest conceivable technical solution for them as being sufficient for the effectiveness of a declared reservation of use. Such an understanding would also be contradicted by the values expressed by the legislator of the DSM Directive, Recital 18 of which does not require the declaration of a reservation ‘in the simplest possible manner’, but only ‘in an appropriate manner’. The German transposing legislator likewise only requires a declaration in a way that is ‘appropriate to the automated processes of text and data mining’ (Explanatory Memorandum to the draft bill BT-Drucks. 19/27426, p. 89).
[101] In this court’s view, there would furthermore be a certain inconsistency if providers of AI models were allowed to develop increasingly powerful text-understanding and text-creating AI models via the exception in Sec. 44b(2) Copyright Act on the one hand, but not required to use existing AI models within the framework of the exception in Sec. 44b(3) second sentence Copyright Act on the other.
[102] The plaintiff has hitherto not demonstrated whether and to what extent sufficient technology for the automated understanding of the content of the disputed reservation of use was already available at the time of the act of reproduction in 2021; in this respect, the plaintiff has only referred to services available in 2023 (response, pp. 14 et seq., Sheets 48 et seq. of the files). However, there are indications that the defendant already had suitable technology at its disposal. According to the defendant’s own submission, the analysis carried out as part of the creation of the L. dataset in the form of a comparison of image content with pre-existing image descriptions clearly also and specifically required the content of these image descriptions to be understood by the software used. Against this background, there is some evidence that systems were already available in 2021 – especially to the defendant – that were capable of automatically understanding a reservation of use formulated in natural language.
3.
[103] However, the defendant can invoke the exception provision of Sec. 60d Copyright Act with regard to the reproduction at issue.
[104] According to this provision, reproductions by research organisations for text and data mining for scientific research purposes are permissible.
a)
[105] As explained above, the reproduction was made for the purpose of text and data mining within the meaning of Sec. 44b(1) Copyright Act. It was also made for the purposes of scientific research within the meaning of Sec. 60d(1) Copyright Act.
[106] Scientific research generally refers to the methodical and systematic pursuit of new knowledge (Spindler, Schuster and Anton, 4th edn. 2019, Copyright Act Sec. 60c para. 3; BeckOK UrhR/Grübler, 42nd edn. 1 May 2024, Copyright Act Sec. 60c para. 5; Dreier, Schulze and Dreier, 7th edn. 2022, Copyright Act Sec. 60c para. 1). The concept of scientific research, which is already satisfied by the methodical-systematic ‘striving’ for new knowledge, is not to be understood so narrowly that it would only cover the work stages directly associated with the acquisition of knowledge; rather, it is sufficient that the work stage in question is aimed at a (later) acquisition of knowledge, as is the case, for example, with numerous data collections that must first be made in order to subsequently draw empirical conclusions. In particular, the concept of scientific research does not presuppose any subsequent research success.
[107] Accordingly, contrary to the plaintiff’s opinion, the creation of a dataset of the type at issue, which can form the basis for training AI systems, can certainly be regarded as scientific research in the aforementioned sense. Although the creation of the dataset as such may not yet be associated with an acquisition of knowledge, it is a fundamental work stage with the objective of using the dataset for the purpose of gaining knowledge at a later date. It can be affirmed that such an objective also existed in the present case. It is sufficient that the dataset was – indisputably – published free of charge and thus made available to researchers (also) in the field of artificial neural networks. Whether the dataset – as the plaintiff claims with regard to the DALL-E 2 and Midjourney services – is also used by commercial undertakings for training or further development of their AI systems is irrelevant because research by commercial undertakings is still also research – even if not privileged as such under Secs. 60c et seq. Copyright Act.
[108] Against this background, the question at issue between the parties as to whether the defendant also pursues scientific research in the form of the development of its own AI models in addition to the creation of corresponding datasets is irrelevant.
b)
[109] The defendant pursues non-commercial purposes within the meaning of Sec. 60d(2) No. 1 Copyright Act.
[110] The question of whether research is non-commercial depends solely on the specific nature of the scientific activity, while the organisation and financing of the institution in which the research is carried out are irrelevant (Recital 42 InfoSoc Directive).
[111] The non-commercial purpose pursued by the defendant in relation to the disputed creation of the L. dataset is already evident from the fact that the defendant indisputably makes it publicly available free of charge. The fact that the development of the dataset at issue would at least also serve the development of the defendant’s own commercial range of products and services (cf. on this criterion BeckOK IT-Recht/Paul, 14th edn. 1 April 2024, Copyright Act Sec. 60d para. 10) has neither been argued by the plaintiff nor is it otherwise apparent. The fact that the dataset at issue may also be used by commercially active undertakings for the training or further development of their AI systems is irrelevant for the qualification of the defendant’s activity. The mere fact that individual members of the defendant also carry out paid activities for such undertakings alongside their work for the association is not sufficient to attribute the activities of these undertakings to the defendant as its own.
c)
[112] Nor is the defendant barred from invoking the exception provision of Sec. 60d Copyright Act by para. 2 No. 3 of the provision.
[113] According to this provision, the exception provision of Sec. 60d Copyright Act cannot be invoked by research organisations that cooperate with a private undertaking which exerts a certain degree of influence on the research organisation and has preferential access to the findings of its scientific research. According to the wording of the provision, the burden of presentation and proof for the actual requirements of this exclusion from the exclusion pursuant to Sec. 60d(2) No. 3 Copyright Act lies with the plaintiff.
[114] (1) The plaintiff’s initial reference in its reply to the fact that the S. AI undertaking had a direct influence on the defendant via the financing of the dataset in question and the filling of ‘relevant positions’ at the defendant by its own employees (Reply p. 18, Sheet 52 of the files) lacks substantiation.
[115] In this respect, the plaintiff merely refers to the fact that one of the defendant’s co-founders, Mr. R. V., is employed by S. AI as ‘Head of Machine Learning Operations’, and that one of the defendant’s members, Mr. R. R., is also employed there as a ‘Research Scientist’ (Reply pp. 4 et seq., Sheets 38 et seq. of the files). However, this activity for the undertaking S. AI of two members of the association alone does not prove any decisive influence of this undertaking on the defendant’s research work.
[116] Leaving this aside, the plaintiff has not even claimed that the defendant granted the undertaking S. AI preferential access to the results of its scientific research, namely the dataset at issue. Rather, he only submits that S. AI had trained its S. D. service with the help of the dataset at issue (Reply pp. 8 et seq., Sheets 42 et seq. of the files).
[117] (2) The plaintiff’s submission of 3 July 2024, referring to a chat on the D. platform that took place in 2021, according to which the defendant’s co-founder, Mr. R. V., is said to have agreed to provide the M. undertaking, in return for a financial contribution of USD 5,000.00, with preferential access to the (then smaller) dataset likewise does not satisfy the exception in Sec. 60d(2) third sentence Copyright Act.
[118] There is no need to determine whether this chat – which is not disputed as such by the defendant (cf. submission of 9 July 2024 p. 3, Sheet 112 of the files) – supports the interpretation drawn by the plaintiff at all. Nor is there any need to determine whether the declaration of such a willingness to grant early access – the plaintiff has not argued whether this was actually granted – can amount to having preferential access to the research results within the meaning of Sec. 60d(2) No. 2 Copyright Act.
[119] In any event, the plaintiff has neither shown nor is it otherwise apparent that the M. undertaking has a decisive influence on the defendant. To the extent that personal ties between the defendant and AI industry undertakings have been shown to exist at all, these are the undertakings S. AI and Google (Reply pp. 4 et seq., Sheets 38 et seq. of the files).
Il.
[120] The decision on costs is based on Sec. 91(1), Sec. 9la(1) first sentence Code of Civil Procedure.
[121] The decision on provisional enforceability is based on Sec. 708 No. 11, Secs. 711 and 709 Code of Civil Procedure.
Translated from the German by David Wright-Policepayeh,
Kilb, Austria.
Footnotes
1
Translator’s note: There appears to be an error in the Court’s original manuscript. Instead of s 40b(3), it should be s 44b(3). We have amended the text accordingly.
© The Authors, 2025. Published by Oxford University Press on behalf of GRUR e.V. All rights reserved. For permissions, please email: journals.permissions@oup.com
This article is published and distributed under the terms of the Oxford University Press, Standard Journals Publication Model (https://academic.oup.com/pages/standard-publication-reuse-rights)