A recent investigation has exposed a significant point of contention within the intersection of artificial intelligence (AI) and intellectual property.
Works by Australian authors, including Peter Carey, Helen Garner, and Tim Winton, have been identified as part of a pirated dataset, known as Books3, which is being used for training generative AI models
This has catalysed an immediate debate, raising critical questions regarding the ethical and legal implications of using copyrighted materials for technological advancements.
After months of concern and speculation, it was horrifying to have confirmed that many Australian authors’ books have been used to train AI without permission. It confirms what we have suspected: that books from pirate sites have been used to train AI.https://t.co/ZKsapzxPpA
— Australian Society of Authors (@asauthors) September 27, 2023
Lack of transparency
Olivia Lanchester, chief executive officer of the Australian Society of Authors (ASA), articulates the dilemma stating: "The lack of transparency over what has been used to train generative AI means that authors haven't known whether their works have been used or not."
This lack of disclosure has agitated authors, many of whom were oblivious to the use of their intellectual property in such a manner.
Reactions from the authorial community span a wide emotional spectrum from mere disappointment to outright indignation.
What is Books3?
Books3, the dataset in question, was formulated by independent developer Shaun Presser.
Its objective was to furnish a resource that would enable smaller developers to compete on equal footing with technology conglomerates such as OpenAI.
Being labelled as "pirated", the dataset has elicited concerns over the unauthorised incorporation of copyrighted literary works.
Notably, this dataset has been employed in the training of several prominent AI models, including Meta's LLaMA, Bloomberg's BloombergGPT, and EleutherAI's GPT-J.
Advocates stricter regulations
Lanchester insists that this issue could have been completely averted had a different approach been adopted.
"There is a plethora of works in the public domain.
"AI developers could have used these or could have approached copyright owners to obtain a licence. It's that simple," she remarked.
The ASA is taking proactive steps to address the issue, advocating for stricter regulations on the use of copyrighted material in AI training.
Legal proceedings in the US
Moreover, the Australian authors' community is closely following legal proceedings in the United States where tech giants such as Meta and OpenAI are entangled in lawsuits for alleged unauthorised usage of copyrighted content.
Such cases could set legal precedents, influencing Australian laws governing intellectual property.
This unfolding situation brings to the fore the need for urgent, balanced dialogue among authors, technology developers, and policy-makers.
There is a pressing need to establish guidelines that safeguard the intellectual property of authors while facilitating technological advancements in AI.
Given the rapid development in AI technologies, time is of the essence in resolving these ethical and legal ambiguities.