Newsletter Subscribe
Enter your email address below and subscribe to our newsletter
[forminator_form id="25163"]

news.mitnews.mit+1news.mit+1A new study from MIT's Computer Science and Artificial Intelligence Laboratory has identified a phenomenon researchers call "attribution decay," finding that as AI image generators are trained on larger datasets, the influence of any individual training image becomes increasingly difficult — or impossible — to detect. The paper, published Tuesday in Nature Communications, carries implications for the ongoing wave of copyright litigation against AI companies.news.mit
Researchers Zheng Dai and David Gifford developed an architecture called a "diffusion ensemble" to answer a deceptively simple question: what would a model produce if it had never seen a particular training image? Rather than retraining models from scratch millions of times, their system trains multiple components on overlapping slices of data, allowing specific influences to be switched off without approximation.news.mit
Testing 24 ensembles on datasets ranging from 256 to more than 160,000 images, they found a consistent pattern: the larger the training set, the less any single image mattered. In some cases, removing an entire artist's body of work or every photograph of a given person produced no measurable change in the generated output.digitaltrends+1
"If you take away a piece of data and the output of the model doesn't change, then that piece of data didn't affect the output," Dai said in MIT's announcement. "So it doesn't make much sense to attribute the output to that piece of data."news.mit
The findings arrive as AI companies face mounting copyright claims. Disney The Walt Disney Company , NBCUniversal Comcast Corporation , and DreamWorks filed an intellectual property lawsuit last year against AI image-generator Midjourney, while The New York Times sued OpenAI and Microsoft in 2023. In the ongoing Andersen v. Stability AI case, plaintiffs have sought access to Midjourney's training datasets, alleging the company scraped images to mimic specific artists' styles.tech.yahoo+1
"We might have to rethink what intellectual property means," Dai told Semafor. "You can't just assume it, and the attribution link sort of vanishes."tech.yahoo
The researchers caution that their findings do not mean models never copy — memorization and reproduction of specific training examples can still occur. The study also examined only diffusion models; whether the same decay applies to large language models remains an open question.news.mit
James Grimmelmann, a law professor at Cornell Law School and Cornell Tech, said the paper "provides reason to think that attribution will fail for interesting models. Instead, technologists and courts will need to resort to other methods for assessing copying."theregister+1