00 · IN THREE MINUTES
The answer in three steps
- 1The researchers compared an output with versions produced after removing a training image or creator.
- 2Across datasets and metrics, the maximum change generally decreased as training data grew.
- 3This challenges single-example attribution; it does not show that training data as a whole is irrelevant or that copying never occurs.
01 · ATTRIBUTION WAS FRAMED AS A COUNTERFACTUAL
Attribution was framed as a counterfactual
A training item counts as causally influential if removing it changes the generated sample while prompt and random input are held fixed. Ordinary models require expensive retraining for every deletion.
02 · AN ENSEMBLE MADE DELETION EXACT
An ensemble made deletion exact
The team trained components on overlapping data subsets. Switching off every component exposed to one unit produced an ablated model without that unit’s influence, allowing many counterfactual outputs to be measured.
03 · INFLUENCE DECAYED WITH SCALE
Influence decayed with scale
Twenty-four ensembles covered datasets from 256 to 162,770 images. Pixel, semantic and additional similarity measures showed decreasing as training sets grew; small brute-force retraining tests supported the pattern.
Attribution was framed as a counterfactual
A training item counts as causally influential if removing it changes the generated sample while prompt and random input are held fixed. Ordinary models require expensive retraining for every deletion.
24 ensemblesAn ensemble made deletion exact
The team trained components on overlapping data subsets. Switching off every component exposed to one unit produced an ablated model without that unit’s influence, allowing many counterfactual outputs to be measured.
256–162,770 imagesInfluence decayed with scale
Twenty-four ensembles covered datasets from 256 to 162,770 images. Pixel, semantic and additional similarity measures showed decreasing counterfactual radius as training sets grew; small brute-force retraining tests supported the pattern.
leave-one-unit-outThe boundary matters
The paper studied attribution to an individual image, person or artist under its counterfactual definition. It does not rule out attribution to larger subsets, dataset-wide dependence, memorized rare outputs or other legal and ethical theories.
scope is specific04 · THE BOUNDARY MATTERS
The boundary matters
The paper studied attribution to an individual image, person or artist under its counterfactual definition. It does not rule out attribution to larger subsets, dataset-wide dependence, memorized rare outputs or other legal and ethical theories.
05 · SIMILARITY CAN MISLEAD
Similarity can mislead
The visually nearest training image need not be the cause of an output. In larger-data experiments, similarity-based attributions more often survived removal of the supposedly responsible example, exposing false attribution.
06 · SOURCES AND EVIDENCE
Sources and evidence
Claims are linked to foundational papers, standards or the primary study behind the update.
- 01Outputs of generative diffusion models are often unattributablePRIMARY STUDY ↗
Supports a defined mechanism, measurement or evidence boundary in this article.
- 02When AI art has no authorRESEARCH EXPLAINER ↗
Supports a defined mechanism, measurement or evidence boundary in this article.
