Edge Rewrite
// request.cf · coarse context

A page that knows where it met you.

Only coarse request metadata is shown. This demo does not display or persist visitor IP addresses.

Country
US
Cloudflare location
CMH
Connection
HTTP/2
Language
Not provided

Ray ID: a21e0a1f2f48a23e

Jump to content

Talk:Model collapse

Page contents not supported in other languages.
Add topic
From Wikipedia, the free encyclopedia
Latest comment: 1 month ago by Gnomingstuff in topic Possible AI use

Not better be a section within overfitting?

[edit]

I have just read the paper, and I am not sure if this might not be better as a section within the overfitting page. The process is essentially recurrent overfitting of the model, due to class imbalance in the input, leading to an exaggerated class imbalance in later stages. Unless I am missing something conceptually. Bastianhornung (talk) 09:32, 26 July 2024 (UTC)Reply

It is about training on synthetic data, including its own output. The media and academic coverage makes it more than notable enough for its own page. Wqwt (talk) 15:15, 2 January 2025 (UTC)Reply

Inbreeding

[edit]

I have heard this phenomenon colloquially referred to as "inbreeding" or "inbred AI" on reddit. Might be worth putting in the article. 68.237.60.88 (talk) 17:28, 30 January 2025 (UTC)Reply

Collapse of a single model or a succession of models?

[edit]

The intro doesn't clear up my confusion on this point: Is this a matter of a single model degrading over time, or of a succession of models, with the later ones performing worse than the earlier ones? My uneducated assumption is that once a model is trained, it becomes relatively fixed (unless retrained), so its performance won't degrade. A subsequent model trained (partially or fully) on the output of the first model, though, will perform worse. Is this correct? It would be good to clarify. Sharpner (talk 23:32, 19 February 2025 (UTC)Reply

Is the core issue 'synthetic vs. human' or 'grounded vs. un-grounded' data?

[edit]

The definition of model collapse often relies on the dichotomy between "synthetic data" and "human-generated data." However, this view is superficial. The fundamental distinction lies not in who generated the data, but in whether the data is "grounded" in reality. Grounded data derives from direct interactions with the world, whereas data generated by an AI model is the output of a statistical model of the world, not of the world itself (Goodfellow et al., 2016; von Helmholtz, 1860).

To strengthen this argument, we can draw parallels to human cognitive, social, and even neurological phenomena. Model collapse is the computational counterpart to what occurs in "echo chambers," where information degrades in closed loops (Arendt et al., 2021). It is even more analogous to the phenomenon of sensory deprivation. When a brain is deprived of new, real-world stimuli, it begins to generate its own perceptions—hallucinations—by recycling internal memories and patterns uncontrollably, a process underpinned by the brain's predictive nature and its reliance on internal models when external data is absent (Friston, 2010; Goldstein & Volkow, 2011).

In all these cases, the principle is the same: the degradation of information occurs in any system—biological or artificial—that is forced to learn from its own representations of the world, instead of from the world itself. Therefore, model collapse is not a problem of "synthetic vs. human data," but rather of isolated learning systems vs. open systems continuously grounded in reality.

References

Raphael2718 (talk) 12:40, 28 August 2025 (UTC)Reply

That is interesting, but seems to be in part original research. Do you have a reference dealing specifically with the topic, so we can verify it? TucanHolmes (talk) 08:02, 29 August 2025 (UTC)Reply

Section on mathematical models

[edit]

The "1D Gaussian" sub-section is, in its current version, difficult to comprehend even with a mathematical background (and might also have been generated using an LLM; I'm not sure). I will try to work out a more succinct description, but I'll have to see how far I get. Any help in rewriting this section, especially from experts currently working on this phenomenon, would be appreciated (since it's still quite recent). TucanHolmes (talk) 14:18, 11 November 2025 (UTC)Reply

Article issues

[edit]

I moved alternate names for the topic ("AI inbreeding", "AI cannabilism" etc) out of the nested footnote and into the main text, as the footnote was getting quite crowded and alternate names should be in the main text anyway.

I'm only a lay person and not well-versed in AI but some issues in the article include:

  • the sources given for the topic's alternate names could be improved. currently they consist only of the source title and the date. no author, publisher, website name etc. (i suspect this article was made prior to the AI ban...)
  • far too many sources given for the topic's alternate names. WP:CITEKILL
  • the sentence in the lede "some techniques have been proposed to mitigate the effect." Nothing about this is mentioned in the body of the article. hence it goes against WP:LEADFOLLOWSBODY

Putting this here as a brief to-do list for future me or any editor interested in contributing. If any of the issues are addressed, I would be very grateful. The good article criteria may also be helpful for editing. 23:41, 28 May 2026 (UTC) Acinonyxjubatusrex (talk) 23:41, 28 May 2026 (UTC)Reply

Possible AI use

[edit]

Concerns were raised previously about AI-generated content in this article, which is prohibited per WP:NEWLLM. This is a procedural talk page section to discuss it. There is no need to ping me in response. Read WP:AISIGNS for more information for why this is tagged; however, note that fixing the issue requires more than just removing the superficial signs. Gnomingstuff (talk) 19:38, 17 June 2026 (UTC)Reply