Talk:Knowledge cutoff
Add topic| A request has been made for this article to be peer reviewed to receive a broader perspective on how it may be improved. Please make any edits you see fit to improve the quality of this article. |
| Knowledge cutoff was nominated as a Engineering and technology good article, but it did not meet the good article criteria at the time (August 30, 2025, reviewed version). There are suggestions on the review page for improving the article. If you can improve it, please do; it may then be renominated. |
| Knowledge cutoff was nominated as a Engineering and technology good article, but it did not meet the good article criteria at the time (August 10, 2025, reviewed version). There are suggestions on the review page for improving the article. If you can improve it, please do; it may then be renominated. |
| Knowledge cutoff was nominated as a Engineering and technology good article, but it did not meet the good article criteria at the time (July 20, 2025, reviewed version). There are suggestions on the review page for improving the article. If you can improve it, please do; it may then be renominated. |
| This article is rated C-class on Wikipedia's content assessment scale. It is of interest to the following WikiProjects: | ||||||||||||||||||||||||||||
| ||||||||||||||||||||||||||||
GA review
[edit]The following discussion is closed. Please do not modify it. Subsequent comments should be made on the appropriate discussion page. No further edits should be made to this discussion.
| GA toolbox |
|---|
| Reviewing |
- This review is transcluded from Talk:Knowledge cutoff/GA1. The edit link for this section can be used to add comments to the review.
Nominator: 16dvnk (talk · contribs) 03:58, 16 July 2025 (UTC)
Reviewer: David Eppstein (talk · contribs) 00:18, 20 July 2025 (UTC)
There are multiple unsourced statements. In some cases this is visible from a lack of footnote; in other cases, the next footnote covers only a later unrelated claim
- One sentence at end of section "Factors behind knowledge cutoffs"
- Two sentences at end of section "Knowledge gaps"
- One sentence at start of section "Historical context"
Twelve of 16 references appear not to meet our standards for reliable sources
- [1] Conductor, [8] ProjectPro (commercial sales site)
- [2-4, 7, 9-11] OpenAI, Anthropic, Google AI, Amazon primary sources about their own products
- [5] Otterly, blog
- [12] WP:FORBES
- [16] non-peer-reviewed preprint
Reference [6] (Brown et al.) appears reliable but cannot be used for the claim that it originated the idea of a knowledge cutoff (in the infobox); we need independent secondary sourcing for that.
Spot-checking the next use of [6], for the claims "Training large language models on static datasets is standard practice. This is necessary for achieving reproducibility and stability in performance evaluation." found no use of the words static and reproducible, and the only uses of the word stable referring to the hardware platform and not performance evaluation.
Reference [6] is also used for the claim "The practice of announcing a cutoff date became an industry standard for transparency after the release of GPT-3 in 2020." as a 2020 publication it seems an unlikely choice for a source about what became standard after 2020 and the word announce does not appear in it.
I conclude that this is very far from Good Article criterion 2 (sourcing) and falls under WP:GAFAIL #1. It was not ready for a Good Article nomination. —David Eppstein (talk) 00:25, 20 July 2025 (UTC)
GA review
[edit]The following discussion is closed. Please do not modify it. Subsequent comments should be made on the appropriate discussion page. No further edits should be made to this discussion.
| GA toolbox |
|---|
| Reviewing |
- This review is transcluded from Talk:Knowledge cutoff/GA2. The edit link for this section can be used to add comments to the review.
Nominator: 16dvnk (talk · contribs) 14:59, 10 August 2025 (UTC)
Reviewer: Grapesurgeon (talk · contribs) 17:56, 10 August 2025 (UTC)
Quick declining. This is the second nomination, and not much improvement since then. Seems AI-generated (draft originally declined for that reason), WP:CRITICISMSECTION, too many sections in article. Also "key people" field in infobox strange choice and unsourced. More issues but not worth getting into; article clearly just not ready yet. grapesurgeon (seefooddiet) (talk) 17:56, 10 August 2025 (UTC)
Another GA nomination
[edit]This article has failed the GA nomination 2 times. I think this hits the requirements, since the major issues like the sourcing issues, the AI issue, and the infobox issues are solved. I am nominating again to further improve this article. I welcome any feedback. Thank you! 16dvnk (talk) 12:01, 21 August 2025 (UTC)
GA review
[edit]The following discussion is closed. Please do not modify it. Subsequent comments should be made on the appropriate discussion page. No further edits should be made to this discussion.
| GA toolbox |
|---|
| Reviewing |
- This review is transcluded from Talk:Knowledge cutoff/GA3. The edit link for this section can be used to add comments to the review.
Nominator: 16dvnk (talk · contribs) 12:00, 21 August 2025 (UTC)
Reviewer: RoySmith (talk · contribs) 18:49, 30 August 2025 (UTC)
I'm afraid I'm going to have to quick-fail this. The first thing that jumped out at me was the use of Fox News as the most cited source in this article. See WP:FOXNEWSSCIENCE. For an article about a technical subject like this at the GA level, I would expect to see mostly sources that specialize in tech, and if any general audience media were used, at least only the highest quality such as the TIME source that was included.
In addition to that, there's an entire paragraph that's lifted almost verbatim from technologyreview.com. And my one foray into fact checking was to look at This is caused by the fact that almost all large language models are trained on static datasets, and training on newer data would cause a major price concern, given that training the most powerful large language models may soon cost over a billion dollars according to Time.[3]
for which I found that the source says nothing about static datasets.
I hate to sound harsh, but this is the third quick fail in a row. I strongly suggest you do not bring this back to WP:GAN. A forum like WP:PR might be a better place to get feedback from other editors. RoySmith (talk) 19:01, 30 August 2025 (UTC)
Peer review
[edit]
Previous GA review: Talk:Knowledge cutoff/GA1
I've listed this article for peer review because a couple of months ago, I GAed the article, but it has failed for obvious reasons. Since then, I kind of gave up on the article for a while. However, I have since gained for experience, and I have realized the sources should not be from news, but instead from reputable sources like peer reviewed articles, or the TIME or IBM sources. However, other than the sources, there are other things that need to be improved but unfortunately I do not have enough experience to pinpoint what exactly we need to improve. My aim is a very solid/ middle B tier/ slightly below GA which only needs small polishing. This wiki topic is pretty important, so that is why I allocate my time into this work. Thank you for your attention.
Sincerely, 16dvnk (talk) 11:48, 31 May 2026 (UTC)
- This article needs improvement following the WP:BCLASS criteria, but I think B-tier is definitely achievable. I'm no AI expert, so I don't know how comprehensive the article is, but it seems like it reasonably covers the topic. However, you should consult available literature to ensure that it is indeed a good summary. The sourcing definitely needs work, as does the writing. A few suggestions:
- Fox News is not (see WP:FOXNEWS) a reliable source for science-related topics, and in this case the article does not appear to support many of the statements it is attached to. In the article, many of the statements referencing a knowledge cutoff are reported statements from ChatGPT itself, and are therefore not reliable as a source. This source should be replaced, and proper citations found.
- Things described as "...a common technique..." or "...mostly-used in reference to..." must be supported by a source that directly identifies them as such.
- Do not include an "Overview" section. The lead should already provide an overview/summary of the article, so there is no need for an explicit section here. The body sections should provide further detail, rather than summarize the content of other sections.
- Suggestion: turn "Overview" into "Description", and explain briefly how LLM models are trained ahead of time, and how this introduces a cutoff based on when it was trained.
- A statement shouldn't be credited to "A research paper on arXiv...". Instead, look at the authors and publisher of the paper, determine its reliability and relevance, and credit the article's conclusions to its authors ("A study by Lorem et al at the University of Ipsum found that..."). Also, arXiv is not a peer-reviewed journal or publisher, but a preprint repository. Consider checking if this paper has been published in a journal, or finding an additional source.
- Avoid short statements like "Knowledge cutoffs create information gaps", as well as overly technical language. This guide provides a good explanation of how this can improve readability. You don't need to go into extreme detail, but jargon terms should be explained. This is especially important for terms like "fine-tuning" that mean specific technical concepts in an LLM context.
- The acronym RAG for retrieval-augmented generation should be explained in the article at first use, not just in the lead.
- Per MOS:SECTIONSTYLE, avoid section headings that reference the title. For example,
Effects of knowledge cutoffs
can simply be titledEffects
.
- H2so4aq (talk) 04:32, 9 June 2026 (UTC)
- Thank you very much for your review. Here is what I have addressed:
- The Fox News source has been replaced with a peer reviewed source. However, the peer reviewed source includes statements like "nursing", but I was unable to find any better source. Please check if it is a good enough source.
- I was unable to find any mention of "common" so I just removed it.
- The "overview" issue has been addressed, but it was still sourced with the nursing source, which is not fully about knowledge cutoffs.
- I have found a peer reviewed paper which was somehow exactly what I was trying to claim. Replaced with that. I have also addressed them in the correct format.
- The fine-tuning (and other examples I can find) example is fixed.
- Acronyms have either been removed/explained
- Have changed "Effects of knowledge cutoffs" to "Mitigation strategies"
- I have checked literature and I do think it is reasonably good.
- I have addressed the points you have given me. Thank you very much for the extensive review! Please review my changes and please tell me if I have to make any other changes. Thank you very much! 16dvnk (talk) 06:13, 12 June 2026 (UTC)
- Thank you for your changes, I've reviewed them to the best of my ability. The article still needs work but this is a great start. I have some additional suggestions:
- More detail is needed in a lot of parts of the article. A good example is the Continual learning section. It goes right into explaining that continual learning is a possible mitigation strategy, then starts explaining LoRA without actually stating what "continual learning" is. The issue with fine-tuning is likewise still not really fixed, as the article still doesn't explain what it is. WP:TECHNICAL has good information here too, as does MOS:JARGON. Basically you should aim to make the article reasonably understandable to someone without a significant AI background who can't click on the links to specific terms. This is one of the B-class criteria, so if you're aiming for that level this is one of the biggest issues to focus on.
- As for the Nursing in Critical Care paper, this is an excellent source. Good job finding this one. It is an academic paper (independent and reliable) that is directly relevant to the topic at hand. It also contains as its introduction a detailed description, written for a non-technical audience, of LLMs and what a knowledge cutoff is, starting by describing LLMs and their training process. These details should be worked into the article's lead and Description section. The rest of the paper is describing the relevance of knowledge cutoffs to a real-world scenario. That could (use your own discretion here) be added to the article, as an example of the impacts of knowledge cutoffs.
- Finding more sources like this will help a lot with the previous issue. I recommend closely reading WP:BACKWARD for a good description of how to approach this issue, and why finding more detailed sources will help immensely.
- If you're stuck, try to look through that paper's references and see if you can find more information that way.
- All sources should be properly formatted and attributed. For instance, the MIT Technology Review source has a credited author and listed publication date, but those aren't currently in the citation. Be sure to check all sources to try and include as much citation information as is reasonable (Author information, publication date, and identifiers are most important).
- Citations also don't need to be applied to every sentence, just to the end of whatever section of content is drawn from them (See WP:CITE and specifically WP:INLINECITE).
- Any relative vocabulary (eg.
While simpler for training...
) should only be used for comparison to something else (eg. While simpler for training than <other method>, ...) or be removed. The example shown could instead read, "While simple for training...". Such judgements should also be supported by a citation if possible. - Per WP:LINK, try not to repeat links too much. Links should be inserted as early as possible, at most once per major section.
- Per the manual of style, the 'See also' section should usually only contain links that are not linked to in the article body. I actually don't think any of the links currently present should be in a 'See also' section, so you may consider removing it altogether if none can be found.
- I don't think the cutoff dates for each model are particularly notable. You may consider whether to keep or remove that information, but I don't have a strong preference.
- Hope this all helps. H2so4aq (talk) 06:24, 15 June 2026 (UTC)
- (1/3)
- Thank you for the constructive feedback! I am also glad to hear that my previous edits worked as intended. I will address these issues right now:
- I have fixed the specific section you have listed AND other places I have found to be ambiguous to the casual users. I think it is good enough, but please review it again.
- Have expanded Lead and Description. Have not added the consequences yet. No new sources are added since I would like to do that in the next iteration. (if required) I would like to tackle the more fundamental issues, before adding new sources to polish the article. (since I feel like sources are good enough? i might be wrong)
- Was something I forgot to do since AFC. Fixed. (other than [13] where I am struggling to wayback machine the source)
- 16dvnk (talk) 12:14, 15 June 2026 (UTC)
- (2/3)
- Easy fix. Done.
- I scanned the article and did not find new instances. Removed a single "r"
- Done
- Done
- I feel like they are notable, and people search Wikipedia to find these dates anyways? I don't know enough about policies, but I will not remove them I will make some further work like More detail later
- 16dvnk (talk) 13:40, 15 June 2026 (UTC)
- (3/3)
- The "detail" issues are mostly fixed
- More high-quality (2) sources are cited
- Once, again thank you very much for the review, and please review my changes. 16dvnk (talk) 16:05, 15 June 2026 (UTC)
- @H2so4aq Thank you very much for your reviews. I think I have finally addressed your points fully, and I cannot think of anything more (other than some kind of expand, or a stronger "implications" section with other examples other than medicare) to improve on. Can you please take another look? Thank you very much! 16dvnk (talk) 13:41, 19 June 2026 (UTC)
- Thank you for your reply; I have seen your edits. I am a bit busy right now but will give a detailed response within a few days. H2so4aq (talk) 03:15, 20 June 2026 (UTC)
- I've taken a look over your edits and I have a few additional suggestions:
- I think restoring the phrase,
hallucinations, where the model generates plausible but verifiably false statements.
is more useful to the reader than "erroneous AI-generated content".- Perhaps a better description would be:
hallucinations, where the model generates confident but false statements.
- Perhaps a better description would be:
- Don't use terms like "real-world impact". Since this page isn't about a piece of fictional media, everything on this page is "real-world".
- Currently, that section is talking about the knowledge gap introduced by knowledge cutoffs, so it should be merged with Knowledge gaps.
- The impacts of knowledge gaps identified by the Cacioli and Eltaybani studies should be further emphasized/discussed. The sources you've found explicitly talk about knowledge cutoffs having serious implications for the use of AI, but I don't really see that importance being reflected in the article.
- For example, Cacioli's study (it's a preprint, keep that in mind) says, "These findings argue that model recency must be treated as a safety-critical attribute on par with alignment or interpretability." Eltaybani's study says, "Although some models can access the internet through browsing tools, their core reasoning and baseline assumptions remain anchored to their original training data."
- I think restoring the phrase,
- Good job improving the article so far. Looking at the WP:BCLASS criteria, I think the article right now meets all of them except for number 6. As far as content and sources, the article's looking pretty good. However, the largest place for improvement is making the language flow better and be more understandable to a non-technical audience.
- At this point, my ability to assess readability for general readers is limited by how much I've looked over the article; a second reviewer would be much more helpful. I recommend keeping the peer review open for a bit longer to see if anyone else offers their opinion, or you could reach out to another user directly to ask that they join the review. H2so4aq (talk) 20:48, 27 June 2026 (UTC)
- Thank you very much! I will try to find someone asap. Thank you for this interesting journey! I had a lot of fun improving this article :D 16dvnk (talk) 04:54, 29 June 2026 (UTC)
- All 3 issues addressed. I will now start to reach out to other users. I will continue improving this article! 16dvnk (talk) 13:41, 29 June 2026 (UTC)
- If anyone is interested in a second opinion, please leave your comments about this article. Thank you very much for your understanding.
- (pinging everyone who has ever reviewed this article, you may safely ignore this if you are not interested) @Reading Beans @TurboSuperA+ @David Eppstein @Grapesurgeon @RoySmith
- Once again, thank you to all these people and anyone who has helped make this article better. I apologize for any inconvenience.
- Sincerely,
- 16dvnk (talk) 12:19, 4 July 2026 (UTC)
- All 3 issues addressed. I will now start to reach out to other users. I will continue improving this article! 16dvnk (talk) 13:41, 29 June 2026 (UTC)
- Thank you very much! I will try to find someone asap. Thank you for this interesting journey! I had a lot of fun improving this article :D 16dvnk (talk) 04:54, 29 June 2026 (UTC)
- @H2so4aq Thank you very much for your reviews. I think I have finally addressed your points fully, and I cannot think of anything more (other than some kind of expand, or a stronger "implications" section with other examples other than medicare) to improve on. Can you please take another look? Thank you very much! 16dvnk (talk) 13:41, 19 June 2026 (UTC)
- (3/3)
- (2/3)
- (1/3)
- Thank you for your changes, I've reviewed them to the best of my ability. The article still needs work but this is a great start. I have some additional suggestions:
- Thank you very much for your review. Here is what I have addressed:
Happy birthday, Knowledge cutoff!
[edit]
Happy birthday, Knowledge cutoff! This is your first ever birthday. May you reach good article status sometime next year. All the best to you! Thank you for your service.
Sincerely,
- Requests for peer review
- Former good article nominees
- C-Class AfC articles
- AfC submissions by date/15 July 2025
- Accepted AfC submissions
- C-Class Artificial Intelligence articles
- High-importance Artificial Intelligence articles
- WikiProject Artificial Intelligence articles
- C-Class Technology articles
- WikiProject Technology articles
