Edge Rewrite
// request.cf · coarse context

A page that knows where it met you.

Only coarse request metadata is shown. This demo does not display or persist visitor IP addresses.

Country
US
Cloudflare location
CMH
Connection
HTTP/2
Language
Not provided

Ray ID: a4602b7059d70555

Jump to content

Wikipedia:WikiProject AI Cleanup/How the AI noticeboard works

From Wikipedia, the free encyclopedia

This is rough guide to how the AI noticeboard works, outlining the rough life cycle of a case. It also contains advice for anyone that wishes to help out; there is more located at WP:AICGUIDE.

Cases (i.e. threads) usually stay on AINB for a few weeks, sometimes more. At time of writing, there are usually ~140 cases on AINB at any one time and only a handful or two of people working there, so if you've made a report please be patient while someone makes their way towards addressing it.

Analysis

[edit]

The goal of analysis is to determine whether someone has habitually misused LLMs, and the time period of usage (including a rough start date), in order to meet the evidential threshold for presumptive removal. Anything before 30 November 2022 (the release of ChatGPT) is extremely unlikely to be LLM-generated. Below is a rough guide to doing this in ascending order of how time-consuming the method is, formatted as a bolded list for your own amusement.

  • Edit filter logs: input their username into this group of edit filters, you can download and use User:Suffusion of Yellow/FilterDebugger so you don't have to scan each diff for what triggered the filter. This is often very useful for answering whether there has been LLM misuse and getting a very rough idea of the time period, though on its own is usually insufficient. You may want to remove the search filter and check for other edit filter hits like adding non-existent templates or citing Wikipedia.
  • Blatant WP:AISIGNS: scan some creations for red-linked categories,[note 1] templates, and "See also" entries; for markdown, chatbot communication, and bolded lists etc. It can be most effective to view the first revision of their creations using XTools (go to their contributions page, scroll to the bottom and click "Articles created", click on the timestamps).
  • Subtle WP:AISIGNS: scan some of their creations more closely for AI signs (unfortunately you just have to learn them). A lot of these also show up in human writing sometimes; it is the presence of several signs at the same time (i.e. in the same article/edit, across several articles/edits) that skyrockets the probability it is AI-generated. Unfortunately, the examples at WP:AISIGNS have become somewhat outdated with respect to recent models; if you're struggling you can use Pangram[note 2] or GPTZero (not withstanding WP:GPTZERO) as another 'sign'/data point. There are also signs of human writing to look out for.
  • Check for other indicators: check whether any of the links in the article appear to be hallucinated (i.e. use Link dispenser to find 404s and do a site search or check the Wayback Machine to see whether they previously existed/worked), check for incorrect ISBNs (can use ISBN Search) etc.
  • Review creations for WP:V: i.e. for instances where the sources don't support the text, implying the existence of hallucinations. This is often the most reliable way to determine whether an article is AI-generated, but also the most time-consuming. You can usually expect an AI-generated article to have over 1⁄3 of its claims fail verification (a rough estimate based on experience).

User conduct

[edit]

WP:AICGUIDE#Engaging with editors suspected of LLM use goes into more detail with respect to advice for engaging with the 'subject' of a case. A summary: don't bite the newcomers, and don't waste time and energy bickering. It is important to remember that being reported to a high-traffic noticeboard, for an issue potentially related to competency, is practically always a terrible experience, especially for newcomers. Other editors need to always stay conscious of this and aim to mitigate it where possible.

While the noticeboard is primarily for cleanup, it can handle conduct issues as well and regularly dishes out blocks for persistent LLM misuse or lack of communication past multiple warnings (you can use {{AINBA}}, but do so responsibly).

LLMPRODs

[edit]


Tracking pages and cleanup of ordinary edits

[edit]

Once a more-or-less exact start date has been established, a tracking page for cleanup (example) can be created using User:DVRTed/AINB-helper. The various situations and methods for removing AI-generated content are outlined (very roughly) below, formatted into an unnecessary table for your own amusement.

Situations and methods for presumptive removal of AI-generated content on Wikipedia
SituationMethod
The presumably-LLM edits are the latest edits to the articleRevert/restore an earlier version of the article[note 3]
The presumably-LLM edits are not the latest edits to the article, but the edits by other editors after it are all minor and 'insignificant' as it wereRestore an earlier version
The presumably-LLM edits are not the latest edits to the article, and the later edits are ‘significant’ (i.e. PAG-compliant additions of text or otherwise time-consuming edits). None of the conditions outlined in the below cell applyOption 1: manually restore an earlier version, copy-paste the later edits in if possible.

Option 2: use Who Wrote That? (WWT) or User:MSK/saucership to highlight content by the LLM-using editor that is still in the article. Edit the current version (in a new tab if using WWT) and remove all content that was highlighted, while replacing it with content from an older revision that was overwritten if necessary

The presumably-LLM edits are not the latest edits to the article, and the later edits are ‘significant’. One of the following also applies:
  • The LLM-generated addition is an important aspect of the topic
  • Edits by others altered the LLM-generated content too much such that it can’t just be removed
  • The version before the LLM edits is of bad quality
Usually these articles just get tagged with {{AI-generated}} because they are too time-consuming or require expertise (i.e. it’s left to local editors or future editors at AINB)

To clean it up you need to preserve the content in some form, reviewing the text and the sources for WP:V and reworking the content to be PAG-compliant. Sometimes you can remove most of it and just keep the basic facts.

On a rare occasion, an article's content was entirely replaced by lengthy AI-generated content a while ago, and it isn't reasonably possible to preserve later edits; here there's unfortunately not many other options but to restore the pre-LLM version in whole.

Archiving and the backlog

[edit]


Notes

[edit]
  1. ↑ You can do an EIA of the user with Bearcat (who thanklessly cleans up non-existent categories). However this won't catch instances when the user fixes it themselves.
  2. ↑ How to use Pangram: copy a passage of plain text from the article and paste it into Pangram, then remove any citations and headings, before clicking "Check for AI". Pangram analyses the text in segments, such that the percentage given is the percentage of the text that is thought to be AI-generated/AI-assisted/human-written (i.e. it is not the probability it is X).
    Pangram is configured to give reliable positive results (i.e. when saying something's AI-generated); as a corollary it gives high false negatives. It is also less reliable for short passages. On the free version, you only get 20 daily credits to use (1 credit = 100 words), so use them wisely.
  3. ↑ You may wish to copy-paste a preloaded edit summary from one of your user subpages, eg. restore earlier version per [[WP:LLMPRV]], see [[Wikipedia:AI noticeboard#SECTION TITLE]] while replacing "earlier" with the rough date of the restored revision.