Wikipedia:Link rot/URL change requests
This page is for requesting modifications to URLs, such as marking dead or changing to a new domain. Some bots are designed to fix link rot; they can be notified here. These bots include InternetArchiveBot and WaybackMedic. This page can be monitored by bot operators from other language wikis since URL changes are universally applicable.
US agencies
[edit]| This section is pinned and will not be automatically archived. |
90+ agencies identified as having web pages deleted during the Trump admin: https://asia.nikkei.com/static/vdata/infographics/deleted-website/
ω Awaiting further developments and time to go through them -- GreenC 16:46, 1 April 2025 (UTC)
White HouseDepartment of Health and Human ServicesDepartment of AgricultureUSAIDNational Park Serviceworker.govDepartment of LaborU.S. Agency for Global Media - usagm.govFederal Mediation and Conciliation Service (United States) - fmcs.govWoodrow Wilson International Center for Scholars - wilsoncenter.orgInstitute of Museum and Library Services - imls.govCommunity Development Financial Institutions Fund - cdfifund.govMinority Business Development Agency - mbda.govDepartment of Transportation - dot.gov- includes 11 agencies: FAA, FHWA, FMCSA, FRA, FTA, GLS, MARAD, NHTSA, OIG, OST, PHMSAEnvironmental Protection Agency - epa.govDepartment of Housing and Urban Development - hud.gov- Centers for Disease Control and Prevention
- Federal Emergency Management Agency
- National Institutes of Health
- General Services Administration
- Department of Homeland Security
- Department of Commerce
- employer.gov
- Office of the Assistant Secretary for Health
- sftool.gov
- Department of Energy
- Department of the Interior
- Department of Education
- NOAA
- Substance Abuse and Mental Health Services Administration
climate.gov- Department of Defense
- Health Resources & Services Administration
- AbilityOne Commission
- Department of State
- United States Patent and Trademark Office
- BOEM
- The Census Bureau
- CISA
- HUD User
- MILLENNIUM CHALLENGE CORPORATION
- performance.gov
- National Archives and Records Administration
- Bureau of Safety and Environmental Enforcement
- Federal Aviation Administration
Food and Drug Administration- House of Representatives
- Department of Justice
National Endowment for the Humanities- Department of the Treasury
- youth.gov
- American Climate Corps
- Federal Trade Commission
- Global Change Research Program
- NASA
- Administration for Community Living
- National Endowment for the Arts
- ATF
- Bureau of Indian Affairs
- Customs and Border Protection
- Consumer Financial Protection Bureau
- Consumer Product Safety Commission
- Office of the Director of National Intelligence
- Economic Development Administration
- Equal Employment Opportunity Commission
- Export-Import Bank of the United States
- FBI
- Federal Committee on Statistical Methodology
- Federal Housing Finance Agency
- geoplatform.gov
- Assistant Secretary for Technology Policy
- IRS
- National Labor Relations Board
- Office of Personnel Management
- Department of Veterans Affairs
- American Battle Monuments Commission
- Agency for Healthcare Research and Quality
americorps.gov- Advanced Research Projects Agency for Health
- Bonneville Power Administration
- cms.gov
- congress.gov
- digital.gov
- ENERGY STAR
- ej.gov
- farmers.gov
- medicalcountermeasures.gov
- peacecorps.gov
- Securities and Exchange Commission
- Social Security Administration
- stopbullying.gov
- Citizenship and Immigration Services
United States Interagency Council on Homelessness- workcenter.gov
Air ForceArmyNavy- Marine Corps
State government websites in the United States
[edit]| This section is pinned and will not be automatically archived. |
Whether or not I am supposed to add this here, I don't know, if I am not supposed to put it here, let me know. But to go along with whoever listed US federal government websites on this page. I also think there needs to be something done about sources citing the governors office websites in several states since the 2024 election. This probably needs IABOT or something to go in and archive all of them. Those include the office of governor websites in:
- Delaware
- Indiana
- Missouri
- New Hampshire
- New Jersey (since 2025)
- North Carolina
- North Dakota
- Possibly South Dakota (since Kristi Noem's resignation to become DHS secretary)
- Virginia (since 2025)
- Washington
- West Virginia
More states might need to be added. The reason why I am requesting these states (and really any state if you think about it) websites be archived by IABOT is because I've been noticing a couple of dead links on governor
In addition, ANY state government website that deals with a particular administration, there needs to be an archive link put up if its used in a citation. That goes for office of governor, their legislative sites, and anything else that may be deleted when the next administration is elected. Hurricane Clyde 🌀my talk page! 02:46, 8 January 2026 (UTC)
- User:Hurricane Clyde absolutely right. I started going through all 50 and completed 10: Wikipedia:Link_rot/URL_change_requests/Archives/2025/June#Five_USA and Wikipedia:Link_rot/URL_change_requests/Archives/2025/July#Five_USA_set_2. I gave up on CA it's too complicated. I pinned this thread so it is not archived and a reminder to keep going. All these government sites are very difficult: they have been in existence since the 1990s, have gone through continues changes and updates with new administrations and technology. They are fragmented with sub-domains under different departments and issues. But every one I did required a lot of updates they all need work. -- GreenC 21:43, 31 January 2026 (UTC)
ift.org.mx and cofece.mx
[edit]| This section is pinned and will not be automatically archived. |
These Mexican government agencies will likely be dissolved this month, and I'm not sure what will happen to references and other materials used within. Sammi Brie (she/her · t · c) 19:00, 10 October 2025 (UTC)
On hold pending dissolution. -- GreenC 04:21, 18 October 2025 (UTC)
- It took place on October 17 for the former and presumably a similar time for the latter. Sites are still up for now (but with the replacing agency's logo). Unclear whether they will use it or another page for their own business. I suspect the domain rpc.ift.org.mx (full of PDFs containing broadcasting technical information) will be retained intact at some other domain at some point. Sammi Brie (she/her · t · c) 06:32, 21 October 2025 (UTC)
- If it's a new agency at the same domain.. what does this mean we should do in terms of archiving URLs? Options are do nothing. Or treat all URLs as dead and add archive URLs. -- GreenC 16:06, 21 October 2025 (UTC)
- With the IFT and Cofece, the sites seem to be up as an archive. I am expecting the domain https://rpc.ift.org.mx/ —which contains most of our IFT citations—to move at some point, so put a pin in this thought for now. I will let you know when that happens. Sammi Brie (she/her · t · c) 18:01, 14 January 2026 (UTC)
- If it's a new agency at the same domain.. what does this mean we should do in terms of archiving URLs? Options are do nothing. Or treat all URLs as dead and add archive URLs. -- GreenC 16:06, 21 October 2025 (UTC)
- It took place on October 17 for the former and presumably a similar time for the latter. Sites are still up for now (but with the replacing agency's logo). Unclear whether they will use it or another page for their own business. I suspect the domain rpc.ift.org.mx (full of PDFs containing broadcasting technical information) will be retained intact at some other domain at some point. Sammi Brie (she/her · t · c) 06:32, 21 October 2025 (UTC)
Migration away from archive.today
[edit]Submitting a formal request to, where possible, replace links to archive.today with archive.org (or other suitable archive). Is this something that can be done with WP:WAYBACKMEDIC? I know that it is not always possible, but can we do what can be done automatically? I also agree with the advice to remove WP:EARLYARCHIVEs, where possible. Best, HouseBlaster (talk • he/they) 21:35, 20 February 2026 (UTC)
- Just here to add a breadcrumb to Wikipedia:Bot requests/Archive 88#archive.today cleanup for anyone finding this and interested in discussing. Dreamyshade (talk) 23:39, 22 February 2026 (UTC)
Not done yet. A handful of editors are picking away at it manually they should be given the opportunity over the next few months to see how they progress. Also, bot work is tricky it will take some time to develop soft-404 filters. -- GreenC 18:48, 27 February 2026 (UTC)
A few dead links
[edit]I was fixing some MCU-related pages and noticed that some of these stopped working. Not sure how common this is or what the fix is other than marking as dead.
Gonnym (talk) 08:43, 4 March 2026 (UTC)
- Some of these are still available under different URLs. For example, the denofgeek source was moved here, the nerdist source was moved to their archive, and the CBR article was moved here. ARandomName123 (talk)Ping me! 05:51, 14 March 2026 (UTC)
This request is kind of a can of worms because it is 5 different domains, each should have a separate request. And ARandomName123 has identified some excellent rules for moving links, there are likely more rules to be discovered. Looking at the numbers:
- comicbookresources.com 3,451 pages some convertible to cbr.com
- nerdist.com 2,182 pages some convertible to archive.nerdist.com
- denofgeek.com 6,799 pages maybe some convertible
- observationdeck.io9.com 31 pages, all dead
- hitfix.com 2,687 pages. This site ceased to exist in 2016, content was purchased by Uproxx Media Group in 2017, in 2018 Warner Music Group acquired Uproxx, in 2024 Uproxx re-acquired it from Warner, and today it is doing business as Uproxx Studios. I don't see the original HitFix content available at https://uproxx.com/ unless someone can find it, the domain is dead.
- I feel the need to note some nerdist sources that have working links at archive.nerdist don't have working content there. (Or more accurately, I had to switch a cite back to dead (actually... I'm going to go switch it to deviated because that's more accurate) because it's an audio interview whose file doesn't exist on its live archive.nerdist page but does exist on its Wayback Machine link.) This, of course, is not the bot's fault; I imagine it has no way to tell the difference between a working link with a working audio file on it and a working link with no audio file on it (and then some nerdist cites (like the Marvel one mentioned above) are text, so...). - Purplewowies (talk) 04:40, 17 April 2026 (UTC)
- "an audio interview whose file doesn't exist on its live archive.nerdist page but does exist on its Wayback Machine link." That's unusual! Usually with media files the Wayback link is broken and the live link works. Had I known this I could have scraped the page for audio content and added a rule to treat it as a dead link (add archive). If you want to show an example I can look into it. -- GreenC 22:56, 17 April 2026 (UTC)
- Amazing work with the bot. Thank you! Gonnym (talk) 08:56, 18 April 2026 (UTC)
- @GreenC: The one the bot changed in this edit, https://archive.nerdist.com/the-mutant-season-34-dylan-sprouse/, has an audio player on it that tries to request https://traffic.libsyn.com/themutantseason/TMS34_Dylan_Sprouse.mp3 (which is a dead link--the download and the play button both try to get this file/URL). The archive.org copy, on the other hand, did successfully archive the file (despite its attempt to load the player itself failing) if one goes through the download link on the archived page (something I was also fully expecting to not work). The specific archived link's download button was doing a 302 response at the time, but archive.org does redirect one to the archived audio file in this instance (here, for reference). I don't know if archive.org managed to properly archive everything like that (since my experience is the same as you regarding the vast majority of media files), but since Dylan and Cole Sprouse is on my watchlist I went down that rabbithole just to make sure the archive it already had was actually useful and all... (Seconding that you do do amazing work with this bot.) - Purplewowies (talk) 04:10, 19 April 2026 (UTC)
- wow surprised it worked. I'm not sure how to approach this because most editors will see the broken player and stop there, and not follow it through to the end, as you say a rabbit hole. Also the archive.org libsyn.com link didn't play for me in Firefox, but it worked in Chrome. If we go back to https://archive.nerdist.com/the-mutant-season-34-dylan-sprouse/ and look at the Download button it points to http://traffic.libsyn.com/themutantseason/TMS34_Dylan_Sprouse.mp3 which is pretty clear it was an MP3 file. The Wayback Machine has an archive page for that, it redirects back to the libsyn.com archive page here. Notice the "&expiration=1529532070" .. that is a self-destruct command, the URL works for a short period then stops working. The WaybackMachine captured it before it expired. This is sometimes done by sites to protect from web scrapers and piracy. So they are on record they don't want the content forever available. A messy problem on a couple levels. -- GreenC 06:06, 20 April 2026 (UTC)
- The bit about Firefox not playing it is interesting, as I did all that checking in Firefox. Huh! - Purplewowies (talk) 17:10, 20 April 2026 (UTC)
- wow surprised it worked. I'm not sure how to approach this because most editors will see the broken player and stop there, and not follow it through to the end, as you say a rabbit hole. Also the archive.org libsyn.com link didn't play for me in Firefox, but it worked in Chrome. If we go back to https://archive.nerdist.com/the-mutant-season-34-dylan-sprouse/ and look at the Download button it points to http://traffic.libsyn.com/themutantseason/TMS34_Dylan_Sprouse.mp3 which is pretty clear it was an MP3 file. The Wayback Machine has an archive page for that, it redirects back to the libsyn.com archive page here. Notice the "&expiration=1529532070" .. that is a self-destruct command, the URL works for a short period then stops working. The WaybackMachine captured it before it expired. This is sometimes done by sites to protect from web scrapers and piracy. So they are on record they don't want the content forever available. A messy problem on a couple levels. -- GreenC 06:06, 20 April 2026 (UTC)
- "an audio interview whose file doesn't exist on its live archive.nerdist page but does exist on its Wayback Machine link." That's unusual! Usually with media files the Wayback link is broken and the live link works. Had I known this I could have scraped the page for audio content and added a rule to treat it as a dead link (add archive). If you want to show an example I can look into it. -- GreenC 22:56, 17 April 2026 (UTC)
comicbookresources.com
[edit]- Enwiki
- Checked 3,579 pages and edited 3,023 pages. Moved 1,492 links to a new URL: 1,336 ruled mapped redirects, 156 ghost mapped redirects, Resolved 5,259 soft-404s. Removed 5
{{dead link}}. Added 24{{dead link}}. Switched 629|url-status=deadto live. Switched 1,038|url-status=liveto dead. Added 3,515 archive URLs (3,515 Wayback).
- Checked 3,579 pages and edited 3,023 pages. Moved 1,492 links to a new URL: 1,336 ruled mapped redirects, 156 ghost mapped redirects, Resolved 5,259 soft-404s. Removed 5
- IABot DB
- Updated about 7,000 links
Done -- GreenC 01:15, 17 April 2026 (UTC)
nerdist.com
[edit]- Enwiki
- Checked 1,985 pages and edited 1,047 pages. Moved 1,092 links to a new URL: 16 normal redirects, 1,076 ruled mapped redirects, Resolved 5 soft-404s. Removed 2
{{dead link}}. Switched 544|url-status=deadto live. Switched 6|url-status=liveto dead. Added 28 archive URLs (28 Wayback).
- Checked 1,985 pages and edited 1,047 pages. Moved 1,092 links to a new URL: 16 normal redirects, 1,076 ruled mapped redirects, Resolved 5 soft-404s. Removed 2
- IABot DB
- Updated about 1,400 links
Done -- GreenC 06:08, 17 April 2026 (UTC)
denofgeek.com
[edit]- Enwiki
- Checked 6,637 pages and edited 2,631 pages. Moved 2,164 links to a new URL: 1,032 normal redirects, 800 ruled mapped redirects, 332 ghost mapped redirects, Resolved 47 soft-404s. Removed 31
{{dead link}}. Added 28{{dead link}}. Switched 93|url-status=deadto live. Switched 357|url-status=liveto dead. Added 436 archive URLs (436 Wayback).
- Checked 6,637 pages and edited 2,631 pages. Moved 2,164 links to a new URL: 1,032 normal redirects, 800 ruled mapped redirects, 332 ghost mapped redirects, Resolved 47 soft-404s. Removed 31
- IABot DB
- Checked about 10,000 and updated 3,200
Done -- GreenC 23:49, 17 April 2026 (UTC)
observationdeck.io9.com
[edit]- Domain is dead and in 31 pages. Invoked IABot it should get most of them. -- GreenC 23:13, 17 April 2026 (UTC)
- Re-done with WaybackMedic on July 14 2026 (see same section name below) -- GreenC 23:56, 14 July 2026 (UTC)
hitfix.com
[edit]- Enwiki
- Checked 1,818 pages and edited 1,136 pages. Added 9
{{dead link}}. Switched 770|url-status=liveto dead. Added 865 archive URLs (865 Wayback).
- Checked 1,818 pages and edited 1,136 pages. Added 9
- IABot DB
- Updated about 1,000 links
Done -- GreenC 16:45, 18 April 2026 (UTC)
Iranian websites
[edit]| This section is pinned and will not be automatically archived. |
Waiting for current world events to resolve. Unclear if the sites will go back online. -- GreenC 04:57, 18 April 2026 (UTC)
aftabnews.ir
[edit]Dead when I visit it. Around 200 links GrapesRock (talk) 10:08, 12 March 2026 (UTC)
mashreghnews.ir
[edit]Around 500 links. GrapesRock (talk) 10:52, 12 March 2026 (UTC)
farsnews.com
[edit]Both the "www." and the "english." version seem broken. Around 1K links.
Edit: looks like it's blocked outside of Iran. GrapesRock (talk) 11:07, 12 March 2026 (UTC)
canoe.ca
[edit]I came across this site, which has been usurpsed by a gambling website that has blocked archivals from archive.org. I see this was discussed in the archives before, and archive.today snapshots were added as replacements for them. However, archive.today is now blocked, and the usurped URL is now present again. It seems this can be solved by changing the .ca to a .com, which archive.org is still able to archive, and is still owned by the original canoe.ca owners. Thanks, ARandomName123 (talk)Ping me! 05:47, 14 March 2026 (UTC)
- ARandomName123, hey that is great news! this (usurped) becomes this (dead) becomes this (archive). It's in over 3,000 pages, many still as archive.today eg. Hulk Hogan. Not sure yet how the bots structure can unwind and rewind, should be possible. -- GreenC 23:09, 17 April 2026 (UTC)
- Lol, forgot that in Jan 2024 I wrote a section about this domain's usurpation: Special:Diff/1191901478/1200150093 -- GreenC 00:41, 21 April 2026 (UTC)
This needs to be done in two passes: PASS 1 will convert canoe.ca -> canoe.com, remove the archive.today link, and remove any {{dead link}} or {{webarchive}} templates. It will only target canoe.ca citations that use archive.today. PASS 2 will add an archive.org URL for the new canoe.com URL, or add a {{dead link}}. -- GreenC 02:31, 21 April 2026 (UTC)
Enwiki - stage 1
- Pass 1 (0001-0050). Checked 50 pages and edited 49 pages. Removed 44 archive.today
- Pass 2 (0001-0050). Checked 50 pages and edited 34 pages. Resolved 60 soft-404s. Added 7
{{dead link}}. Added 38 archive URLs (38 Wayback).
- Pass 1 (0051-0500). Checked 450 pages and edited 358 pages. Removed 776 archive.today
- Pass 2 (0051-0500). Checked 450 pages and edited 326 pages. Resolved 834 soft-404s. Added 74
{{dead link}}. Added 725 archive URLs (725 Wayback).
- Pass 1 (0501-6289). Checked 5,789 pages and edited 4,606 pages. Removed 18
{{dead link}}. Removed 9,496 archive.today - Pass 2 (0501-6289). Checked 5,789 pages and edited 4,256 pages. Moved 1 links to a new URL: 1 ruled mapped redirects, Resolved 10,063 soft-404s. Added 1,190
{{dead link}}. Switched 2|url-status=liveto dead. Added 8,575 archive URLs (8,575 Wayback).
Enwiki - stage 2
This stage resets cases where |url= was already canoe.com but |archive-url= still has an archive.today
- Pass 1. edited 66 pages. Removed 105 archive.today
- Pass 2. Checked 61 pages and edited 60 pages. Resolved 307 soft-404s. Added 17
{{dead link}}. Added 88 archive URLs (88 Wayback).
Enwiki - stage 3
This stage targets archive.today links in {{usurped}} templates for conversion to canoe.com and archive.org
- Pass 1. Checked 2,338 pages and edited 1,738 pages. Changed 806 citation metadata. Unwound 1,422
{{usurped}}templates. - Pass 2. Checked 1,161 pages and edited 1,160 pages. Resolved 1,469 soft-404s. Added 537
{{dead link}}. Added 928 archive URLs (928 Wayback).
Enwiki - stage 4 Targets citations with a primary URL of slamwrestling.net and an archive.today URL of canoe.ca
- Checked 698 pages and edited 40 pages. Made live 57 URLs. Per WP:ATODAY the archive.today URLs are deleted - because the primary URLs are live no replacement archive.org is added.
Enwiki - stage 5 Manual fix edge cases
- Repaired about 10 citations
Done - Eliminated over 10,000 archive.today about 90% converted to archive.org the rest replaced with {{dead link}}. Includes citations that are {{usurped}}, CS1|2 templates, {{webarchive}}, and bare links. This domain represented 1.5% of archive.today links on enwiki. -- GreenC 16:27, 25 April 2026 (UTC)
Comments
[edit]I just noticed this via my watchlist. I'm pretty sure most, if not all these links are to Slam Wrestling. That site moved to slamwrestling.net quite some time ago. The regulars at WP:PW should have identified and/or fixed this problem long ago, since Slam is their top go-to source. Don't know why that didn't happen. Also don't know what advantage there is to preserving ancient archive links when the same content is found and can be archived at the current stable address. RadioKAOS / Talk to me, Billy / Transmissions 02:35, 22 April 2026 (UTC)
- Thanks for the feedback. Do you have an example of an old URL and new URL at slamwrestling.net ? I tried searching for some old article titles at slamwrestling.net without luck. Canoe had a lot of material including music, hockey etc.. and yes a lot of wrestling. -- GreenC 03:11, 22 April 2026 (UTC)
- Yes, my guess is that only about 5% were wrestling articles that can be found at slamwrestling.net, but it could be as high as 20%. Examples:
- Chris_Chetti#cite_note-23 at https://slamwrestling.net/report/overbooking-convicts-guilty-as-charged/ The original url was http://www.canoe.ca/SlamWrestlingArchive/jan10_guiltyascharged.html but GreenC bot removed it.
- Chris_Chetti#cite_note-31 at https://slamwrestling.net/report/confusion-reigns-at-guilty-as-charged/ --Jahalive (talk) 19:30, 22 April 2026 (UTC)
- I don't see a way to automate conversion, there is no information in the old URL that points to the new. -- GreenC 02:23, 23 April 2026 (UTC)
- The new URL's seem to be https://slamwrestling.net/report/ followed by the title, all lowercase, with a dash in between words. That conversion could be automated.--Jahalive (talk) 20:19, 13 May 2026 (UTC)
- I don't see a way to automate conversion, there is no information in the old URL that points to the new. -- GreenC 02:23, 23 April 2026 (UTC)
I'm getting whiplash watching your bots change things (related to canoe.ca/canoe.com), add archives that don't work, remove archives that do work, add dead-link and cbignore, and then remove them. It was bad enough dealing with the first wave of mistakes the other day, but this second wave is ridiculous.
These 3 just came in on my watchlist:
- [1] removed a perfectly good archive-url to wayback.
- [2] removed a perfectly good archive-url to wayback
- [3] Removed
{{Dead link|date=April 2026|fix-attempted=yes}}{{cbignore|bot=medic}}that some bot put there yesterday. This article had a series of edits made to it 22 April, multiple bots, most messing up, and I spent a lot of time fixing it. And now this that makes no sense.
What cbignore or bots-deny can I use that will stop your bots from continuing their mayhem? ▶ I am Grorp ◀ 20:41, 23 April 2026 (UTC)
- As explained above, it is a two step process. Step 1. Step 2. Final Diff. Your reverts of step 1 are actually creating a problem. I have manually repaired the problem your reverts created. -- GreenC 21:11, 23 April 2026 (UTC)
- Also, I just eliminated over 10,000 archive.today links for canoe.ca by replacing them with archive.org links to canoe.com .. this is a complex operation internally for a bot not designed for changing both the domain name and archive provider during a single edit. Thus it required making two edits. I can't tell where you got involved in the process, but it looks you have been undoing and reverting the bot which caused unexpected results. -- GreenC 21:20, 23 April 2026 (UTC)
- Huh, that's a ton of archive.today links actually. Looks like slam/jam are #6/#24 on the most archived domains at Wikipedia:Archive.today_guidance#Stats. Also thanks for writing the explanation up at Special:Diff/1191901478/1200150093, that's actually where I got this idea from. ARandomName123 (talk)Ping me! 23:04, 23 April 2026 (UTC)
- @GreenC: Your bot run was running at the same time as this bot run, and this edit messed up royally. How could I know that these two bots weren't running in tandem and I was supposed to wait a period of time for two runs of yours? All I saw was mess up after mess up with valid archive links being removed, or replaced with ones which didn't lead to valid source articles. I didn't click on the edit summary link that led to this discussion until today when these 3 articles showed up again in my watchlist. From this point, I'll stay hands-off for 48 hours for these 3 articles and then I'll check their results. ▶ I am Grorp ◀ 00:42, 24 April 2026 (UTC)
- Ok fair enough. My edit summary could have said 2-steps. The other problem is there are so many links, the lag between step 1 and step 2 was too long and other processes got involved. Should have had smaller batches. Lesson learned next time 2-step is needed for a large domain. -- GreenC 02:04, 24 April 2026 (UTC)
- The shorter-batch idea sounds good. If I'd seen two back to back edits, I'd have checked them together, fixed what was broken, and that would have been the end of that. Yes, the other bot did introduce a random edit that mucked up the double-run (and made me micro-check your other bot edits). ▶ I am Grorp ◀ 02:58, 24 April 2026 (UTC) Another suggestion would be to use two different edit summaries, where the first one mentions "1st of 2 bot runs" (or something similar). ▶ I am Grorp ◀ 03:00, 24 April 2026 (UTC)
- Ok fair enough. My edit summary could have said 2-steps. The other problem is there are so many links, the lag between step 1 and step 2 was too long and other processes got involved. Should have had smaller batches. Lesson learned next time 2-step is needed for a large domain. -- GreenC 02:04, 24 April 2026 (UTC)
- @GreenC: Your bot run was running at the same time as this bot run, and this edit messed up royally. How could I know that these two bots weren't running in tandem and I was supposed to wait a period of time for two runs of yours? All I saw was mess up after mess up with valid archive links being removed, or replaced with ones which didn't lead to valid source articles. I didn't click on the edit summary link that led to this discussion until today when these 3 articles showed up again in my watchlist. From this point, I'll stay hands-off for 48 hours for these 3 articles and then I'll check their results. ▶ I am Grorp ◀ 00:42, 24 April 2026 (UTC)
- Huh, that's a ton of archive.today links actually. Looks like slam/jam are #6/#24 on the most archived domains at Wikipedia:Archive.today_guidance#Stats. Also thanks for writing the explanation up at Special:Diff/1191901478/1200150093, that's actually where I got this idea from. ARandomName123 (talk)Ping me! 23:04, 23 April 2026 (UTC)
- Also, I just eliminated over 10,000 archive.today links for canoe.ca by replacing them with archive.org links to canoe.com .. this is a complex operation internally for a bot not designed for changing both the domain name and archive provider during a single edit. Thus it required making two edits. I can't tell where you got involved in the process, but it looks you have been undoing and reverting the bot which caused unexpected results. -- GreenC 21:20, 23 April 2026 (UTC)
irishexaminer.com/archives/
[edit]Some of these redirect to new links, but some don't. This redirect works for 2012 Waterford Crystal Cup. This should redirect to here for Cork Premier Senior Hurling Championship, but it's broken. If ghost redirects can be found, that'd be great. ~600 Thank you! MrLinkinPark333 (talk) 02:43, 26 March 2026 (UTC)
Enwiki
- Checked 625 pages and edited 592 pages. Moved 1,133 links to a new URL: 298 normal redirects, 813 ruled mapped redirects, 22 ghost mapped redirects, Added 29
{{dead link}}. Switched 7|url-status=deadto live. Switched 5|url-status=liveto dead. Added 293 archive URLs (293 Wayback).
IABot
- Checked about 800 URLs and updated 249
Done -- GreenC 01:55, 29 April 2026 (UTC)
dawn.com
[edit]Various links redirect to new URLs. Others are broken or working.
Redirects
- beta.dawn.com
- archives.dawn.com
Broken
- dawn.com/year/month/day/
Working
- herald.dawn.com
I think it'd be easier to check the entire website. 14k Thank you! MrLinkinPark333 (talk) 23:33, 29 March 2026 (UTC)
Enwiki
- Batch 1: 00001-03000: Checked 3,000 pages and edited 1,182 pages. Moved 2,049 links to a new URL: 173 normal redirects, 1,824 ruled mapped redirects, 52 ghost mapped redirects, Resolved 14 soft-404s. Added 61
{{dead link}}. Switched 78|url-status=deadto live. Switched 19|url-status=liveto dead. Added 92 archive URLs (92 Wayback).
- Batch 2: 03001-15090: Checked 12,090 pages and edited 4,820 pages. Moved 8,287 links to a new URL: 650 normal redirects, 7,442 ruled mapped redirects, 195 ghost mapped redirects, Resolved 53 soft-404s. Removed 1
{{dead link}}. Added 257{{dead link}}. Switched 339|url-status=deadto live. Switched 84|url-status=liveto dead. Added 499 archive URLs (499 Wayback).
Done -- GreenC 01:38, 30 April 2026 (UTC)
tvtonight.com.au
[edit]the link http://www.tvtonight.com.au/2013/11/timeshifted-wednesday-20-november-2013.html does not lead anymore to the actual rankings that https://web.archive.org/web/20141104212522/http://www.tvtonight.com.au/2013/11/timeshifted-wednesday-20-november-2013.html shows. https://tvtonight.com.au/2021/06/timeshifted-ratings-have-moved.html#search-open says that searching for "Timeshifted: Wednesday 20 November 2013" would work but it doesn't, show I'm not sure if these are still available on the site. Gonnym (talk) 09:00, 18 April 2026 (UTC)
- There are 233 pages with timeshifted URLs. Spot checks show anything with a date of November 2019 or later works. Prior dates do not work. I can add archives for the older dates. -- GreenC 01:53, 30 April 2026 (UTC)
Enwiki
- Checked 223 pages and edited 175 pages. Added 148
{{dead link}}. Switched 226|url-status=liveto dead. Added 1,516 archive URLs (1,516 Wayback).
IABot
- Checked about 2,000 and updated about 1,300
Done -- GreenC 23:24, 30 April 2026 (UTC)
chesterchronicle.co.uk
[edit]Most of them seem to just redirect normally to cheshire-live.co.uk (e.g. this to this from Diana, Princess of Wales). Some don't such as this from here where I found the new link by Googling the title (and there don't seem to be any helpful ghost redirects).
Around 450 articles in the domain. GrapesRock (talk) 15:15, 20 April 2026 (UTC)
- It looks like the Wayback Machine has not crawled the new site, or it's being blocked. I could match most of those dead link and archive URLs to a new live URL via a CDX inference check, but without Wayback captures not possible. For example this is not in the Wayback Machine; if it was, I could match the slug of the new URL ("following-in-brynleys-footsteps") with the original URL here (and in this case making adjustment for the 's = -s). -- GreenC 19:40, 2 May 2026 (UTC)
Enwiki
- Checked 455 pages and edited 436 pages. Moved 513 links to a new URL: 131 normal redirects, 381 ruled mapped redirects, 1 ghost mapped redirects, Resolved 151 soft-404s. Removed 1
{{dead link}}. Added 50{{dead link}}. Switched 11|url-status=deadto live. Switched 10|url-status=liveto dead. Added 77 archive URLs (77 Wayback).
IABot DB
- Updated about 350 URLs
Done -- GreenC 02:29, 4 May 2026 (UTC)
informador.com.mx
[edit]These URLs need a slight adjustment, from informador.com.mx to informador.mx, in order to redirect to their new url. For example, changing this to that redirects to the new URL here for Follow the Leader (Wisin & Yandel song). 730 articles. Thanks! MrLinkinPark333 (talk) 01:56, 22 April 2026 (UTC)
Enwiki
- Checked 732 pages and edited 673 pages. Moved 857 links to a new URL: 854 ruled mapped redirects, 3 ghost mapped redirects, Removed 2
{{dead link}}. Added 6{{dead link}}. Switched 160|url-status=deadto live. Switched 1|url-status=liveto dead. Added 7 archive URLs (7 Wayback).
IABot DB
- Checked and updated about 2,500 URLs
Done -- GreenC 15:03, 4 May 2026 (UTC)
bleacherreport.com
[edit]This has its content here. (changing "/amp/" to "/articles/" and removing the ".amp"). Some are on "syndication.bleacherreport.com", and for those also removing the "syndication." fixes it such as this to this.
There's 140 (with the /amp) (but some of them just have their archives with the /amp). Overall, there are 10.4K articles on the domain. GrapesRock (talk) 11:52, 22 April 2026 (UTC)
- I'll do the whole thing including the amp rules. -- GreenC 15:24, 4 May 2026 (UTC)
Enwiki
- Batch 00001-03000: Checked 3,000 pages and edited 1,131 pages. Moved 1,111 links to a new URL: 16 normal redirects, 1,070 ruled mapped redirects, 25 ghost mapped redirects, Resolved 21 soft-404s. Removed 23
{{dead link}}. Added 23{{dead link}}. Switched 5|url-status=deadto live. Switched 42|url-status=liveto dead. Added 249 archive URLs (249 Wayback).
- Batch 03001-10413: Checked 7,413 pages and edited 2,856 pages. Moved 3,036 links to a new URL: 142 normal redirects, 2,891 ruled mapped redirects, 3 ghost mapped redirects, Resolved 41 soft-404s. Removed 42
{{dead link}}. Added 44{{dead link}}. Switched 58|url-status=deadto live. Switched 88|url-status=liveto dead. Added 388 archive URLs (388 Wayback).
IABot DB
- Updated about 1,410 links
Done -- GreenC 21:30, 5 May 2026 (UTC)
indyweek.com/api
[edit]All dead, but there do exist missing redirects for seemingly all of them (as long as we have the title). I manually went through and changed all the ones that were tagged as dead to the correct links (e.g. this edit). I don't know whether your bot can actually find the correct urls, but if it tags all the ones that don't have archives, I'd be happy to manually redirect those.
90 or so articles. GrapesRock (talk) 12:44, 22 April 2026 (UTC)
GrapesRock: The bot tagged 40 URLs with {{dead link}}. -- GreenC 01:38, 6 May 2026 (UTC)
- Thanks, I've dealt with them all! Having the list of articles made it more convenient :-) GrapesRock (talk) 11:18, 6 May 2026 (UTC)
- Great, and thanks also. -- GreenC 23:17, 6 May 2026 (UTC)
Enwiki
- Checked 91 pages and edited 89 pages. Moved 85 links to a new URL: 85 ghost mapped redirects, Removed 5
{{dead link}}. Added 40{{dead link}}. Switched 9|url-status=deadto live. Added 8 archive URLs (8 Wayback).
IABot DB
- Updated about 190 URLs
Done -- GreenC 02:20, 6 May 2026 (UTC)
jmripl.com
[edit]Usurped. Legoktm (talk) 22:57, 24 April 2026 (UTC)
- jmripl.com --> repository.law.uic.edu for the John Marshall Review of Intellectual Property Law. Added to WP:JUDI. ClumsyOwlet (talk) 14:00, 27 April 2026 (UTC)
- ω Awaiting next judi batch. -- GreenC 01:46, 6 May 2026 (UTC)
louisville.com
[edit]Some are dead like this. Adding "archive." to the start of the url like this gives a live version of the link. There's also live links on the www. sub-domain, but I haven't been able to find any examples on Wikipedia.
Only around 100 articles, and some of them already have the "article." at the start GrapesRock (talk) 16:11, 28 April 2026 (UTC)
Enwiki
- Checked 99 pages and edited 82 pages. Moved 85 links to a new URL: 1 normal redirects, 84 ruled mapped redirects, Removed 3
{{dead link}}. Added 1{{dead link}}. Switched 46|url-status=deadto live. Added 3 archive URLs (3 Wayback).
IABot DB
- Checked and updated 140 URLs
Done -- GreenC 23:15, 6 May 2026 (UTC)
tribuna.com
[edit]This link from Jens Lehmann has its content here when the title is "La Flop XI degli ultimi 30 anni: chi di loro non avresti mai voluto vedere al Milan?" and the date is 24 April 2020. Seems plausible that some inferred mapped redirects could be found? Thanks!
There's 1K or so articles on the domain. GrapesRock (talk) 16:37, 5 May 2026 (UTC)
- This might take a couple days to retrieve the domain's CDX data because the API is very slow and domain is very large. If the target slug ("la-flop-xi-degli-ultimi-30..") was positioned /slug-name-date instead of /name-date-slug it wouldn't require a CDX download and could be done in hours instead of days. Just a limitation how the CDX API works. The inferred mapped redirect could work for some but I am seeing inconsistencies like this article dated 26 May 2019 but it has a URL date of 2020-03-06 .. the CDX method is grounded only to the slug it will be more precise. Also, this domain appears to be a mess of different types of problems, will see how it goes. -- GreenC 03:42, 7 May 2026 (UTC)
- The site has a CloudFlare security layer enabled which made it difficult. It's oddly maintained with inconsistencies which made programming rules difficult. I was able to get some but not as much as expected. -- GreenC 15:23, 10 May 2026 (UTC)
Enwiki
- Batch 001-500: Checked 500 pages and edited 172 pages. Moved 114 links to a new URL: 98 normal redirects, 11 ruled mapped redirects, 5 ruled inferred mapped redirects, Removed 1
{{dead link}}. Added 7{{dead link}}. Switched 4|url-status=deadto live. Switched 4|url-status=liveto dead. Added 78 archive URLs (78 Wayback).
- Batch 501-972: Checked 472 pages and edited 170 pages. Moved 125 links to a new URL: 107 normal redirects, 13 ruled mapped redirects, 5 ruled inferred mapped redirects, Removed 3
{{dead link}}. Added 8{{dead link}}. Switched 5|url-status=deadto live. Switched 7|url-status=liveto dead. Added 53 archive URLs (53 Wayback).
IABot DB
- Updated about 750 links
Done -- GreenC 15:23, 10 May 2026 (UTC)
chemicalland21.com
[edit]This website which is used as a reference in ~150 articles has apparently been usurped or hijacked and now goes to msbscakery.hk. All external links to chemicalland21.com should be archived or removed, I think. Marbletan (talk) 12:35, 6 May 2026 (UTC)
- ω Awaiting next WP:JUDI batch -- GreenC 15:25, 10 May 2026 (UTC)
dtnext.in
[edit]Take this url from Anna University. If you lowercase-ify everything, remove the numbers and the .vpf you get this which is where the content is. You also get a working link with identical content, if you just remove the .vpf (e.g. this). The simplest rule seems to be just remove the .vpf and lowercase-ify everything (since that's where it will 301 to).
Edit to add this paragraph: I don't think I realized that the url in the previous paragraph already redirects to a working url (the url where you remove the numbers and the .vpf and lowercase-ify everything).
Unfortunately some urls such as this from Bruno Mars don't have the full article title in the url. In that case, doing instead this url works. That being said, from the spattering I've looked, all the incomplete urls all end with -.vpf. Thank you :-)!
There are 400 or so articles with the .vpf (though this overcounts since not all results have the dtnext.in url with the .vpf) and 1170 total articles on the domain. GrapesRock (talk) 12:43, 7 May 2026 (UTC)
- Will do, the whole domain. I did it before in Nov 2022, but the bot and techniques back then were more primitive. The Bruno Mars link has four paths, 1 dead and 3 live: 1, 2, 3, 4. I'll randomly choose a working link based on what is available in the CDX records. -- GreenC 16:59, 10 May 2026 (UTC)
Enwiki
- Checked 1,170 pages and edited 298 pages. Moved 300 links to a new URL: 13 normal redirects, 210 ruled mapped redirects, 7 ghost mapped redirects, 70 ruled inferred mapped redirects, Resolved 351 soft-404s. Removed 22
{{dead link}}. Added 10{{dead link}}. Switched 256|url-status=deadto live. Switched 3|url-status=liveto dead. Added 5 archive URLs (5 Wayback).
IABot DB
- Checked 590 links and updated 68
Done -- GreenC 20:54, 10 May 2026 (UTC)
americanradiohistory.com
[edit]They all seem to 301 to links on worldradiohistory.com. I thought it was a simple replace, but this redirects to this which turns turns "Archive-Billboard" into "Archive-All-Music/Billboard". However, the smattering that I've checked all seem to successfully redirect.
Over 8500 articles. GrapesRock (talk) 16:45, 7 May 2026 (UTC)
- 3500 dead links was a lot more than I was expecting! I would've examined the domain closer had I known. I have now done this closer examination :-).
- It seems like ones with OCR-Page near the end of the url cause problems. But some of them certainly are available. For instance this is dead, but its content seems to be here. It is also available here. Not all OCR-Page links are dead such as this one. All of these are from this edit.
- This link from Betty Clooney exists here. The only thing that prevented it from redirecting was the lowercase b in billboard, with this url successfully redirecting.
- This link from Bess Johnson is available here.
- This link from Best I Ever Had (Grey Sky Morning) is available here ("12-02a.pdf" -> "12-02.pdf"). I edited this one directly since I feel like that a was probably accidentally inserted at some point.
- This link from Charles E. Apgar exists here (this is the only wireless age one so I just changed it directly)
- This url from Charles & Diana: A Royal Love Story exists here.
- This exists here.
- This from Denise Darcel exists here. (I can fix these manually, there's only 18 articles)
- This exists here.
- This exists here.
- Here are some rules I notice from above that worked on the first relevant link I found:
- americanradiohistory.com/hd2/IDX-Business/Music/Archive-Billboard-IDX/IDX/ into www.worldradiohistory.com/Archive-All-Music/Billboard/ with the "OCR-Page-XXXX.pdf" turned into ".pdf"
- americanradiohistory.com/Archive-BC/ into www.worldradiohistory.com/Archive-All-BC/Broadcasting-Magazine
- americanradiohistory.com/Archive-TV-Radio-Age/Issues/ to worldradiohistory.com/Archive-TV-Radio-Age/
- Removing "-Page-XXXX" sometimes fixes things.
- lowercase-ifying .PDF to .pdf sometimes fixes things
- sometimes ".pdf" -> "-N.pdf" fixes things (between July 31 1993 and December 17, 1994 every other issue of Billboard Magazine was a newspaper I think?)
- sometimes ".o.pdf" -> ".pdf" fixes things
- I don't know how hard it would be, but it would be nice if the page information was added to the cite with page=whatever.
- It seems like this domain is simply a huge mess; I think ignoring the page shenanigans, the last bit stays constant and it is just the path that changes. I don't know how useful knowing that is though. GrapesRock (talk) 17:46, 11 May 2026 (UTC)
- 18531+3500 = 22031 then 3500/22031 = 16% or a 84% conversion rate (80/20 Rule) on first pass. I'll try the new rules next. -- GreenC 19:28, 11 May 2026 (UTC)
The results for Pass 2 (pages with {{dead link}}) are disappointingly low. Possibly due to my rules, and I think the site is just lots of edge cases. If you want to keep trying, below is the code I used for the rules. You can post the code to AI. Ask it to add new rules; send the new rules code block to me, and I'll re-run it. -- GreenC 05:38, 12 May 2026 (UTC)
@GreenC: My AI is telling me that the code was bugged since the nested sub is modifying the original rather than the modified copy and it had me change a line like "sub("(?i)-?(OCR-)?Page-[0-9]{4}", "", d)" into "sub("(?i)-?(OCR-)?Page-[0-9]{4}", "", dd)". I've also added a couple new rules:
New Code
|
|---|
subs(".PDF", ".pdf", newurl)
if awk.match(newurl, "americanradiohistory[.]com/hd2/IDX-Business/Music/Archive-Billboard-IDX/IDX/", d) > 0:
var old = d
var s = newurl
subs(d, "worldradiohistory.com/Archive-All-Music/Billboard/", d)
subs(old, d, s)
if awk.match(s, "(?i)-?(OCR-)?Page-[0-9]{4}[.]pdf", dd) > 0:
var old2 = dd
var ss = s
sub("(?i)-?(OCR-)?Page-[0-9]{4}", "", dd)
subs(old2, dd, ss)
mapredirectTable[ss] = 1
else:
mapredirectTable[s] = 1
# ---------------------------------------------------------
# NEW LOGIC: Archive-BC-IDX to Archive-BC (with dynamic years)
# Handles: /Archive-BC-IDX/94-OCR/BC-1994-11-21-Page-0008.pdf
# Handles: /Archive-BC-IDX/63-OCR/1963-01-14-BC-0009.pdf
# ---------------------------------------------------------
if awk.match(newurl, "americanradiohistory[.]com/Archive-BC-IDX/", dummy) > 0:
var s = newurl
# 1. Swap the domain (adding the www. as requested)
subs("americanradiohistory.com", "www.worldradiohistory.com", s)
# 2. Rebuild the directory path and dynamically extract the year.
# Use gsub here because awk.nim's sub does not evaluate regex capture groups ($1, $3)
# Group 1 ($1): Captures the full filename prefix (e.g., "BC-1994-11-21" or "1963-01-14-BC")
# Group 2 ($2): Captures "BC-" if it exists at the start
# Group 3 ($3): Captures exactly the 4-digit year (e.g., "1994" or "1963")
# Group 4 ($4): Captures "-BC" if it exists at the end of the date
gsub("/Archive-BC-IDX/[0-9]+-OCR/((BC-)?([0-9]{4})-[0-9]{2}-[0-9]{2}(-BC)?)", "/Archive-BC/BC-$3/$1", s)
# 3. Strip the trailing page number
# sub is fine here since there are no capture groups in the replacement string
# The (Page-)? makes the word "Page-" optional, so it catches both
# "-Page-0008.pdf" and "-0009.pdf" smoothly.
sub("(?i)-?(OCR-)?(Page-)?[0-9]{4}[.]pdf", ".pdf", s)
# Add to table
mapredirectTable[s] = 1
if awk.match(newurl, "americanradiohistory[.]com/Archive-BC/", d) > 0:
var old = d
var s = newurl
subs(d, "worldradiohistory.com/Archive-All-BC/Broadcasting-Magazine/", d)
subs(old, d, s)
if awk.match(s, "(?i)-?(OCR-)?Page-[0-9]{4}[.]pdf", dd) > 0:
var old2 = dd
var ss = s
sub("(?i)-?(OCR-)?Page-[0-9]{4}", "", dd)
subs(old2, dd, ss)
mapredirectTable[ss] = 1
# NEW LOGIC: Catch the -BC.pdf ending and generate both variants
elif awk.match(s, "(?i)[0-9]{4}-[0-9]{2}-[0-9]{2}-BC[.]pdf", dd) > 0:
var old2 = dd
# 1. Dashed version (e.g., BC-1982-10-04.pdf)
var s1 = s
var d1 = dd
sub("(?i)-BC", "", d1)
d1 = "BC-" & d1
subs(old2, d1, s1)
mapredirectTable[s1] = 1
# 2. %20 version (e.g., BC%201982%2010%2004.pdf)
var s2 = s
var d2 = dd
sub("(?i)-BC", "", d2)
subs("-", "%20", d2)
d2 = "BC%20" & d2
subs(old2, d2, s2)
mapredirectTable[s2] = 1
else:
mapredirectTable[s] = 1
if awk.match(newurl, "americanradiohistory[.]com/Archive-TV-Radio-Age/Issues/", d) > 0:
var old = d
var s = newurl
subs(d, "worldradiohistory.com/Archive-TV-Radio-Age/", d)
subs(old, d, s)
if awk.match(s, "(?i)-?(OCR-)?Page-[0-9]{4}[.]pdf", dd) > 0:
var old2 = dd
var ss = s
sub("(?i)-?(OCR-)?Page-[0-9]{4}", "", dd)
subs(old2, dd, ss)
mapredirectTable[ss] = 1
else:
mapredirectTable[s] = 1
# https://www.americanradiohistory.com/hd2/IDX-Site-Early-Radio/Archive-Wireless-World-IDX/80s/Wireless-World-1986-02-OCR-Page-0028.pdf
# to https://www.americanradiohistory.com/Archive-Wireless-World/80s/Wireless-World-1986-02.pdf
if awk.match(newurl, "americanradiohistory[.]com/hd2/", d) > 0:
var s = newurl
# Note: Depending on your specific regex engine implementation under the hood,
# the backreference might need to be "\\1" instead of "$1".
gsub("/hd2/[^/]+/([^/]+)-IDX/", "/$1/", s)
sub("(?i)-?(OCR-)?Page-[0-9]{4}", "", s)
mapredirectTable[s] = 1
if awk.match(newurl, "americanradiohistory[.]com/Archive-Billboard/", d) > 0:
var s = newurl
# Safely target the filename prefix by including the slash and hyphen
sub("/BB-", "/Billboard-", s)
mapredirectTable[s] = 1
subs("americanradiohistory.com", "worldradiohistory.com", newurl)
# https://www.americanradiohistory.com/hd2/IDX-Business/Music/Archive-Billboard-IDX/IDX/30s/1936/BB-1936-12-05-OCR-Page-0034.pdf
if awk.match(newurl, "(?i)-?(OCR-)?Page-[0-9]{4}[.]pdf", d) > 0:
var old = d
var s = newurl
sub("(?i)-?(OCR-)?Page-[0-9]{4}", "", d)
subs(old, d, s)
mapredirectTable[s] = 1
if awk.match(newurl, "(?i)[0-9][.][a-z][.]pdf", d) > 0:
var old = d
var s = newurl
sub("(?i)[.][a-z]", "", d)
subs(old, d, s)
mapredirectTable[s] = 1
if awk.match(newurl, "(?i)[0-9][a-z][.]pdf", d) > 0:
var old = d
var s = newurl
sub("(?i)[a-z][.]", ".", d)
subs(old, d, s)
mapredirectTable[s] = 1
if newurl ~ "/billboard":
var s = newurl
subs("/billboard", "/Billboard", s)
mapredirectTable[s] = 1
# ---------------------------------------------------------
# NEW LOGIC: Billboard Newspaper Variants (-N.pdf)
# Date range: July 31 1993 to December 17 1994
# ---------------------------------------------------------
# First, ensure this is actually a Billboard URL
if awk.match(newurl, "(?i)billboard", dummy) > 0:
# Match the specific YYYY-MM-DD.pdf format at the end of the URL
if awk.match(newurl, "[0-9]{4}-[0-9]{2}-[0-9]{2}[.]pdf", d) > 0:
# 'd' will hold exactly "1993-08-07.pdf".
# We slice the first 10 characters to isolate the date string
var dateStr = d[0..9]
# Lexicographical comparison works perfectly for ISO-formatted dates
if dateStr >= "1993-07-31" and dateStr <= "1994-12-17":
var s = newurl
var old = d
var d_new = d
# Swap .pdf for -N.pdf in our matched filename chunk
# (Using (?i) to make it case-insensitive in case of .PDF)
sub("(?i)[.]pdf", "-N.pdf", d_new)
# Inject the updated filename back into the full URL
subs(old, d_new, s)
# Add the new variant to the test table
mapredirectTable[s] = 1
|
I believe the new rules that I've added are
- a more general handling of hd2 (which indicates specific pages cited)
- A new rule for Archive-BC domains
- Try replacing BB- with Billboard-
- Try adding -N.pdf when it is between the relevant dates for Billboard Magazine.
Also, one thing that it was unsure about is whether Nim/Awk used \1 or $1. This was only used in 2 places (and it said "Note:" before those instances).
Could you try it out when you get the chance? Thanks! GrapesRock (talk) 10:50, 14 May 2026 (UTC)
- GrapesRock: This is great, thank you! I wrote the awk.nim library and sub() does not support capture groups but gsub() does. I need to fix sub(). Also this rules code is not my best day, it was a quick fix and that d vs. dd is a clear bug. Thanks for the running it through AI and the new updates. I'll make a Pass 3. -- GreenC 21:24, 14 May 2026 (UTC)
- Pass 3 done. Not sure how much further you want to go. Could be many edge cases spread thinly. Or legit dead links. -- GreenC 05:42, 15 May 2026 (UTC)
Rules code
|
|---|
subs(".PDF", ".pdf", newurl)
if awk.match(newurl, "americanradiohistory[.]com/hd2/IDX-Business/Music/Archive-Billboard-IDX/IDX/", d) > 0:
var old = d
var s = newurl
subs(d, "worldradiohistory.com/Archive-All-Music/Billboard/", d)
subs(old, d, s)
if awk.match(s, "(?i)-?(OCR-)?Page-[0-9]{4}[.]pdf", dd) > 0:
var old2 = dd
var ss = s
sub("(?i)-?(OCR-)?Page-[0-9]{4}", "", d)
subs(old2, dd, ss)
mapredirectTable[ss] = 1
else:
mapredirectTable[s] = 1
if awk.match(newurl, "americanradiohistory[.]com/Archive-BC/", d) > 0:
var old = d
var s = newurl
subs(d, "worldradiohistory.com/Archive-All-BC/Broadcasting-Magazine/", d)
subs(old, d, s)
if awk.match(s, "(?i)-?(OCR-)?Page-[0-9]{4}[.]pdf", dd) > 0:
var old2 = dd
var ss = s
sub("(?i)-?(OCR-)?Page-[0-9]{4}", "", d)
subs(old2, dd, ss)
mapredirectTable[ss] = 1
else:
mapredirectTable[s] = 1
if awk.match(newurl, "americanradiohistory[.]com/Archive-TV-Radio-Age/Issues/", d) > 0:
var old = d
var s = newurl
subs(d, "worldradiohistory.com/Archive-TV-Radio-Age/", d)
subs(old, d, s)
if awk.match(s, "(?i)-?(OCR-)?Page-[0-9]{4}[.]pdf", dd) > 0:
var old2 = dd
var ss = s
sub("(?i)-?(OCR-)?Page-[0-9]{4}", "", d)
subs(old2, dd, ss)
mapredirectTable[ss] = 1
else:
mapredirectTable[s] = 1
subs("americanradiohistory.com", "worldradiohistory.com", newurl)
# https://www.americanradiohistory.com/hd2/IDX-Business/Music/Archive-Billboard-IDX/IDX/30s/1936/BB-1936-12-05-OCR-Page-0034.pdf
if awk.match(newurl, "(?i)-?(OCR-)?Page-[0-9]{4}[.]pdf", d) > 0:
var old = d
var s = newurl
sub("(?i)-?(OCR-)?Page-[0-9]{4}", "", d)
subs(old, d, s)
mapredirectTable[s] = 1
if awk.match(newurl, "(?i)[0-9][.][a-z][.]pdf", d) > 0:
var old = d
var s = newurl
sub("(?i)[.][a-z]", "", d)
subs(old, d, s)
mapredirectTable[s] = 1
if awk.match(newurl, "(?i)[0-9][a-z][.]pdf", d) > 0:
var old = d
var s = newurl
sub("(?i)[a-z][.]", ".", d)
subs(old, d, s)
mapredirectTable[s] = 1
if newurl ~ "/billboard":
var s = newurl
subs("/billboard", "/Billboard", s)
mapredirectTable[s] = 1
# http://www.americanradiohistory.com/Archive-BC/BC-1982/1982-10-04-BC.pdf
if awk.match(newurl, "(?i)[0-9]{4}-[0-9]{2}-[0-9]{2}-BC[.]pdf", d) > 0:
var old = d
var s = newurl
sub("(?i)-BC", "", d)
d = "BC-" & d
subs(old, d, s)
mapredirectTable[s] = 1
|
Enwiki
- Pass 1 Checked 8,721 pages and edited 8,652 pages. Moved 18,531 links to a new URL: 11,647 normal redirects, 6,884 ruled mapped redirects, Resolved 6 soft-404s. Removed 177
{{dead link}}. Added 3,501{{dead link}}. Switched 25|url-status=deadto live. Switched 13|url-status=liveto dead.
- I'll make an additional pass or passes on the 3,501
{{dead link}}- some might be fixable by looking for worldradiohistory.com versions at the Wayback Machine.
- I'll make an additional pass or passes on the 3,501
- Pass 2 Checked 1,718 pages and edited 93 pages. Moved 127 links to a new URL: 127 ruled mapped redirects, Resolved 4 soft-404s. Removed 85
{{dead link}}. Added 3,399{{dead link}}.
- Pass 3 Checked 1,718 pages and edited 359 pages. Moved 689 links to a new URL: 689 ruled mapped redirects, Resolved 4 soft-404s. Removed 625
{{dead link}}. Added 2,713{{dead link}}.
IABot DB
- Checked about 18,000 links and updated about 3,200
rappler.com
[edit]At some point they switched from Day-Month-Year to -Month-Day-Year in their urls. For instance this is now present here. The final word may need to pluralized such as this to this (this could be somebody accidentally deleted "s-" though? not sure). Seems like they sometimes change the sections too (nation -> philippines on one of them), but from what I've checked there's 301's available after switching to MDY.
Around 9.3K articles on the domain. GrapesRock (talk) 20:53, 7 May 2026 (UTC)
Enwiki
- Checked 9,347 pages and edited 5,840 pages. Moved 10,983 links to a new URL: 2,929 normal redirects, 8,047 ruled mapped redirects, 7 ghost mapped redirects, Resolved 43 soft-404s. Removed 2
{{dead link}}. Added 75{{dead link}}. Switched 426|url-status=deadto live. Switched 19|url-status=liveto dead. Added 209 archive URLs (209 Wayback).
IABot DB
- Not done, due to pct of dead links vs. number working via redirects.
Done -- GreenC 02:01, 13 May 2026 (UTC)
money.cnn.com
[edit]Fortune magazine links were converted in 2024 to money.cnn.com. However, this subdomain is now broken. I did not find any new URLs. ~7600. Thank you! MrLinkinPark333 (talk) 00:27, 10 May 2026 (UTC)
- Major loss. According to AI: Natively, the content can only be found in two places: The Fortune Digital Archive: Available exclusively to paid subscribers who log into fortune.com and navigate to the "E-Magazine" viewer (which is powered by PressReader). Institutional Databases: Licensed out cover-to-cover to academic and library repositories, primarily EBSCO's Fortune Magazine Archive. That leaves Archives as the last option. The archives might be under archive.fortune.com or money.cnn.com .. the path portion of the URL remains the same. -- GreenC 20:52, 10 May 2026 (UTC)
- archive.fortune.com leads to a paywall, so I don't know if they're accessible there. Money.cnn.com redirects to CNN Business. You might get lucky with ghost archives if they were moved to either site. MrLinkinPark333 (talk) 21:58, 10 May 2026 (UTC)
- Treated as a dead site, everything archived. Did a sampling test run, no ghost redirs detected. -- GreenC 19:04, 14 May 2026 (UTC)
- archive.fortune.com leads to a paywall, so I don't know if they're accessible there. Money.cnn.com redirects to CNN Business. You might get lucky with ghost archives if they were moved to either site. MrLinkinPark333 (talk) 21:58, 10 May 2026 (UTC)
Enwiki
- Checked 7,656 pages and edited 7,013 pages. Added 357
{{dead link}}. Switched 2,194|url-status=liveto dead. Added 6,765 archive URLs (6,765 Wayback).
IABot DB
- Checked 12,022 links and updated 12,016
Done -- GreenC 01:35, 15 May 2026 (UTC)
bt.com.bn
[edit]Redirects to a gambling site. Most of them already have archives, but a few like Telisai–Lumut Highway don't. I request adding this to the JUDI list. 460. Thank you! MrLinkinPark333 (talk) 20:24, 10 May 2026 (UTC)
- ω Awaiting next WP:JUDI batch. -- GreenC 16:30, 15 May 2026 (UTC)
archaeology.org
[edit]Articles before December 10, 2012 are broken. Luckily, they can be changed from archaeology.org to archive.archaeology.org. Example is this to that for Archaeology of Israel. This also works for their magazine archives like this to that for Rome. I think it'd be easier to check the entire site and changing the urls to archive.archaeology.org to see if itll fix the ones that are broken. Some of these already have archived copies. ~1270. Thanks! MrLinkinPark333 (talk) 21:52, 10 May 2026 (UTC)
Enwiki
- Checked 1,882 pages and edited 896 pages. Moved 984 links to a new URL: 340 normal redirects, 639 ruled mapped redirects, 5 ghost mapped redirects, Resolved 170 soft-404s. Added 3
{{dead link}}. Switched 42|url-status=deadto live. Switched 2|url-status=liveto dead. Added 44 archive URLs (44 Wayback).
IABot DB
- Checked about 1,000 URLs and updated about 250
Done -- GreenC 00:45, 16 May 2026 (UTC)
shtetlinks.jewishgen.org
[edit]They redirect to kehilalinks.jewishgen.org. If there is www. in the subdomain then it doesn't automatically redirect, and you must remove the www. such as this into this which now redirects here. Thanks!
One thing I noticed about the broader jewishgen.org domain is that there was at least one cite to a mere Wikipedia mirror (namely this which was on Olivia Newton-John). I don't think anything particularly needs to be done about that fact, it just felt remiss not to mention it.
Only around 160 on the shetlinks sub-domain. There are also around 1700 articles on the broader domain. GrapesRock (talk) 11:41, 11 May 2026 (UTC)
Enwiki
- Checked 162 pages and edited 152 pages. Moved 178 links to a new URL: 178 ruled mapped redirects, Removed 1
{{dead link}}. Switched 13|url-status=deadto live. Added 4 archive URLs (4 Wayback).
IABot DB
- Checked and updated 124 URLs
Done -- GreenC 02:03, 16 May 2026 (UTC)
thisislondon.co.uk
[edit]The Standard (London newspaper) has rebranded from thisislondon.uk to standard.co.uk.
DiophantineEquation (talk) 16:50, 12 May 2026 (UTC)
- This was fairly difficult because there are multiple layers over time where they made changes. I was able to make live a significant portion, at the Standard site. Those that remain there is no clear way to find the new link, or the content no longer exists. -- GreenC 19:42, 16 May 2026 (UTC)
Enwiki
- Pass 1: Checked 1,922 pages and edited 523 pages. Moved 180 links to a new URL: 136 ruled mapped redirects, 44 ruled inferred mapped redirects, Resolved 1,950 soft-404s. Removed 1
{{dead link}}. Added 60{{dead link}}. Switched 34|url-status=deadto live. Switched 25|url-status=liveto dead. Added 293 archive URLs (293 Wayback).
- Need two passes to search Wayback CDX records for slug segments eg. tackle-the-nhs-morphine-crisis which could be available in one domain (as a 301) or the other (as a 200).
- Pass 2: Checked 1,922 pages and edited 712 pages. Moved 771 links to a new URL: 1 normal redirects, 770 ruled inferred mapped redirects, Resolved 1,786 soft-404s. Removed 3
{{dead link}}. Added 3{{dead link}}. Switched 595|url-status=deadto live. Switched 1|url-status=liveto dead. Added 4 archive URLs (4 Wayback).
- Pass 3: Checked 1,922 pages and edited 400 pages. Moved 434 links to a new URL: 434 ruled inferred mapped redirects, Resolved 1,080 soft-404s. Removed 5
{{dead link}}. Switched 309|url-status=deadto live. Added 3 archive URLs (3 Wayback).- Pass 3 w/ additional conversion rules.
IABot DB
- Updated about 1,400 URLs
Done -- GreenC 20:55, 16 May 2026 (UTC)
pe.com
[edit]If you change "pe.com" to "pressenterprise.com" then for some they successfully 301, some they are at the url already, and some don't seem to exist.
- For already being at the url: this to this is just the new url.
- For not existing this.
- For 301-ing: this turned into this redirects here.
Around 1250 articles. GrapesRock (talk) 11:21, 14 May 2026 (UTC)
Enwiki
- Checked 1,157 pages and edited 853 pages. Moved 943 links to a new URL: 943 ruled mapped redirects, Resolved 3 soft-404s. Removed 18
{{dead link}}. Added 88{{dead link}}. Switched 212|url-status=deadto live. Switched 18|url-status=liveto dead. Added 184 archive URLs (184 Wayback).
IABot DB
- Checked about 2,000 URLs
Done -- GreenC 00:27, 17 May 2026 (UTC)
marvel.com
[edit]I've come across a lot of marvel.com pages which are dead such as this. Not sure there is a new URL for these. Gonnym (talk) 11:33, 14 May 2026 (UTC)
Enwiki
- Checked 3,222 pages and edited 1,502 pages. Moved 465 links to a new URL: 64 normal redirects, 272 ruled mapped redirects, 129 ghost mapped redirects, Resolved 731 soft-404s. Added 98
{{dead link}}. Switched 19|url-status=deadto live. Switched 626|url-status=liveto dead. Added 1,263 archive URLs (1,263 Wayback).
IABot DB
- Checked 12,615 and updated 6,553
Done -- GreenC 16:35, 17 May 2026 (UTC)
abc.com
[edit]I've come across a lot of abc.com pages which are dead such as this. Not sure there is a new URL for these. Gonnym (talk) 11:33, 14 May 2026 (UTC)
- For better or worse, ABC changed it's URL structure to 36-char UUID: this is now this. ABC is part of a large holding company (Disney) with many properties, they do this for a number of good technical reasons, but it creates problems for us at Wikipedia. The URL is no longer humanly associated with the content. They can change the content of the page without needing to change the URL itself ie. renaming the show, date of publication or where the page routes. This makes fixing content drift and link rot more difficult since the URL is not decipherable. For example, I can no longer search the WaybackMachine index for "good-morning-america" to find a new URL location after it moved. They may be doing the same thing at Disney+, Hulu, ESPN+, FX. -- GreenC 17:40, 17 May 2026 (UTC)
Enwiki
- Checked 600 pages and edited 412 pages. Moved 387 links to a new URL: 328 normal redirects, 58 ruled mapped redirects, 1 ghost mapped redirects, Resolved 28 soft-404s. Removed 1
{{dead link}}. Added 10{{dead link}}. Switched 19|url-status=deadto live. Switched 11|url-status=liveto dead. Added 87 archive URLs (87 Wayback).
IABot DB
- Checked 741 URLs and updated 314
Done -- GreenC 21:04, 17 May 2026 (UTC)
insidetv.ew.com
[edit]Pages using insidetv.ew.com such as this are completely dead. Gonnym (talk) 11:34, 14 May 2026 (UTC)
Enwiki
- Checked 1,909 pages and edited 1,524 pages. Added 15
{{dead link}}. Switched 534|url-status=liveto dead. Added 1,637 archive URLs (1,637 Wayback).
IABot DB
- Checked and updated 3,335 URLs
Done -- GreenC 02:08, 18 May 2026 (UTC)
rawstory.com
[edit]If you change "rawstory.com/rs" to just "rawstory.com/" then from what I've checked it successfully redirects. For instance this to this (it also gets rid of the day, but I haven't found a case where it doesn't successfully redirect from Y/M/D to Y/M so it's probably unnecessary to code that in?). Thanks!
Around 200 articles on the /rs path and around 700 articles on rawstory.com total. GrapesRock (talk) 11:36, 14 May 2026 (UTC)
Enwiki
- Checked 675 pages and edited 339 pages. Moved 276 links to a new URL: 11 normal redirects, 263 ruled mapped redirects, 2 ghost mapped redirects, Resolved 4 soft-404s. Removed 1
{{dead link}}. Added 7{{dead link}}. Switched 101|url-status=deadto live. Switched 19|url-status=liveto dead. Added 74 archive URLs (74 Wayback).
IABot DB
- Checked 1,380 URLs and updated about 800
Done -- GreenC 15:42, 18 May 2026 (UTC)
fivethirtyeight.com
[edit]A former senior editor at FiveThirtyEight just reported that "ABC News has now taken all FiveThirtyEight articles completely offline. They now redirect to abcnews dot com.politics". Here's a few examples of this redirect in Wiki sources:
https://fivethirtyeight.com/features/is-your-dd-character-rare/https://fivethirtyeight.com/features/world-cup-stoppage-time-is-wildly-inaccuratehttps://projects.fivethirtyeight.com/coronavirus-polls/
Thanks! Sariel Xilo (talk) 17:18, 15 May 2026 (UTC)
Note: ABC previously did a partial takedown of articles & those links were fixed (Wikipedia:Link rot/URL change requests/Archives/2025/March#fivethirtyeight.com) Sariel Xilo (talk) 17:22, 15 May 2026 (UTC)
This is what AI had to say about the shutdown:
Reason for shutdown
|
|---|
|
The sudden removal of FiveThirtyEight's historical articles today is a massive, frustrating example of corporate link rot. Former staffers, including Nathaniel Rakich, flagged the purge this morning, calling it a "needless erasure of thousands of pages of knowledge," and founder Nate Silver has publicly criticized ABC for wiping out a decade's worth of data-driven journalism. Here is the breakdown of why ABC News pulled the plug and how we got here: 1. The Final Stage of the 2025 Shutdown This move is the final nail in the coffin following the events of March 2025. During a wave of broader cost-cutting measures and layoffs at Disney, ABC News effectively shut down the FiveThirtyEight brand. Active publishing ceased, and while the archives remained accessible for the last year, the brand itself was defunct. 2. Infrastructure and CMS Consolidation From a corporate engineering perspective, keeping historical archives online isn't free. FiveThirtyEight relied heavily on custom interactive graphics, complex data visualizations, and likely a heavily modified CMS. Maintaining that legacy infrastructure—keeping it patched, secure, and compatible with modern web standards—costs money. Disney decided that absorbing those maintenance costs for static pages that no longer generate new ad revenue was unjustifiable. 3. SEO Traffic Funneling Instead of maintaining the archives or properly migrating them to a static, read-only format, ABC implemented a blanket URL redirect. By pointing every old FiveThirtyEight link to abcnews.com/politics, ABC News retains the residual SEO value and inbound link juice. It abruptly forces anyone seeking historical polling analysis into their active, monetized news ecosystem. What Survives the Purge This event perfectly highlights the inherent fragility of relying on centralized, corporate-owned entities for digital preservation. The moment a collection is no longer profitable, it vanishes behind a 301 redirect. While the narrative articles and specific analyses are now dependent on whatever snapshots were successfully captured and rendered by web archiving toolchains before the redirect went live, not all the data was lost. The New York Times managed to adopt and salvage some of FiveThirtyEight's publicly available polling databases (like the presidential approval trackers) shortly after the initial shutdown in early 2025. However, the vast majority of the context, the methodology explanations, and the historical election post-mortems are now gone from the live web. |
-- GreenC 22:45, 15 May 2026 (UTC)
- These articles were under two domains (links to IA listings): fivethirtyeight.com from the blog's start in 2008 until September 2023, and then abcnews.go.com/538 from then until it shut down. There are also some third-level domains like projects.fivethirtyeight.com and data.fivethirtyeight.com, and probably others. Antony–22 (talk⁄contribs) 04:50, 17 May 2026 (UTC)
Enwiki
- Checked 1,928 pages and edited 1,017 pages. Added 14
{{dead link}}. Switched 341|url-status=liveto dead. Added 1,133 archive URLs (1,133 Wayback).
IABot DB
- Updated 2,700 URLs
abcnews.go.com/538
[edit]- Enwiki
- Checked 82 pages and edited 80 pages. Added 2
{{dead link}}. Switched 23|url-status=liveto dead. Added 74 archive URLs (74 Wayback).
- Checked 82 pages and edited 80 pages. Added 2
- IABot DB
- Updated 58 URLs
Done -- GreenC 00:40, 19 May 2026 (UTC)
londongardensonline.org.uk
[edit]This was a charity directory that changed domain. The domain is now owned by a company. URLs in the format http://www.londongardensonline.org.uk/gardens-online-record.asp?ID=THM001 should be changed to https://londongardenstrust.org/conservation/inventory/site-record/?ID=THM001 as here. MRSC (talk) 18:22, 17 May 2026 (UTC)
Enwiki
- Checked 388 pages and edited 385 pages. Moved 443 links to a new URL: 443 ruled mapped redirects, Removed 15
{{dead link}}. Switched 127|url-status=deadto live. Switched 1|url-status=liveto dead. Added 2 archive URLs (2 Wayback).
IABot DB
- Checked and updated 385 URLs
Done -- GreenC 18:35, 19 May 2026 (UTC)
bbm.ca usurped
[edit]Links to pages like https://www.bbm.ca/_documents/top_30_tv_programs_english/2013-14/2013-14_09_23_TV_ME_NationalTop30.pdf should be marked as usurped as the live bbm.ca is a completely different website (see archive https://web.archive.org/web/20131005081516/https://www.bbm.ca/_documents/top_30_tv_programs_english/2013-14/2013-14_09_23_TV_ME_NationalTop30.pdf). BBM.ca changed to Numeris. Gonnym (talk) 06:26, 18 May 2026 (UTC)
- ω Awaiting next WP:JUDI batch -- GreenC 02:37, 19 May 2026 (UTC)
www.newsday.com
[edit]I found that http://www.newsday.com/entertainment/tv/arrow-is-off-to-a-fast-start-1.4089299 is now located at https://www.newsday.com/entertainment/tv/arrow-is-off-to-a-fast-start-d33660. Gonnym (talk) 07:07, 18 May 2026 (UTC)
- 5,985 pages + download Wayback CDX records (apx 12 hrs)
- This will have two passes. The first will add archive URLs to dead links. The second will attempt to move dead links to live where possible. This process has elements of probabilistic (vs. deterministic). So the bot errs on the side of false negatives (ie. skips conversions) to avoid making false positives (converting to the wrong URL). The former has the backstop of an archive URL, the later has no backstop, it is a wrong URL. End result: it can't make as many conversions as otherwise would be possible. -- GreenC 03:03, 27 May 2026 (UTC)
Enwiki
- Pass 1: Checked 5,998 pages and edited 4,021 pages. Moved 27 links to a new URL: 7 normal redirects, 20 ruled mapped redirects, Resolved 1,052 soft-404s. Removed 1
{{dead link}}. Added 408{{dead link}}. Switched 2|url-status=deadto live. Switched 926|url-status=liveto dead. Added 3,941 archive URLs (3,941 Wayback).
- Pass 2: Checked 5,998 pages and edited 2,217 pages. Moved 2,878 links to a new URL: 1 normal redirects, 3 ruled mapped redirects, 2,874 ruled inferred mapped redirects, Resolved 74 soft-404s. Removed 17
{{dead link}}. Added 387{{dead link}}. Switched 2,657|url-status=deadto live. Switched 8|url-status=liveto dead.
IABot DB
- Updated 8,112 URLs
Done -- GreenC 19:04, 27 May 2026 (UTC)
www.boston.com
[edit]https://www.boston.com/ae/tv/2012/10/09/arrow-falls-short-bullseye/r1BsuzTM8jDY923XdlOJFI/story.html is now at https://www.boston.com/culture/tv/2012/10/09/arrow-falls-short-of-a-bullseye/ Gonnym (talk) 07:08, 18 May 2026 (UTC)
- 14,929 pages + download Wayback CDX records (12+ hrs)
Enwiki
- Checked 14,736 pages and edited 12,098 pages. Moved 10,396 links to a new URL: 316 normal redirects, 2,659 ruled mapped redirects, 7,421 ruled inferred mapped redirects, Resolved 1,863 soft-404s. Removed 10
{{dead link}}. Added 450{{dead link}}. Switched 1,167|url-status=deadto live. Switched 808|url-status=liveto dead. Added 4,158 archive URLs (4,158 Wayback).
IABot DB
- Checked 28,000 URLs and updated 20,000
Done - this was a mountain of URLs, most needed fixing, and 2/3rds was able to move to a new URL via CDX inference mapping. -- GreenC 16:02, 29 May 2026 (UTC)
Per the TFD this template will need to be "unwound" with archive URLs added. Primefac (talk) 12:14, 18 May 2026 (UTC)
- Checked 3,310 pages and edited 3,310 pages. Converted 3,339 templates. Added 3,125 archive URLs (3,125 Wayback).
Done -- GreenC 01:52, 30 May 2026 (UTC)
- Awesome, thanks. Primefac (talk) 10:21, 30 May 2026 (UTC)
mahasz.hu
[edit]These Hungarian charts no longer work but most are at slagerlistak.hu. Most can be converted over to slagerlistak.hu/chart-name/year/week or slagerlistak.hu/archivum/eves-osszesitett-slagerlistak/chart-name/year with the following:
Weekly charts (year/week)
- Single Top 40: here to there for I Want to Know What Love Is
- Radio Top 40: here to there for You and I (Lady Gaga song)
- Albums Top 40: here to there for Wrecking Ball (Bruce Springsteen album)
Year-end
- Radio: here to here for Set Fire to the Rain
- Album: here to there for Minutes to Midnight (Linkin Park album)
- Editor's choice: here to there for Teenage Dream (Katy Perry song)
- Best selling: here to there for In the Zone
Broken:
I didn't include ~60 of them because they need manual adjustments. If I can't replace them, I'll make a new request. 270 Thanks! MrLinkinPark333 (talk) 20:09, 18 May 2026 (UTC)
Enwiki
- Checked 390 pages and edited 334 pages. Moved 380 links to a new URL: 380 ruled mapped redirects, Removed 7
{{dead link}}. Switched 60|url-status=deadto live. Added 3 archive URLs (3 Wayback).
IABot DB
- Checked and updated 2200 URLs
Done -- GreenC 03:49, 30 May 2026 (UTC)
- Which three didn't work? MrLinkinPark333 (talk) 03:56, 30 May 2026 (UTC)
From the logs:
This Boy's Fire----https://web.archive.org/web/20120219203200/http://www.mahasz.hu/ ---- fixbadstatus1.1 (old logbadstatus1) No Quarter: Jimmy Page and Robert Plant Unledded----https://web.archive.org/web/20100826043127/http://www.mahasz.hu/ ---- fixbadstatus1.1 (old logbadstatus1) Édes méreg----https://web.archive.org/web/20051018090219/http://www.mahasz.hu/m/hu/arany_keres.php?EV=2005 ---- fixbadstatus3.3 (old barelink-modify)
-- GreenC 02:38, 31 May 2026 (UTC)
archive.org/details
[edit]This is a bot process to remove or move links at archive.org/details/ID when the ID is no longer available.
For example this:
- Seligman, M. E. P. (2011). Flourish: A Visionary New Understanding of Happiness and Well-Being. New York: Free Press. ISBN 978-1-4391-9076-0.
Becomes this:
- Seligman, M. E. P. (2011). Flourish: A Visionary New Understanding of Happiness and Well-Being. New York: Free Press. ISBN 978-1-4391-9076-0.
Because this:
No longer works.
- Notes
- Estimate about 5% of archive.org/details need removal or modification.
- This is a 1-time batch job.
- Books are unavailable for many reasons technical and policy.
- IDs sometimes move - same book, new ID.
- Google Books has the same but nobody maintains them systematically, the rate is higher than 5%, and the quantity of GB exceeds IA by at least 2x
-- GreenC 00:30, 19 May 2026 (UTC)
haaretz.com
[edit]Many links seem to redirect to other links and some of them may have been dead at some point. For instance this from 2011 in Israel was marked dead way back in September 2011. It now 301's here. Thanks!
Just over 12000 articles. GrapesRock (talk) 11:18, 19 May 2026 (UTC)
- Perhaps the example was marked dead because the live redirect is a paywall, while the dead has a working archive. -- GreenC 00:39, 28 May 2026 (UTC)
- I did a dry run of the entire set, and almost all the links are of the type this ie. ending in 1.xxxxx .. these links were at one time not behind a paywall, and so are currently available at archive.org in full .. but most do not have archive.org links on Wikipedia. So if I moved the primary link to the new redirect here, we would no longer know what the archive.org link was for the old 1.xxx - the page becomes locked behind a paywall at the new link. I will instead treat all the 1.xxx as dead, so archive URLs are added, and the content remains visible. In this case upgrading to a live link is a step backwards. -- GreenC 17:45, 30 May 2026 (UTC)
- A common exception: when the url contains ".premium" (example), meaning the 1.xxx url was always paywalled and the archive.org version will also be paywalled. In these cases the url will be moved to the new live URL here, and no archive added. -- GreenC 18:44, 30 May 2026 (UTC)
- This solution had about a 2:1 ratio: for every 2 links that have a working archive, 1 link was moved to a live (paywalled) URL. -- GreenC 17:52, 31 May 2026 (UTC)
- A common exception: when the url contains ".premium" (example), meaning the 1.xxx url was always paywalled and the archive.org version will also be paywalled. In these cases the url will be moved to the new live URL here, and no archive added. -- GreenC 18:44, 30 May 2026 (UTC)
- I did a dry run of the entire set, and almost all the links are of the type this ie. ending in 1.xxxxx .. these links were at one time not behind a paywall, and so are currently available at archive.org in full .. but most do not have archive.org links on Wikipedia. So if I moved the primary link to the new redirect here, we would no longer know what the archive.org link was for the old 1.xxx - the page becomes locked behind a paywall at the new link. I will instead treat all the 1.xxx as dead, so archive URLs are added, and the content remains visible. In this case upgrading to a live link is a step backwards. -- GreenC 17:45, 30 May 2026 (UTC)
Enwiki
- Batch 00001-00100: Checked 100 pages and edited 67 pages. Moved 4 links to a new URL: 4 normal redirects, Added 1
{{dead link}}. Switched 18|url-status=liveto dead. Added 110 archive URLs (110 Wayback).
- Batch 00101-00200: Checked 100 pages and edited 82 pages. Moved 48 links to a new URL: 48 normal redirects, Added 1
{{dead link}}. Switched 6|url-status=deadto live. Switched 21|url-status=liveto dead. Added 81 archive URLs (81 Wayback).
- Batch 00201-01000: Checked 800 pages and edited 609 pages. Moved 355 links to a new URL: 347 normal redirects, 2 ruled mapped redirects, 6 ghost mapped redirects, Resolved 3 soft-404s. Added 22
{{dead link}}. Switched 4|url-status=deadto live. Switched 151|url-status=liveto dead. Added 683 archive URLs (683 Wayback).
- Batch 01001-11968: Checked 10,968 pages and edited 8,393 pages. Moved 5,219 links to a new URL: 5,115 normal redirects, 10 ruled mapped redirects, 94 ghost mapped redirects, Resolved 35 soft-404s. Removed 10
{{dead link}}. Added 149{{dead link}}. Switched 123|url-status=deadto live. Switched 1,597|url-status=liveto dead. Added 8,745 archive URLs (8,745 Wayback).
IABot DB
- Checked 36,119 URLs and updated 23,100
Done -- GreenC 01:21, 2 June 2026 (UTC)
- Dear GreenC, I am confused. I have recently encountered a few of these bot edits (citing the present talk page discussion) and in all of them, the redirect on the haaretz website is live and the correct article. (and indeed: "paywalled" now: you need to make a free account to continue reading). I think something must have gone wrong. I have reverted one or more of these automated edits. If I should not have, please inform me. Slomo666 (talk) 14:01, 4 June 2026 (UTC)
- Despite the alleged exception for the premium articles, I have seen at least one .premium article that the bot did set to dead. This saga baffles me, and I just hope that articles accidentally being set to dead is the worst of it. Slomo666 (talk) 14:25, 4 June 2026 (UTC)
- Right it was intentional. Discussed above. Freewall, Paywall, etc.. they all have the same effect: The Wayback Machine can't save them. Had I switched it over to the live walled link, what happens in a few years when the link dies (it will): there will be no way to view the content again. There is no archive URL, because the Wayback Machine can't save walled pages. The way it was done, the old links are still available at the old archive URLs, they were saved at Wayback before Haaretz put up walls. We are lucky the Wayback captures exist. When the live walled links die in a few years, there will be nothing to save them from becoming permanent 404. -- GreenC 08:07, 5 June 2026 (UTC)
- No, what I am saying is that your bot set a bunch of already archived citation URL's from "live" to "dead" after the respective archive URL's (which were not inaccessible!) were already added to the citation template.
- I understand the bot cannot retroactively archive things that are no longer live, (or in this case: behind a registration-wall) but that is not what I am saying it should have done or what it did.
- If you want examples, look at this: I reverted your bot's edit, because all of those URL's, while they redirect to the new haaretz website, actually do link to a live version of the same articles, and there was thus no need to reset the "url-status=live" to "url-status=dead".
- I think this is problematic, because we generally do not want to set url-status to dead unless it actually is. I think the bot should not be doing what I described, and instead, only set, for those refs where the archive-link is already filled, but the new url is behind the wall, the 'url-access' to something. (in these examples, I set it to 'registration', because that was what it was, but I can imagine it may be difficult for the bot to tell the difference between 'limited', ' registration' and 'subscription')
- Slomo666 (talk) 15:51, 5 June 2026 (UTC)
- I had not really thought through all the permutations at play this was a complicated site with a lot of variation and edge cases. But I am also not too worried about it because I know from experience redirects are fragile and the first to go, those redirects to the live site will disappear sooner than later. Still, in the example you provided: would you prefer to go here or here? The first takes 99% of readers to a 0% readable page (signup required). The second takes 100% of readers to a 100% readable page (no signup required). Setting to dead will be problematic to haaretz owners, and those who have accounts on the website. It will be pragmatic for most everyone else. Had I done it the other way, it would have been problematic for the majority as they would be redirected to walled off content. You could say none of that matters only that dead is dead and live is live, but not sure I agree that's always in the best interest, and time will resolve it anyway. -- GreenC 06:57, 6 June 2026 (UTC)
- Personally, I prefer to go to a source that is available, of course. However, I assumed (and maybe I am wrong about this) that marking URL's that are not dead as dead was against one or more of our guidelines. I recall there being some resistance to expansion of the 'url-access-level' parameter, (I might be confusing it, and maybe it was a discussion on adding a new option to the url-status param) based on the idea that it is not up to wikipedia to censure sources for not being accessible freely. I do think we should have the ability to work into the template that a source link now redirects, but I think the current options ('deviated' and 'unfit') are really only for when it redirects to a harmful URL or a website that does not support the same content anymore. Is there a venue where I could ask about how this policy works? Teahouse maybe? Slomo666 (talk) 12:24, 8 June 2026 (UTC)
- I had not really thought through all the permutations at play this was a complicated site with a lot of variation and edge cases. But I am also not too worried about it because I know from experience redirects are fragile and the first to go, those redirects to the live site will disappear sooner than later. Still, in the example you provided: would you prefer to go here or here? The first takes 99% of readers to a 0% readable page (signup required). The second takes 100% of readers to a 100% readable page (no signup required). Setting to dead will be problematic to haaretz owners, and those who have accounts on the website. It will be pragmatic for most everyone else. Had I done it the other way, it would have been problematic for the majority as they would be redirected to walled off content. You could say none of that matters only that dead is dead and live is live, but not sure I agree that's always in the best interest, and time will resolve it anyway. -- GreenC 06:57, 6 June 2026 (UTC)
- Right it was intentional. Discussed above. Freewall, Paywall, etc.. they all have the same effect: The Wayback Machine can't save them. Had I switched it over to the live walled link, what happens in a few years when the link dies (it will): there will be no way to view the content again. There is no archive URL, because the Wayback Machine can't save walled pages. The way it was done, the old links are still available at the old archive URLs, they were saved at Wayback before Haaretz put up walls. We are lucky the Wayback captures exist. When the live walled links die in a few years, there will be nothing to save them from becoming permanent 404. -- GreenC 08:07, 5 June 2026 (UTC)
- Despite the alleged exception for the premium articles, I have seen at least one .premium article that the bot did set to dead. This saga baffles me, and I just hope that articles accidentally being set to dead is the worst of it. Slomo666 (talk) 14:25, 4 June 2026 (UTC)
archive.*.co.uk
[edit]A bunch of British newspaper archives changed their url schemes and now none of them link directly to the source and instead redirect to a page with all the articles from the relevant day. Since they all have small numbers of pages, I've been manually replacing marked dead links with alternative links. However, the corresponding IABot entries should be updated to mark these domains as dead I think? Thanks!
These are the ones that I've found so far; there's probably others, but these are the ones I could find with "insource:"co.uk" insource:"archive" insource:/archive\.[a-z0-9-]+\.co\.uk\/[0-9]{4}/"
List of domains
|
|---|
|
GrapesRock (talk) 11:22, 20 May 2026 (UTC)
- GrapesRock, unlike the WP:JUDI process, the WaybackMedic process can not do batches of domains. Each domain has to be individually configured and processed. (There is bespoke work to catch soft-404s and redirect rules by monitoring logs, it's impossible to automate certain things safely.) I checked a couple and some redirect properly to the original article. I'm also not sure marking them dead in IABot is right. So, I'm not sure how to approach this. It's a bit much in quantity of domains, and low link count for each domain. -- GreenC 04:45, 28 May 2026 (UTC)
uk-sport-web.prod.oceanusorigin.com
[edit]Replace that unwieldy domain with skysports.com. For example this to this.
Around 70 pages. GrapesRock (talk) 17:14, 21 May 2026 (UTC)
Enwiki
- Checked 71 pages and edited 72 pages. Moved 99 links to a new URL: 99 ruled mapped redirects, Removed 48
{{dead link}}.
IABot DB
- Checked and updated about 300 URLs.
Done -- GreenC 04:56, 9 July 2026 (UTC)
news.yahoo.co.jp/articles
[edit]sigh… this widely used news aggregator website loves to kills its links in just few days. (1
). Basically any source older than 7 days can be dead, it will almost certainly dead by one month.
Note I proposed of a bot that automatically archive this sorts of website when added at the WP:IDEALAB. See Wikipedia:Village pump (idea lab)#Bot for automatic source rescue Warm Regards, Miminity (Talk?) (me contribs) 06:16, 23 May 2026 (UTC)
Enwiki
- Checked 1,651 pages and edited 1,007 pages. Added 260
{{dead link}}. Switched 298|url-status=liveto dead. Added 1,125 archive URLs (1,125 Wayback).
IABot DB
- Updated 5,728 links
Done -- GreenC 06:11, 11 July 2026 (UTC)
post-gazette.com
[edit]I requested an archive check on some of these links a few months ago. Recently, found a broken magazine article by them for Midnight Ramble (film). I think all articles by this newspaper should be checked now. ~8700. Thank you! MrLinkinPark333 (talk) 19:08, 23 May 2026 (UTC)
- This domain certainly needed help. It works, only moves slow due to extra network steps to verify the links. -- GreenC 21:16, 12 July 2026 (UTC)
Enwiki
- Checked 8,713 pages and edited 5,032 pages. Moved 3,178 links to a new URL: 22 normal redirects, 2,587 ruled mapped redirects, 569 ghost mapped redirects, Resolved 906 soft-404s. Removed 3
{{dead link}}. Added 271{{dead link}}. Switched 117|url-status=deadto live. Switched 566|url-status=liveto dead. Added 3,802 archive URLs (3,802 Wayback).
IABot DB
- Checked 15,125 and updated 6,837 URLs
Done -- GreenC 19:19, 13 July 2026 (UTC)
mixonline.com
[edit]This broken link is now here for Flow (Terence Blanchard album). Unfortunately, it does not redirect. I would like to request a full check. 1290. Thank you! MrLinkinPark333 (talk) 19:13, 23 May 2026 (UTC)
Enwiki
- Checked 1,305 pages and edited 270 pages. Moved 244 links to a new URL: 18 normal redirects, 220 ruled mapped redirects, 6 ghost mapped redirects, Resolved 7 soft-404s. Added 3
{{dead link}}. Switched 21|url-status=deadto live. Switched 1|url-status=liveto dead. Added 71 archive URLs (71 Wayback).
IABot DB
- Checked 1,710 and updated 780 URLs
Done -- GreenC 14:57, 14 July 2026 (UTC)
catholicnewsagency.com
[edit]Links to individual Catholic News Agency articles now all redirect to the home page https://www.ewtnnews.com/ (example here), following EWTN's decision to redirect the Catholic News Agency URLs to the new EWTN News website.
Could you please fix this? Veverve (talk) 09:33, 25 May 2026 (UTC)
- This could have been a temporary under construction as the example now redirects correctly. But I'll run the domain through and see what they might have missed, sites often drop things during a move. -- GreenC 15:12, 14 July 2026 (UTC)
Enwiki
- Checked 2,387 pages and edited 2,355 pages. Moved 3,300 links to a new URL: 3,211 normal redirects, 73 ruled mapped redirects, 16 ghost mapped redirects, Resolved 268 soft-404s. Removed 1
{{dead link}}. Added 10{{dead link}}. Switched 91|url-status=deadto live. Switched 39|url-status=liveto dead. Added 184 archive URLs (184 Wayback).
IABot DB
- Checked 5.048 and updated 2,221 URLs
Done -- GreenC 20:10, 14 July 2026 (UTC)
justice.gov/usao-dc/
[edit]Department of Justice Press Releases deleted by the Trump Administration related to January 6 riots.[4] -- GreenC 20:47, 25 May 2026 (UTC)
Not done The links on enwiki appear to still work (unrelated to 1/6) -- GreenC 23:37, 14 July 2026 (UTC)
wtvq.com
[edit]Website of WTVQ-DT, a local TV station in Kentucky. Domain is parked and the content on it is not loading, likely because the station has been sold. About 160 pages. Sammi Brie (she/her · t · c) 02:32, 26 May 2026 (UTC)
- If I pick a few at random they seem to work. Perhaps it was temporary? Neils51 (talk) 21:09, 27 May 2026 (UTC)
Not done - spot checked links still work. -- GreenC 23:39, 14 July 2026 (UTC)
philrobson.net
[edit]Website of Phil_Robson, a Jazz musician, loads a Korean blog. While I don't read Korean, it seems to have nothing to do with Phil, Jazz, or even Music. He has another website instead: https://www.philrobsonmusic.com/ — Preceding unsigned comment added by Bevande (talk • contribs) 06:45, 30 May 2026 (UTC)
Not done - links still work -- GreenC 23:40, 14 July 2026 (UTC)
observationdeck.io9.com
[edit]http://observationdeck.io9.com/ is completely dead (see link). Gonnym (talk) 16:49, 1 June 2026 (UTC)
Enwiki
- Checked 25 pages and edited 10 pages. Added 1
{{dead link}}. Switched 7|url-status=liveto dead.
IABot DB
- Checked and updated 31 links
Done -- GreenC 23:55, 14 July 2026 (UTC)
tv.com
[edit]http://www.tv.com is completely dead. (see link). Gonnym (talk) 16:59, 1 June 2026 (UTC)
- TV.com went offline in 2021. First time reported here. Appears to have a lot of dead links. -- GreenC 00:19, 15 July 2026 (UTC)
- The archives also aren't loading for some reason (see link). Gonnym (talk) 17:07, 1 June 2026 (UTC)
- That is a technical error in Wayback replay. It's URLs in the /shows path eg. [5] - I will add archives in the hopes archive.org finds and fixes the error with time, better than a dead link. -- GreenC 00:19, 15 July 2026 (UTC)
Enwiki
- Checked 4,676 pages and edited 2,226 pages. Added 709
{{dead link}}. Switched 494|url-status=liveto dead. Added 1,967 archive URLs (1,967 Wayback).
IABot DB
- Checked and updated 26,297 URLs
Done -- GreenC 22:13, 15 July 2026 (UTC)
brokenpencil.com
[edit]brokenpencil.com- 48
Broken Pencil shutdown in 2024; all of the website's links prompt a sign in that if declined goes to a 401 Authorization Required. Haven't done an in-depth check but ProQuest seems to have some of the magazine saved if the Wayback doesn't (or if the Wayback has saved a subscription needed page). For example, https://brokenpencil.com/article/journeys-on-paper-jeeyon-shim-and-the-personal-resonance-of-tabletop-rpgs/ can be found at ProQuest 3108820147. Sariel Xilo (talk) 17:44, 2 June 2026 (UTC)
- At this writing (over 1½ months on), domain-parked at BigScoots.com. --Slgrandson (How's my egg-throwing coleslaw?) 12:10, 22 July 2026 (UTC)
Enwiki
- Checked 48 pages and edited 38 pages. Switched 4
|url-status=liveto dead. Added 37 archive URLs (37 Wayback).
IABot DB
- Checked and updated 66 URLs
Done -- GreenC 22:28, 15 July 2026 (UTC)
monumentaustralia.org.au
[edit]Has been usurped by a gaming website.
It appears that http://monumentaustralia.org still works, and the pages I tested follow the same url format, so having the bot remove the .au should fix things. In solidarity, nil nz 04:35, 4 June 2026 (UTC)
Enwiki
- Checked 1,097 pages and edited 1,088 pages. Moved 1,421 links to a new URL: 1,421 ruled mapped redirects, Removed 2
{{dead link}}. Switched 18|url-status=deadto live. Switched 1|url-status=liveto dead. Added 1 archive URLs (1 Wayback).
IABot DB
- Checked and updated 1,378 links
Done -- GreenC 03:00, 16 July 2026 (UTC)
- Appreciated! In solidarity, nil nz 03:14, 16 July 2026 (UTC)
thenortheasttoday.com
[edit]Found this soft 404 to the main page at Lai Haraoba. Looking around their website, only some of the old articles are there. I think a full check is needed. ~240. Thanks! MrLinkinPark333 (talk) 23:21, 5 June 2026 (UTC)
Enwiki
- Checked 246 pages and edited 101 pages. Moved 65 links to a new URL: 6 normal redirects, 51 ruled mapped redirects, 8 ghost mapped redirects, Resolved 48 soft-404s. Removed 5
{{dead link}}. Added 37{{dead link}}. Switched 25|url-status=deadto live. Added 6 archive URLs (6 Wayback).
IABot DB
- Checked 313 links and updated 79
Done -- GreenC 03:40, 16 July 2026 (UTC)
sec.state.ma.us/ele/
[edit]State of Massachusetts Election pages
- https://www.sec.state.ma.us/ele/eledist/reps11idx.htm -> https://www.sec.state.ma.us/divisions/elections/voting-information/district/2022-representative.htm (157 pages)
- https://www.sec.state.ma.us/ele/eledist/sen11idx.htm -> https://www.sec.state.ma.us/divisions/elections/voting-information/district/2022-senatorial.htm (34 pages)
- Any other URL on this domain that redirects to https://www.sec.state.ma.us/divisions/elections/elections-and-voting.htm is effectively soft 404 and should be treated as if it were dead. I spot checked a few other patterns that used this domain and only the election division seems to be the one using this soft 404 pattern; other parts of the website hard 404 properly.
* Pppery * (alt) in solidarity 23:28, 10 June 2026 (UTC)
Enwiki
- Pass 1: Checked 381 pages and edited 348 pages. Moved 263 links to a new URL: 246 ruled mapped redirects, 17 ghost mapped redirects, Resolved 28 soft-404s. Added 1
{{dead link}}. Switched 9|url-status=deadto live. Switched 8|url-status=liveto dead. Added 132 archive URLs (132 Wayback). - Pass 2: Checked 381 pages and edited 110 pages. Moved 10 links to a new URL: 10 ruled mapped redirects, Resolved 131 soft-404s. Added 2
{{dead link}}. Switched 15|url-status=liveto dead. Added 117 archive URLs (117 Wayback).
IABot DB
- Pass 1:
Checked 163 and updated 91 links - Pass 2: Checked and updated 163 links
- Looks like you missed some. For example https://en.wikipedia.org/w/index.php?diff=prev&oldid=1364370763 changed a HTTP to HTTPs but the HTTPS page is still soft 404. * Pppery * in solidarity 04:36, 16 July 2026 (UTC)
- Ah thanks for checking. Looks like the bot was policy blocked (403) and it fell back to another technique that "worked", but the method failed to return the redirect URL _elections-and-voting.htm - It all looked OK because it was making edits with status 200 URLs, but they were just http->https. I should have looked more carefully. I'll try again with a different method to get past the 403 and false redirect info. -- GreenC 20:06, 16 July 2026 (UTC)
Done - better in Pass 2 with headless browser. -- GreenC 04:00, 17 July 2026 (UTC)
- Ah thanks for checking. Looks like the bot was policy blocked (403) and it fell back to another technique that "worked", but the method failed to return the redirect URL _elections-and-voting.htm - It all looked OK because it was making edits with status 200 URLs, but they were just http->https. I should have looked more carefully. I'll try again with a different method to get past the 403 and false redirect info. -- GreenC 20:06, 16 July 2026 (UTC)
tvguide.com
[edit]Links lead to page not found, see this example. For that example I've found this:
- https://www.tvguide.com/News/Hot-List-2013-Agents-SHIELD-1072981.aspx -> https://www.tvguide.com/news/Hot-List-2013-Agents-SHIELD-1072981/
but couldn't find one for
Gonnym (talk) 10:47, 15 June 2026 (UTC)
Gonnym: I think this change is a regression: live vs. archive. I could process the 16,000 pages looking for cites that end in .aspx that do not have an archive URL available and try to switch those, but I think after going through all those gates it might end up a lot of effort for little return. -- GreenC 04:14, 16 July 2026 (UTC)
- Actually more of these then I thought. Will do ie. cites that end in .aspx, and that do not have an archive URL available, switch to the live URL, otherwise keep/add the better archive version. -- GreenC 04:52, 16 July 2026 (UTC)
Enwiki
- Batch 00001-03000: Checked 3,000 pages and edited 1,573 pages. Moved 223 links to a new URL: 130 normal redirects, 93 ruled mapped redirects, Resolved 126 soft-404s. Removed 1
{{dead link}}. Added 91{{dead link}}. Switched 11|url-status=deadto live. Switched 207|url-status=liveto dead. Added 1,629 archive URLs (1,629 Wayback).
- Batch 03001-09000: Checked 6,000 pages and edited 3,064 pages. Moved 475 links to a new URL: 294 normal redirects, 181 ruled mapped redirects, Resolved 269 soft-404s. Added 136
{{dead link}}. Switched 19|url-status=deadto live. Switched 444|url-status=liveto dead. Added 2,914 archive URLs (2,914 Wayback).
- Batch 09001-17201: Checked 8,201 pages and edited 4,266 pages. Moved 621 links to a new URL: 373 normal redirects, 248 ruled mapped redirects, Resolved 261 soft-404s. Added 186
{{dead link}}. Switched 33|url-status=deadto live. Switched 594|url-status=liveto dead. Added 4,210 archive URLs (4,210 Wayback).
IABot DB
- Checked and updated 34,430 links
Done -- GreenC 20:26, 19 July 2026 (UTC)
tvline.com
[edit]- http://tvline.com/2012/11/27/shield-tv-show-casts-brett-dalton/ leads to the main page (couldn't find the page in the archives)
- http://tvline.com/2014/02/20/marvels-agents-of-shield-casting-spoilers-scandal-parenthood/ redirects to https://www.tvline.com/casting-news/marvels-agents-of-shield-casting-spoilers-scandal-parenthood-494548/
Gonnym (talk) 10:53, 15 June 2026 (UTC)
Enwiki
- Checked 7,035 pages and edited 6,118 pages. Moved 10,628 links to a new URL: 8,117 normal redirects, 2,335 ruled mapped redirects, 176 ghost mapped redirects, Resolved 2,113 soft-404s. Removed 1
{{dead link}}. Added 28{{dead link}}. Switched 192|url-status=deadto live. Switched 826|url-status=liveto dead. Added 968 archive URLs (968 Wayback).
IABot DB
- Checked 14,893 URLs and updated 5,247
Done -- GreenC 05:15, 23 July 2026 (UTC)
thejc.com
[edit]Found that this redirects here for Tom Hardy. However, this doesn't redirect to the new URL here for Manchester Reform Synagogue. I think the whole website should be checked just in case. ~3000. Thanks! MrLinkinPark333 (talk) 02:41, 23 June 2026 (UTC)
Enwiki
- Checked 3,072 pages and edited 2,373 pages. Moved 2,647 links to a new URL: 2,507 normal redirects, 74 ruled mapped redirects, 66 ghost mapped redirects, Resolved 741 soft-404s. Removed 1
{{dead link}}. Added 29{{dead link}}. Switched 82|url-status=deadto live. Switched 133|url-status=liveto dead. Added 868 archive URLs (868 Wayback).
IABot DB
- Checked 5,101 URLs and updated 1,774
Done -- GreenC 03:06, 24 July 2026 (UTC)
mojim.com
[edit]mojim.com (魔鏡歌詞網, "Magic Mirror Lyrics") was a major Taiwanese Chinese-language lyrics database. The domain went offline c. 2023–2024 and is now dead (NXDOMAIN/parked). A community recovery archive at mojimlyrics.com preserves the same content at matching paths; the site is live and confirmed operational.
~82 confirmed replaceable links across ~66 en.wiki articles. All verified live. URLs wrapped in to avoid spam-filter triggering — operators can verify each before applying.
Previously filed at Wikipedia:Bot requests on 15 June 2026 by User:~2026-35133-34 and redirected here by Qwerfjkltalk + Primefac (20 June 2026).
URL list
|
|---|
|
Group A — Song-page links (tw.mojim.com → mojimlyrics.com/song/…):
Group B — Artist-page links (tw.mojim.com → mojimlyrics.com/artist/…):
Group C — http:// links (mojim.com without subdomain):
|
User:~2026-35133-34 ~2026-36521-46 (talk) 04:30, 24 June 2026 (UTC)
Not done - was already done by the requester. -- GreenC 03:13, 24 July 2026 (UTC)
mtnmath.com
[edit]I'm not entirely sure how this works. But I've observed that mtnmath.com seems to have been linking to weird places lately, and the last time it worked well was 2021. Since then it has linked to weird spam, and more recently linked to weird domains that show Cloudflare errors.
It was present on Ordinal arithmetic until I pointed it out to someone, but I now see the same domain linked on non-Main namespaces on the English Wikipedia and still the same link on other language Wikipedias.
SgeoTC 22:34, 28 June 2026 (UTC)
Not done - links exist on two pages - adjust as needed. -- GreenC 03:16, 24 July 2026 (UTC)
screen4screen.com
[edit]The website which once also served as a database for films and had their release dates has dropped that aspect completely, leaving hundreds of error 404s. So all S4S links may be tagged as dead. In fact, their archive.today archives may also be replaced with Wayback links per WP:ARCHIVETODAY. Kailash29792 (talk) 06:34, 4 July 2026 (UTC)
- Job in two passes. Pass 1 removes archive.today links. Pass 2 adds archive.org links. -- GreenC 04:25, 24 July 2026 (UTC)
Enwiki
- Pass 1: Checked 777 pages and edited 105 pages. Removed 104 archive.today URLs.
- Pass 2: Checked 777 pages and edited 749 pages. Added 3
{{dead link}}. Switched 624|url-status=liveto dead. Added 159 archive URLs (159 Wayback).
IABot DB
- Checked and updated 1,014 URLs
Done -- GreenC 15:26, 24 July 2026 (UTC)
thestandard.com.ph
[edit]The Standard rebranded to Manila Standard back in 2016, with the website "thestandard.com.ph" moving to "manilastandard.net" sometime later. All news articles from before the domain change are preserved under the new URL. For example, this article is now located here. MarcusAbacus (talk) 10:12, 6 July 2026 (UTC)
- The new site "manilastandard.net" is protected by a bot-challenge system I have not encountered before. With some work I was able to get through, and added the new algorithm to a growing toolchest. -- GreenC 20:01, 24 July 2026 (UTC)
Enwiki
- Checked 231 pages and edited 229 pages. Moved 285 links to a new URL: 17 normal redirects, 268 ruled mapped redirects, Removed 1
{{dead link}}. Switched 151|url-status=deadto live.
IABot DB
- Updated 337 URLs
Done -- GreenC 20:38, 24 July 2026 (UTC)
onesports.ph
[edit]The "onesports.ph" website was redesigned sometime in mid-March 2026 and had the side effect of removing the authors from any article made before the redesign. Fortunately, all articles after the redesign are authored so it is more of a bug than anything suspicious.
Still, I think it would be useful if every article from before the redesign is archived, since a lot of them are authored. For a more specific timeframe, the earliest authored articles I could find after the redesign were from March 15, 2026, so anything before that will likely need to be archived. MarcusAbacus (talk) 14:29, 11 July 2026 (UTC)
- MarcusAbacus, I don't understand the request. Can you provide some examples? Thanks. -- GreenC 20:40, 24 July 2026 (UTC)
- Let's take this article announcing a recent major signing, published on January 1, 2026. Accessing the article now will have a "by One Sports" tag. Using the Internet Archive, accessing the same article will have an author attached, Kiko Demigillo. For an earlier example, this article from February 15, 2024, when put through the Archive, is found to be made by Ohmer Bautista. The earliest articles that weren't affected by the redesign date back to March 15, 2026, such as this article, which still has the author attached. MarcusAbacus (talk) 01:56, 25 July 2026 (UTC)
- This is a fuzzy situation. With Eya Laure live page: 1. Missing author in the header section. 2. Includes an embedded video 3. Good quality formatting 4. Has the author in the footer. With the archive version: 1. Includes author in header section. 2. Missing embedded video. 3. Low quality formatting 4. Has author in the footer. .. so both pages have the author in the footer. The main problem is the missing video, and lower quality formatting in the archive. I am inclined not to move it around without a clear improvement. -- GreenC 03:45, 25 July 2026 (UTC)
- Let's take this article announcing a recent major signing, published on January 1, 2026. Accessing the article now will have a "by One Sports" tag. Using the Internet Archive, accessing the same article will have an author attached, Kiko Demigillo. For an earlier example, this article from February 15, 2024, when put through the Archive, is found to be made by Ohmer Bautista. The earliest articles that weren't affected by the redesign date back to March 15, 2026, such as this article, which still has the author attached. MarcusAbacus (talk) 01:56, 25 July 2026 (UTC)
monthlyreview.org
[edit]Majority of these redirect to new URLs, like this to that for List of ghost towns in Pennsylvania. However, this one is broken at Lyndall Urwick as it redirects to search results. ~240. Thanks! MrLinkinPark333 (talk) 00:39, 12 July 2026 (UTC)
Enwiki
- Checked 541 pages and edited 501 pages. Moved 578 links to a new URL: 278 normal redirects, 190 ruled mapped redirects, 110 ghost mapped redirects, Resolved 28 soft-404s. Added 2
{{dead link}}. Switched 23|url-status=deadto live. Switched 3|url-status=liveto dead. Added 61 archive URLs (61 Wayback).
IABot DB
- Checked 992 URLs and updated 342
Done -- GreenC 03:25, 25 July 2026 (UTC)
nhl.com
[edit]Many nhl.com links are now 404ing. For example, each citation (and links to recaps) for 2009 Stanley Cup playoffs are now broken. DiophantineEquation (talk) 18:55, 14 July 2026 (UTC)
- This domain (> 17,000 pages) has problems, not least there was a mass casualty event that eliminated 10s of thousands of URLs displaying a blank page. this (blank page) can now be found here. This rule works most of the time, but some not. In others the slug changed like this is now here (nesns vs. nesn-s) .. there is a method to find these searching the Wayback CDX records. -- GreenC 17:31, 25 July 2026 (UTC)
Enwiki
- Batch 00000-03000: Checked 3,000 pages and edited 1,373 pages. Moved 8,697 links to a new URL: 3,169 normal redirects, 4,984 ruled mapped redirects, 544 ruled inferred mapped redirects, Resolved 4,228 soft-404s. Added 551
{{dead link}}. Switched 237|url-status=deadto live. Switched 656|url-status=liveto dead. Added 2,984 archive URLs (2,984 Wayback).
- Batch 03000-09000: Checked 6,000 pages and edited 2,810 pages. Moved 20,113 links to a new URL: 6,643 normal redirects, 12,270 ruled mapped redirects, 1,200 ruled inferred mapped redirects, Resolved 8,650 soft-404s. Added 1,200
{{dead link}}. Switched 605|url-status=deadto live. Switched 849|url-status=liveto dead. Added 6,705 archive URLs (6,705 Wayback).
- Batch 09000-17331: Checked 8,331 pages and edited 3,809 pages. Moved 25,320 links to a new URL: 8,431 normal redirects, 15,475 ruled mapped redirects, 1,414 ruled inferred mapped redirects, Resolved 12,175 soft-404s. Removed 1
{{dead link}}. Added 1,934{{dead link}}. Switched 613|url-status=deadto live. Switched 1,070|url-status=liveto dead. Added 9,946 archive URLs (9,946 Wayback).
IABot DB
In progress checking 116,787 URLs .. will take a while -- GreenC 06:46, 27 July 2026 (UTC)
nasonline.org
[edit]The entire National Academy of Sciences website has been reorganized. We found that the PDF links previously on http://www.nasonline.org/publications/biographical-memoirs are hosted on https://www.nasonline.org/wp-content/uploads/2024/06/ (mostly, though some are in other wp-content folders). We haven't found the other new locations yet. Ping @Jethro8 @syntaxTerror. Thanks, Thanks, Framawiki (please notify me when you reply) 08:06, 15 July 2026 (UTC)
- We could use some help reaching out to their webmaster.
- Other urls aren't easily corrected, for instance:
- Member pages:
- Old urls:
- http://www.nasonline.org/member-directory/members/20004883.html
- http://www.nasonline.org/member-directory/members/56393.html
- New urls:
- https://www.nasonline.org/directory-entry/haim-brezis-a6fxxh/
- https://www.nasonline.org/directory-entry/ruth-s-defries-xygotf/
- News items:
- Old url:
- http://www.nasonline.org/news-and-multimedia/news/2021-nas-election.html
- New url:
- https://www.nasonline.org/news/national-academy-of-sciences-elects-new-members-including-a-record-number-of-women-and-international-members/
- Jethro8 (talk) 08:16, 15 July 2026 (UTC)
The URLs are so divergent, there is not enough information to create conversion rules. It would either require a map provided by NAS, or even better, NAS creates HTTP/S redirects which is a standard practice. The backstop is convert to Wayback archives. -- GreenC 03:41, 26 July 2026 (UTC)
- Good news. The trailing slug ("/haim-brezis-a6fxxh") is optional: https://www.nasonline.org/directory-entry/haim-brezis .. thus any old "/member-directory/" path can be inferred based on name .. they account for about 40%. The /news path is not recoverable but it only accounts for about 15%. The /biographical-memoirs like this is now here (new path, same filename). The /member-directory like this is now here (need inference check with first-last and first-middle-last). -- GreenC 22:29, 26 July 2026 (UTC)
- Hello, I managed to replace the majority of http://www.nasonline.org/member-directory/members urls on frwiki using the index provided by the sitemap files [6] and a distance comparaison script.
- If somebody wants to run the same here on enwiki, you can download all "directory-entry-*" sitemaps in a folder, and run the following phab:P95321. The only error I identified is for George F. Gao [7]. Thanks, Framawiki (please notify me when you reply) 09:09, 28 July 2026 (UTC)
- Pages prefixed by http://www.nasonline.org/member-directory/deceased-members/ was moved to the same destination, so the same script is working too for them. Thanks, Framawiki (please notify me when you reply) 14:06, 28 July 2026 (UTC)
- Good news. The trailing slug ("/haim-brezis-a6fxxh") is optional: https://www.nasonline.org/directory-entry/haim-brezis .. thus any old "/member-directory/" path can be inferred based on name .. they account for about 40%. The /news path is not recoverable but it only accounts for about 15%. The /biographical-memoirs like this is now here (new path, same filename). The /member-directory like this is now here (need inference check with first-last and first-middle-last). -- GreenC 22:29, 26 July 2026 (UTC)
De Imperatoribus Romanis
[edit]The old URLs roman-emperors.org and www.roman-emperors.org of online encyclopedia De Imperatoribus Romanis have been usurped and should be replaced with the current roman-emperors.sites.luc.edu. I also link a related discussion opened 26 June 2026 at Wikipedia talk:WikiProject Classical Greece and Rome#Hijacked source, content relocated - but is it RS?. If you need more info, please kindly ping me. Thank you, Rotideypoc41352 (talk · contribs) 14:50, 15 July 2026 (UTC)
MySwar
[edit]Please change from myswar.com or myswar.in since the present domain link is https://myswar.co. Kailash29792 (talk) 05:07, 16 July 2026 (UTC)
amws.com.au
[edit]It has been usurped. Jerod Lycett (talk) 00:12, 18 July 2026 (UTC)
iffhs.com
[edit]IFHSS (https://iffhs.com) has updated its website structure, affecting both the visual layout and the URL links.
Consequently, many references need to be updated, and the question is whether a bot can handle these updates or if everything must be changed manually, e.g. on IFFHS World's Best Player for the ref. with "IFFHS WORLD'S BEST MAN PLAYER OF THE DECADE 2011-2020" as title:
- old : https://iffhs.com/index.php/posts/954
- new: https://iffhs.com/en/news/iffhs-best-man-player-world-of-the-decade-2011-2020-954
The number at the end (in this case, "954") in the old url-structure is retained in the same way for all URLs in the new structure; the only difference is that a title now precedes it. Miria~01 (talk) 08:33, 21 July 2026 (UTC)
bac-lac.gc.ca
[edit]These links to RPM music charts have recently been changed. Per Template talk:Single chart, I think all of these links will need archives. This is because the new URLs have an ID number added to them and is not consistent. ~5k. Thanks! MrLinkinPark333 (talk) 20:37, 21 July 2026 (UTC)
statsnz.maps.arcgis.com
[edit]Please update all references for https://statsnz.maps.arcgis.com/apps/webappviewer/index.html?id=6f49867abe464f86ac7526552fe19787 to https://www.stats.govt.nz/geographic-boundary-viewer/ the app for the old link is superceded by a new app and is referenced on hundreds of pages ~2026-40810-66 (talk) 03:34, 22 July 2026 (UTC)
magic.wizards.com
[edit]magic.wizards.com- 215
I'm not sure if all links with the en/articles/archive & en/content path are dead or if there are other dead paths in the Magic part of the Wizards of the Coast website. wizards.com/Magic/Magazine was previously addressed. Sariel Xilo (talk) 18:14, 23 July 2026 (UTC) Update: found one that is a redirect instead of a 404. Sariel Xilo (talk) 18:40, 23 July 2026 (UTC)
barbadosadvocate.com
[edit]Former national newspaper of the eponymous Caribbean island (1895-2023), which ceased operations amid ownership disputes and salary issues. Nowadays represented by archive feeds at FB/IG, as well as the University of Florida's incomplete dLOC collection. (Paging @CaribDigita/WP:BARBADOS...) --Slgrandson (How's my egg-throwing coleslaw?) 22:14, 27 July 2026 (UTC)