Links Weren't Linking
For more than a decade I have had a Python link checker I occasionally run against my Links :: vanderwal.net page. I regularly update the links page and the additions are manual inserts and reposting the page when I find something I’m regularly going back to, or want to go back to. About 18 months back my long used python script using Beautiful Soup (HTML parser) - Wikipedia needed updating the BeautifulSoup package and tweaking the script a little bit.
Over these last 18 months the bot / scraper / AI crawler wars have become a hot mess and running a link checker script, unknowingly was getting blocked by hosts / servers that are working to not get slammed by the orders of magnitude of increased traffic from non-organic sources. One method (among many) server owners deplay is blocking access to websites from requests with no user agent or an old user agent claim (as this is what a lot of the aggressive inorganic traffic uses). My old school link checker feels like a web version of Wall-E going around and checking to clean up my broken links and getting caught in the wake from large AI industrial retrieval ships.
Thursday I wanted to run my link checker and realized it wasn’t going to be all that helpful. I thought a quick edit of the script to add a more current user agent to the what is presented to the link’s server would help. I went with: “Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/605.1.15 ” “(KHTML, like Gecko) Version/18.0 Safari/605.1.15”. This one minor change when running the link checker removed about 200 false positive broken links.
There are How Many Links?
With many things I have run and been doing for ages, my estimation of how much or how many of something is wildly off. Another minor tweak I made was to list the number of links tested on my Links :: vanderwal.net page. I have had a links page for two or three years longer than I have had a web domain. I set up this domain in 1997 and the site started with a homepage and a links page of things I wanted near in reach. My first link check I made on Thursday I had 943 links. Well, that was about double of my rough estimate.
Updating to Get Better Data
When I last updated the Python script to a new version of BeautifulSoup I didn’t really change much else. Now that things were less than optimal I noticed my Python script just used a default check, but wasn’t using GET. Adding GET would provide feedback as to the type of HTTP response codes I was getting and that would help reduce my manual checks of the list of “broken links”.
A quick test after adding GET had more updates and parsing needed. If I’m getting the actual responses back, I want to log the returned codes and information, but I really want that organized. The initial output of the checking is a long list of links with non–200 status with their state, the link, error code, if a redirect it tried the redirect and if a 200 it shows the “200” and the new link.
The initial output from running BeautifulSoup is the list in a temporary text document. The last part of the script runs through the text document and turns that list into a grouped list and table markdown file in my Notes directory that Obsidian watches over.
Structured Documents Made the Next Steps Easier
The structured Markdown note has sub-headings for the different returned link states. This works like a really good triage list. I could easily double-check the links marked as broken (404, 409, 500, and no DNS) in the table and added another column to the table for the checked status I provided. I had 52 of these and I went through them in about 15 minutes.
I made a back-up of the links page and verified git was current so I could rollback. I handed off the update to Claude Code to follow my notation of the section I just checked, my structured outcome note, and to follow one of the four processes I posted. It made the updates in seconds and I ran diff against the two files and the outcome was perfect.
I moved the outcome page to production and ran a commit. I ran the Python checker again and tackled a couple of small sections in the same manner. On my 3rd running of the checker I looked at the roughly 430 redirections captured. Since Claude had managed the two prior updates perfectly I had it handle the redirection updates. This too was processed quickly and the output links page was perfect on the manual check. But Claude Code running Opus 5.5 had additional notes from the update that included some redirects were not similar to the page historically that I linked to (in some cases the domain had changed hands years ago and Claude derived context from surrounding links that it was no longer a fit), noted that sites had changed languages, sites were now ads from domain resellers, and more. It had a list of links that could now be problematic. I wasn’t expecting this additional feedback and the feedback I checked was spot on and insanely helpful.
I ran the Python link checker a fourth time and I was down to 916 links on the page. Broken links gone, with many of my “archive” links gone too, which are the gray links as the linked site went from not being updated in years to now fully gone from the living internet. I have left things with this fourth pass of the checker on my to-do list. The links that are large sites with large organizations behind them are in the blocked connection or have “are you human” test in front of them, but still function. I have about 70 or so links that I need to manually go through and log what to do. Some no longer have the content I have long pointed to.
Outcomes
I’m happy with where the Python link checker ended up. The web is a lot more fragile than it seems like it ever was. I was surprised how many links had redirects (I have 50 that still need manual reviews). While I’m bummed many “archive” links are not on the page they were gone and they could be let go.
The two sections on the links page that took the biggest hit are the tech section and the “Urban Planning and Cities” section. Finding sites that fit the urban planning and cities perspective have always been tough to find. Losing some means I need to dig again.
Web Mentions
This post can receive webmentions. Webmentions are currently not displayed, but they are received.