So I'm working on a (English) site which at one point had many translated pages (using Weglot).
Wasn't working suitably, so all the pages and the folders they were in (/it. /es, /fr etc etc) were deleted.
But....those pages had already been indexed before deletion. They're not reachable in any way. Not just the pages/posts of course, but all the usual rubbish URLs that get generated alongside (/?sid, ?/action etc etc etc)
They're showing up in various error/warning reports in GSC (thousands overall).
I'm assuming that overall they may not be hurting (though could in theory be contributing to an overall thin content percentage measure??). But equally, since G appears to think they still exist, they could be interfering with crawl budget and, resulting from that, having an overall negative effect.
The question is how to get rid of them so they're not a problem either way.
I think (assuming that none are linked to by other pages/sites and the sitemap(s) don't contain them either) that instructing robots.txt to return a 410 on every offending URL will do the trick. Still might take time, though.
Anyone done this before? (successfully or not)
Any other approach I'm missing?
Is it possible I'm overthinking it anyway, and just ignoring it is a viable way to proceed?
If the 410 approach is the right one, what code do you actually put in the robots.txt? (I can probably get this from AI, but would be nice to know if someone has used in anger and seen it work).


LinkBack URL
About LinkBacks
Reply With Quote


