How to Automatically Find Broken Links and Missing Images
Build an automated crawler that finds broken links and missing images, logs each issue to Notion, and routes fixes to an owner, so you skip the manual audits.
You can catch broken links and missing images automatically by running a scheduled crawl, checking every page for missing cover images and non-200 link responses (including bot-blocked URLs), then logging issues into a remediation queue.
Automated crawler checking a site for broken links and missing images. Photo by Mohammad Rahmani on Unsplash
Why automate broken link and missing image checks?
Manual audits are reactive and incomplete: pages ship every day, links rot quietly, and a “quick check” rarely covers every hub and blog page.
A scheduled crawler gives you three things that matter:
Consistency: every published page gets checked on the same cadence.
Proof: each issue includes the exact page + failing URL/image.
Workflow: issues become trackable work (not a one-off spreadsheet).
What this crawler should detect (minimum viable checks)
Start with checks that are easy to measure and easy to fix.
1) Missing cover images on hub/blog pages
If your site is generated from a CMS (including Notion/Bullet-style databases), “missing images” often means:
cover image not set
image URL present but returns 404/403
image blocked by hotlink protection or bot filtering
What to log per issue
page URL
what image field is missing (or the failing image URL)
HTTP status (if applicable)
first-seen timestamp
2) Broken or bot-blocked external links
A link can be “bad” even when it’s not a 404:
403 / 429: bot protection, rate limiting, or WAF rules
301/302 chains: still works, but adds latency and sometimes breaks later
timeouts: destination is down or slow
For SEO and user experience, the most important rule is simple: a visitor (and a crawler) should be able to reach the destination without errors.
What to log per issue
source page URL
destination URL
HTTP status + final resolved URL (after redirects)
whether the issue is “hard fail” (404/410) vs “blocked” (403/429)
The crawl → detect → log → remediate pattern
We run this exact loop on the Connex site. The architecture looks like this:
Pull the crawl seed list
Start with your known published pages list (preferred). If you don’t have one, build it from your sitemap.
Fetch each page and parse it
Extract: canonical URL, cover image reference (if you can access it), and all outbound links.
Validate assets and links
For images: verify the URL exists and returns a successful response.
For links: request the URL, follow redirects, record status.
Write results to a Notion database
One row per issue (not per crawl run).
Update “last seen” if the issue persists.
Trigger remediation
Assign issues to an owner, or route them by type (cover image vs link).
Avoid noisy results (so the crawler stays useful)
A crawler that flags the same “expected” issues forever becomes background noise. Add two controls early:
Muting (suppress known exceptions)
Some URLs will always look “broken” to bots (common with affiliate links, gated content, or certain CDNs). Add a Muted toggle with a reason.
Retries + thresholds
Don’t create an issue on the first failure. Use:
2–3 retries with backoff
only log after N consecutive failures
Implementation checklist (technical walkthrough)
Define what counts as “published page” and build the seed list
Add a Notion database for issues (fields: Page URL, Issue type, Target URL, Status code, Muted, First seen, Last seen, Owner)
Implement crawling with rate limits and robots.txt respect
Detect missing covers for hub/blog pages
Validate external links (follow redirects, classify blocked vs broken)
De-duplicate issues (one issue row per page+target)
Add mute + retry logic
Add a remediation view for “New issues” and “Needs work”
Where to run this crawler
You can run the crawl as:
a daily scheduled job (best default)
an on-demand button for a “pre-publish QA” run
If your content operations already live in Notion, storing issues in Notion keeps the workflow tight.
Tools that fit this workflow
If your CMS content is managed in Notion, pairing a crawler with Notion automations can keep the maintenance loop short. For integrations and routing, tools like Notion and Zapier are common building blocks.
Get help building your link checker
Building a crawler like this usually breaks at the noise stage: bot-blocked URLs and one-off timeouts flood the issue queue until nobody looks at it. If you have hit that wall, book a ZoomFlow session. One of our consultants will build the crawl, the Notion issue database, and the mute and retry logic with you live, and you own the workflow when the call ends.
Build a scheduling form that creates Google Calendar events, Trello cards, and QuickBooks customers automatically, capturing customer info once, not 3 times.