Why 'does it answer?' is not enough
A supplier's image address can fail in at least eight different ways, and most only look like success if you check the status code. The address may return the image. It may be missing, with a 404 or 410. It may redirect once to the real image, or many times, or in a circle. It may return a normal web page saying not found, with a success status. It may return a file whose real type does not match its extension. It may return an image far too small for your store. Or the host may never answer at all. The HTTP specification says a client should detect circular redirects and that a server's declared media type can be wrong, because servers are often misconfigured. A check has to look at what actually comes back, not at whether something did.
If the matrix is wider than the box, scroll horizontally to read every column. Keyboard: focus the matrix and use Left/Right.
case result to record
working image (200, image/jpeg) ok
missing file (404 or 410) missing
redirect to an image (301/308) ok, final address noted; permanent
redirect loop fail, stop after a fixed number of hops
error page returned with 200 not an image (type is text/html)
file type does not match extension mismatch, record both
image below the minimum size too small
host does not answer in time timeout, product not blocked
(synthetic case list; no real supplier host)Redirects, and what to store
A permanent redirect (301 or 308) means the image has a new address, and storing the final address saves a hop on every page view. A temporary one (302 or 307) says keep the original, so do not overwrite it. A chain of redirects should be limited to a small fixed number and a loop should be cut off as a failure. The specification mentions that older editions suggested a limit of five; choose a number and write it down.
HEAD or GET, and timeouts
A HEAD request is cheaper because it returns no content. The specification says the headers should match a GET, but a server may leave out values it only knows when it builds the content, such as Content-Length, so some hosts answer HEAD differently. A robust check falls back to a real fetch of the start of the file when the answer to HEAD lacks the type or size. Every request needs a time limit, otherwise one dead host holds up the whole import. Results are cached per address within a run so a thousand products that share one placeholder are not a thousand requests.
- Respect the supplier's terms and robots file; do not automate requests to a host that forbids it.
- Keep the request rate low. Ask the supplier for a bulk image archive where available.
Refuse addresses that are not on the public internet
The addresses come from a supplier's file, so a check that fetches them is your server making requests that someone else chose. A feed that lists an address pointing at your own internal network, at the machine itself or at a cloud metadata address such as 169.254.169.254 could make your server request that address and record what comes back. OWASP calls this server-side request forgery, names those ranges among the destinations to block, and warns that following redirects can get around a check made only on the first address.
So the check should resolve each address before it connects and refuse it, reporting it as refused, if it is a loopback, private-network or link-local address. It should repeat that check on every redirect, since a public address can redirect to a private one, and connect only to the address it checked. OWASP prefers allow-lists to deny-lists, which are easy to get round, so where you know which hosts the supplier uses for images, request only those hosts.
- Test it with a fixture: a loopback address, a private-network address, a link-local address and a public address that redirects to one of them. None of the four should cause a request to reach the target.
- A local test server on the loopback address is needed for the other tests, so make that exception a test-only setting that is off in the delivered configuration.
What to do with a bad image
Agree one rule and apply it everywhere: hold the product for review, or import it without an image and flag it. Do not drop it silently, and do not substitute a placeholder image into a shopping feed; Google's image rules disallow placeholders, as the Merchant Center image guide explains. Write a one-line report per product with the original address, the final address, the status, the declared and detected type, the size and the action taken, so a person can see what happened.
A safe first investigation
Paste three image addresses from the feed into a browser: one that works, one that is missing in your store and one that redirects. Note what each shows. A page that says not found but loads normally is the case a status check misses. Do not use real supplier credentials for this and do not send us any.
How the paid job is accepted
The job feed-image-url-validation-before-import is £345 for one import path and one supplier layout. A local fixture of eight cases, like the list above, must classify each as expected; four further fixture cases (a loopback address, a private-network address, a link-local address and a public address that redirects to one of them) must each be refused without a request reaching the target; the agreed 20-row sample must finish within the time limit even though one host never answers; the per-product report must list the fields above; and existing tests must pass. We check only public addresses you confirm may be requested, we do not copy or edit images, and a working address does not mean a shopping platform will accept the image. Prices are untested proposals, and payment follows the agreed checks and your sign-off. Nothing is booked or charged by an enquiry.
Sources and limits
- RFC 9110: HTTP Semantics Checked 2026-10-11.
- A HEAD response carries no content and should have the same header fields as GET, but a server may omit fields known only while generating content, such as Content-Length.
- A client should detect and intervene in cyclical redirections; 301 and 308 are permanent and 302 and 307 temporary; 307 and 308 keep the request method.
- 410 means the resource is probably permanently gone, while 404 does not say whether the absence is temporary or permanent.
- Content-Type states the sender's declared media type, and servers are often misconfigured, so declared and real types can differ.
- Google Merchant Center Help: image_link Checked 2026-10-11.
- Placeholder and generic images are not allowed, with listed exceptions for some product categories.
- OWASP Server-Side Request Forgery Prevention Cheat Sheet Checked 2026-10-11.
- Server-side request forgery abuses an application to interact with the internal or external network or the machine itself, and mishandled user-supplied URLs are a main enabler.
- Its block list names loopback ranges, the RFC 1918 private ranges, IPv6 unique-local and link-local ranges and the cloud metadata address 169.254.169.254, and warns that deny-lists are bypass-prone and allow-lists are preferred.
- Following redirects can bypass input validation; the resolved IP address must be checked, and the client should connect only to validated addresses.