Extract Images from HTML

Paste a page's HTML to list its images, see which ones lack alt text, and export the URLs.

To get a page's HTML, open it, right-click and choose View page source (Ctrl + U in most browsers), then select all and copy. For images added by JavaScript, copy the rendered page instead: in DevTools, right-click the <html> element and choose Copy, then Copy outerHTML. This tool doesn't fetch pages. Everything runs in your browser and nothing is uploaded.

Turns relative URLs into full ones. A <base href> in the HTML takes priority, as in browsers.
.html or .htm, up to 20 MB. You can also drop a file on the text box.
Ctrl + Enter extracts. Results update as you type, except for HTML over 5 MB.

What the tool finds

Paste a page's HTML and the tool lists every image reference in the markup, with the place it came from:

  • <img src> and each candidate in srcset, with its descriptor such as 480w or 2x, plus <source srcset> inside <picture>.
  • Lazy-load attributes used by JavaScript libraries: data-src, data-srcset and data-lazy-src.
  • Social images from og:image and twitter:image, favicons, Apple touch icons and images loaded early with <link rel="preload" as="image">.
  • url(...) in inline style attributes, <input type="image">, SVG <image href> and video posters.

Relative URLs are resolved against the page URL you enter, or against the page's own <base href>, as a browser would do. Images embedded as data: URLs appear as "data: URL (image/gif, 42 bytes)" instead of a wall of base64. The HTML is parsed in your browser into an inert document, so its scripts don't run and its images don't load. Nothing is uploaded.

How to use it

  1. Open the page, choose View page source (Ctrl + U in most browsers), select all and copy. Paste it into the box, or open a saved .html file.
  2. Enter the page URL so relative paths become full URLs.
  3. Read the audit, then narrow the table: pick a source type, tick Missing alt only, or search for a folder or file name. Unique URLs only merges repeats and adds a Count column.
  4. Download CSV saves every matching row (UTF-8 with a byte order mark, so Excel shows accents correctly). Copy URLs copies each URL once, one per line.

The include options switch whole groups on or off: srcset candidates, lazy-load attributes, icons and social images, and inline style backgrounds. Thumbnails are off by default, because showing them makes your browser request each image from its server. When you turn them on, they load lazily and without a referrer.

A worked example

With the page URL https://www.example.com/blog/, this HTML:

<img src="/img/logo.svg" alt="Example Outdoors" width="140" height="40">
<img src="grip.jpg" srcset="grip-480.jpg 480w, grip-960.jpg 960w"
     alt="Close-up of a lugged outsole" loading="lazy">
<img src="divider.png" alt="" width="600" height="8">
<img data-src="lacing.jpg" class="lazyload">

gives six rows:

URLSourceAltSize attributesLoading
https://www.example.com/img/logo.svgimg srcPresent140 × 40Not set
https://www.example.com/blog/grip.jpgimg src, srcset: 2 candidatesPresentMissinglazy
https://www.example.com/blog/grip-480.jpgimg srcset, 480wPresentMissinglazy
https://www.example.com/blog/grip-960.jpgimg srcset, 960wPresentMissinglazy
https://www.example.com/blog/divider.pngimg srcEmpty600 × 8Not set
https://www.example.com/blog/lacing.jpgLazy-load attribute, data-srcMissingMissingdata-src

The audit reports four <img> elements: one without alt (fix it), one with empty alt (check that it's decorative), and two without width and height.

Checking alt text

Screen readers read alt text aloud, and browsers show it when an image fails to load. Google says in its image SEO guide that it uses alt text along with computer vision and the page content to understand an image. The same guide warns that filling alt attributes with keywords gives a poor experience and may make a site look like spam. Its example ranks alt="puppy" as better and alt="Dalmatian puppy playing fetch" as best. The W3C's alt text tips ask for the most concise description of the image's purpose, and say words like "image" or "picture" are usually unnecessary because screen readers already announce an image.

The tool marks alt text that looks like a file name (IMG_2041.jpg), is one generic word, or starts with "image of" as something to review, not as an error.

Empty alt is not an error

A missing alt attribute is a real problem: the W3C decorative images tutorial notes that some screen readers then announce the file name. But alt="" is the correct markup for a decorative image, such as a divider, because it lets assistive technology skip it. That is why the audit labels missing alt "Fix" and empty alt "Check". One exception: an image that is the only content of a link needs real alt text, or the link has no name.

Width, height and layout shift

Modern browsers use the width and height attributes to work out an image's aspect ratio and reserve its space before the file arrives. Without them, content jumps as images load, which adds to Cumulative Layout Shift (CLS). web.dev recommends always setting both, or reserving the space with CSS aspect-ratio. The tool only reads attributes, so an image sized by a stylesheet still counts as missing: treat that count as a list to check.

srcset and lazy loading

srcset offers several files for one image: a w descriptor gives each file's width and works with sizes, and an x descriptor gives the pixel density. The browser picks one, so each candidate gets its own row. Google recommends keeping a fallback URL in src, because some crawlers don't understand srcset.

loading="lazy" delays offscreen images until the visitor scrolls near them. Don't use it on images visible when the page opens, especially the largest one: web.dev advises lazy-loading only images outside the initial viewport. Older sites keep the real URL in data-src and swap it in with a script, often with a placeholder in src, which is why the tool lists both.

Limits

  • It reads only the HTML you give it. Images that JavaScript adds later are not in the page source. Copy the rendered DOM instead: in DevTools, right-click the <html> element and choose Copy outerHTML.
  • Images set in stylesheets or <style> blocks aren't listed, only inline style attributes. Google doesn't index CSS images anyway, so an important image belongs in an <img>.
  • It doesn't check that the URLs work. The HTTP Header Checker shows an image URL's status code, redirects and Content-Type.
  • The table shows the first 1,000 rows; the CSV and Copy URLs include them all.

To audit the links on the same page, paste the HTML into Extract links from HTML.

Frequently asked questions

How do I get all image URLs from a web page?

Copy the page's HTML with View page source, paste it here and enter the page URL. Every image is listed with a full URL, and Copy URLs gives you the list one per line.

How do I find images without alt text?

Paste the HTML and tick Missing alt only. The audit also counts images with an empty alt, which is fine for decorative images, and alt text that looks like a file name.

Is alt="" the same as leaving out the alt attribute?

No. An empty alt marks the image as decorative, so screen readers skip it. With no alt attribute at all, some screen readers read the file name aloud instead.

Can the tool fetch a page from its URL?

No. It only reads HTML that you paste or open as a file. The Page URL field is used to turn relative image paths into full URLs.

Why are some images on the page not listed?

Images added by JavaScript after the page loads are not in the page source, and background images set in CSS files or style blocks are not read. For images added by scripts, copy the rendered HTML from DevTools with Copy outerHTML and paste that.

Is my HTML uploaded anywhere?

No. The HTML is parsed in your browser and never sent to a server. Input under 200 KB is kept in your browser's local storage so it is still there when you come back.