Extract Links from HTML

Paste a page's HTML to get every link with its anchor text and rel values, in a table you can export.

This tool doesn't fetch pages. To get the HTML, open the page and press Ctrl + U (Cmd + Option + U in Chrome on a Mac) to view the source, then select all and copy. For links added by JavaScript, open DevTools (F12), right-click the <html> element in Elements and choose Copy, then Copy outerHTML. The HTML is read with the browser's DOMParser, which doesn't run scripts or load images, and nothing is uploaded.

Where the HTML came from. Used to resolve relative links and to tell internal from external.
.html or .htm, up to 50 MB. You can also drop a file on the text box.
Links update as you type. Ctrl + Enter extracts now.

What the link extractor finds

Paste a page's HTML and the tool lists every link in it, in document order: each <a href> and each <area href> from an image map. For each link you get the resolved URL, the href as written, the anchor text, the rel values, the target and a type. With Include <link> tags checked, it also lists canonical, alternate (hreflang), prev and next links from the head. An <a> without an href, such as a named anchor, isn't a link, so it is counted in the status line but not listed.

The HTML is read with the browser's DOMParser, which doesn't run scripts or load images, and nothing is uploaded.

How to get a page's HTML

The tool doesn't fetch pages, so you paste the HTML yourself:

  • Page source. Press Ctrl + U (Cmd + Option + U in Chrome on a Mac), select all and copy. This is the HTML the server sent, before any JavaScript ran.
  • Rendered DOM. Open DevTools (F12), right-click the <html> element in the Elements panel and choose Copy, then Copy outerHTML. This includes links that scripts added after the page loaded.

Or open a saved .html file. Either way, enter the page's address in Page URL so relative links can be resolved.

How links are resolved and classified

Relative hrefs such as ../guides/ or /pricing are resolved against the page URL, the way a browser resolves them. A <base href> in the HTML takes priority, and a relative base is itself resolved against the page URL; the status line says when this happens. Protocol-relative links such as //cdn.example.org/file.pdf take the page's scheme.

TypeMeaning
Internal Same host as the page URL. With Treat www and non-www as the same site on (the default), www.example.com and example.com match. Other subdomains are external.
ExternalAny other host.
AnchorA #fragment on the same page, such as a skip link.
Email, Phone, JavaScriptmailto:, tel: and javascript: links.
OtherOther schemes, such as ftp: or data:, and hrefs that can't be resolved.
Relative, Absolute Only when there is no page URL: relative hrefs stay as written, and absolute ones can't be split into internal and external.

A worked example

With the page URL https://example.com/blog/post, this HTML:

<a href="/pricing">Pricing</a>
<a href="https://partner.example.net/" rel="sponsored">Partner</a>
<a href="#faq">FAQ</a>
<a href="/team/"><img src="team.jpg" alt=""></a>

gives these rows:

URLAnchor textTypeRel
https://example.com/pricingPricingInternal
https://partner.example.net/PartnerExternalsponsored
https://example.com/blog/post#faqFAQAnchor
https://example.com/team/[image: no alt], marked EmptyInternal

Internal links, external links and SEO

Search engines find pages by following links. According to Google's link best practices, Google can only crawl a link reliably if it is an <a> element with an href that resolves to a real web address; the JavaScript type flags javascript: links. Google also recommends that every page you care about has a link from at least one other page on your site.

Choose Internal with Unique URLs only to see which pages a template links to, then compare the list with your sitemap (or build one with the sitemap generator) to spot sections it never links to. Check external links for redirects and errors with the HTTP header checker.

What nofollow, sponsored and ugc mean

Google's guide to qualifying outbound links recommends rel="sponsored" for ads and paid placements, rel="ugc" for user-generated content such as comments and forum posts, and rel="nofollow" when the others don't apply and you'd rather Google not associate your site with the linked page or crawl it from your site. Values can be combined, as in rel="ugc nofollow".

Google introduced sponsored and ugc on September 10, 2019, and said that all three values now work as hints for ranking rather than as strict instructions. For crawling and indexing, nofollow became a hint on March 1, 2020. Google says these links will generally not be followed, but the linked page can still be found in other ways, such as sitemaps. The tool counts a link as qualified if it has any of the three. It also shows noopener and noreferrer, which are about security and privacy, not ranking.

Empty anchor text and accessibility

A link with no text gives screen reader users nothing useful to hear, and search engines no anchor text. The tool takes the link's text first, then the alt text of an image inside the link, shown as [image: alt], then an aria-label or title attribute. A link with none of these is counted under Empty anchor text and marked Empty. To fix it, add visible text, alt text on the image, or an aria-label. Google documents using image alt text as anchor text and the title attribute as a fallback. To audit the images themselves, use Extract Images from HTML.

Filters, export and limits

  • Filters apply to every link. The table shows the first 1,000 rows; Download CSV and Copy URLs include all matching rows. The CSV is UTF-8 with a byte order mark, so Excel shows accents correctly.
  • Links update as you type for HTML up to about 5 MB. For larger HTML, press Extract links.
  • Links added by JavaScript appear only if you copy the rendered DOM from DevTools. URLs inside CSS, scripts, JSON data or iframes aren't listed.
  • Broken markup is parsed as a browser would, so an unclosed tag doesn't hide later links. The tool doesn't check whether links work.

Frequently asked questions

How do I extract all links from a web page?

Copy the page's HTML, either the page source (Ctrl + U) or the rendered DOM from DevTools (Copy outerHTML), and paste it here with the page URL. You get every link with its anchor text, rel values and type, and can download them as CSV or copy the URLs.

Why are my links listed as Relative instead of Internal?

There is no page URL, so relative hrefs like /about can't be resolved and there is no site to compare hosts with. Enter the address the HTML came from and the links are resolved and split into internal and external.

Does it find links that JavaScript adds to the page?

Only if they are in the HTML you paste. The page source from Ctrl + U doesn't include them, so copy the rendered HTML from the Elements panel in DevTools instead.

What is the difference between nofollow, sponsored and ugc?

Google asks for sponsored on ads and paid links, ugc on links in comments and forum posts, and nofollow when neither applies and you'd rather Google not associate your site with the linked page. Google treats all three as hints rather than strict rules.

How do I get a list of unique URLs?

Check Unique URLs only. Each URL is kept once, at its first position, and the Count column shows how many times it appears. Copy URLs then gives one URL per line.

Is the HTML I paste uploaded?

No. It is parsed in your browser with DOMParser, which doesn't run the page's scripts or load its images. HTML under 200 KB is saved in your browser's local storage so it is still there when you come back.