Extract Links from HTML
Paste a page's HTML to get every link with its anchor text and rel values, in a table you can export.
This tool doesn't fetch pages. To get the HTML, open the page and press Ctrl + U (Cmd + Option + U in Chrome on a Mac) to view the source, then select all and copy. For links added by JavaScript, open DevTools (F12), right-click the <html> element in Elements and choose Copy, then Copy outerHTML. The HTML is read with the browser's DOMParser, which doesn't run scripts or load images, and nothing is uploaded.
Links
- Total links
- 0
- Internal
- 0
- External
- 0
- Nofollow, sponsored or ugc
- 0
- Empty anchor text
- 0
- Unique URLs
- 0
| # | URL | Anchor text | Type | Rel | Count |
|---|
URLs are shown as text, not links, so nothing from the pasted page can open by accident. Count is how many of the matching links point to that URL. The CSV and the copied URLs contain every matching row, including rows beyond the first 1,000 shown. CSV cells that start with =, +, - or @ get a leading apostrophe so spreadsheets treat them as text.
What the link extractor finds
Paste a page's HTML and the tool lists every link in it, in document order: each <a href> and each
<area href> from an image map. For each link you get the resolved URL, the href as written, the
anchor text, the rel values, the target and a type. With Include <link> tags checked, it also
lists canonical, alternate (hreflang), prev and next links from the head. An <a> without an href,
such as a named anchor, isn't a link, so it is counted in the status line but not listed.
The HTML is read with the browser's DOMParser, which doesn't run scripts or load images, and nothing is uploaded.
How to get a page's HTML
The tool doesn't fetch pages, so you paste the HTML yourself:
- Page source. Press Ctrl + U (Cmd + Option + U in Chrome on a Mac), select all and copy. This is the HTML the server sent, before any JavaScript ran.
- Rendered DOM. Open DevTools (F12), right-click the
<html>element in the Elements panel and choose Copy, then Copy outerHTML. This includes links that scripts added after the page loaded.
Or open a saved .html file. Either way, enter the page's address in Page URL so relative links can be resolved.
How links are resolved and classified
Relative hrefs such as ../guides/ or /pricing are resolved against the page URL, the way a
browser resolves them. A <base href> in the HTML takes priority, and a relative base is itself
resolved against the page URL; the status line says when this happens. Protocol-relative links such as
//cdn.example.org/file.pdf take the page's scheme.
| Type | Meaning |
|---|---|
| Internal | Same host as the page URL. With Treat www and non-www as the same site on (the default), www.example.com and example.com match. Other subdomains are external. |
| External | Any other host. |
| Anchor | A #fragment on the same page, such as a skip link. |
| Email, Phone, JavaScript | mailto:, tel: and javascript: links. |
| Other | Other schemes, such as ftp: or data:, and hrefs that can't be resolved. |
| Relative, Absolute | Only when there is no page URL: relative hrefs stay as written, and absolute ones can't be split into internal and external. |
A worked example
With the page URL https://example.com/blog/post, this HTML:
<a href="/pricing">Pricing</a>
<a href="https://partner.example.net/" rel="sponsored">Partner</a>
<a href="#faq">FAQ</a>
<a href="/team/"><img src="team.jpg" alt=""></a> gives these rows:
| URL | Anchor text | Type | Rel |
|---|---|---|---|
| https://example.com/pricing | Pricing | Internal | |
| https://partner.example.net/ | Partner | External | sponsored |
| https://example.com/blog/post#faq | FAQ | Anchor | |
| https://example.com/team/ | [image: no alt], marked Empty | Internal |
Internal links, external links and SEO
Search engines find pages by following links. According to
Google's link best practices,
Google can only crawl a link reliably if it is an <a> element with an href that resolves to a real
web address; the JavaScript type flags javascript: links. Google also recommends that every page you
care about has a link from at least one other page on your site.
Choose Internal with Unique URLs only to see which pages a template links to, then compare the list with your sitemap (or build one with the sitemap generator) to spot sections it never links to. Check external links for redirects and errors with the HTTP header checker.
What nofollow, sponsored and ugc mean
Google's guide to
qualifying outbound links
recommends rel="sponsored" for ads and paid placements, rel="ugc" for user-generated content
such as comments and forum posts, and rel="nofollow" when the others don't apply and you'd rather Google
not associate your site with the linked page or crawl it from your site. Values can be combined, as in
rel="ugc nofollow".
Google
introduced sponsored and ugc
on September 10, 2019, and said that all three values now work as hints for ranking rather than as strict
instructions. For crawling and indexing, nofollow became a hint on March 1, 2020.
Google says these links will generally not be followed, but the linked page can still be found in other ways, such as
sitemaps. The tool counts a link as qualified if it has any of the three. It also shows noopener and
noreferrer, which are about security and privacy, not ranking.
Empty anchor text and accessibility
A link with no text gives screen reader users nothing useful to hear, and search engines no anchor text. The tool
takes the link's text first, then the alt text of an image inside the link, shown as
[image: alt], then an aria-label or title attribute. A link with none of these is counted under
Empty anchor text and marked Empty. To fix it, add visible text, alt text on the image, or an
aria-label. Google documents using image alt text as anchor text and the title attribute as a fallback. To audit the
images themselves, use Extract Images from HTML.
Filters, export and limits
- Filters apply to every link. The table shows the first 1,000 rows; Download CSV and Copy URLs include all matching rows. The CSV is UTF-8 with a byte order mark, so Excel shows accents correctly.
- Links update as you type for HTML up to about 5 MB. For larger HTML, press Extract links.
- Links added by JavaScript appear only if you copy the rendered DOM from DevTools. URLs inside CSS, scripts, JSON data or iframes aren't listed.
- Broken markup is parsed as a browser would, so an unclosed tag doesn't hide later links. The tool doesn't check whether links work.
Frequently asked questions
How do I extract all links from a web page?
Copy the page's HTML, either the page source (Ctrl + U) or the rendered DOM from DevTools (Copy outerHTML), and paste it here with the page URL. You get every link with its anchor text, rel values and type, and can download them as CSV or copy the URLs.
Why are my links listed as Relative instead of Internal?
There is no page URL, so relative hrefs like /about can't be resolved and there is no site to compare hosts with. Enter the address the HTML came from and the links are resolved and split into internal and external.
Does it find links that JavaScript adds to the page?
Only if they are in the HTML you paste. The page source from Ctrl + U doesn't include them, so copy the rendered HTML from the Elements panel in DevTools instead.
What is the difference between nofollow, sponsored and ugc?
Google asks for sponsored on ads and paid links, ugc on links in comments and forum posts, and nofollow when neither applies and you'd rather Google not associate your site with the linked page. Google treats all three as hints rather than strict rules.
How do I get a list of unique URLs?
Check Unique URLs only. Each URL is kept once, at its first position, and the Count column shows how many times it appears. Copy URLs then gives one URL per line.
Is the HTML I paste uploaded?
No. It is parsed in your browser with DOMParser, which doesn't run the page's scripts or load its images. HTML under 200 KB is saved in your browser's local storage so it is still there when you come back.