BeautifulSoup getting href

Question

I have the following soup:

<a href="some_url">next</a>
<span class="class">...</span>

From this I want to extract the href, "some_url"

I can do it if I only have one tag, but here there are two tags. I can also get the text 'next' but that's not what I want.

Also, is there a good description of the API somewhere with examples. I'm using the standard documentation, but I'm looking for something a little more organized.

Please post a code sample to show how you're trying to do it — seb, Apr 28 '11 at 08:33
Alright, I figured it out: soup.find('a')['href'] The thing that confused me was that I was using django (html) to see it, which actually removes the href before presenting it: soup.find('a') becomes only 'next' — dkgirl, Apr 28 '11 at 08:38
True, this question is a duplicate. Yet the beauty of @MarkLongair's answer makes it precious, even a few years later. — Giampaolo Ferradini, Sep 02 '19 at 18:46

score 528 · Answer 1 · edited Jun 09 '22 at 22:09

528

You can use find_all in the following way to find every a element that has an href attribute, and print each one:

# Python2
from BeautifulSoup import BeautifulSoup
    
html = '''<a href="some_url">next</a>
<span class="class"><a href="another_url">later</a></span>'''
    
soup = BeautifulSoup(html)
    
for a in soup.find_all('a', href=True):
    print "Found the URL:", a['href']

# The output would be:
# Found the URL: some_url
# Found the URL: another_url

# Python3
from bs4 import BeautifulSoup

html = '''<a href="https://some_url.com">next</a>
<span class="class">
<a href="https://some_other_url.com">another_url</a></span>'''

soup = BeautifulSoup(html)

for a in soup.find_all('a', href=True):
    print("Found the URL:", a['href'])

# The output would be:
# Found the URL: https://some_url.com
# Found the URL: https://some_other_url.com

Note that if you're using an older version of BeautifulSoup (before version 4) the name of this method is findAll. In version 4, BeautifulSoup's method names were changed to be PEP 8 compliant, so you should use find_all instead.

If you want all tags with an href, you can omit the name parameter:

href_tags = soup.find_all(href=True)

edited Jun 09 '22 at 22:09

JayRizzo

3,234
3
33
49

answered Apr 28 '11 at 08:38

Mark Longair

446,582
72
411
327

4

can you get the single href with the class "class="class"" – yoshiserry May 19 '14 at 01:24
12

@yoshiserry soup.find('a', {'class': 'class'})['href'] – rleelr Jan 08 '17 at 16:15
1

How do you attenuate false positives and unwanted results (i.e. `javascript:void(0)`, `/en/support/index.html`, `#smp-navigationList`)? – voices Feb 12 '18 at 10:28
Hello how can I get the 'next' value in href. ```NEXT``` – abdoulsn Oct 11 '19 at 13:14
1

@abdoulsn `soup.find('a').contents[0]` – CharukaHS Jan 04 '20 at 02:18
I believe in your last line, href needs to be a string 'href' – Uralan Jul 09 '20 at 19:31

BeautifulSoup getting href

1 Answers1

Linked

Related