Fetch all href link using selenium in python

ID : 131374

viewed : 8

Tags : pythonseleniumselenium-webdriverweb-scrapingpython

Top 5 Answer for Fetch all href link using selenium in python

vote vote

100

Well, you have to simply loop through the list:

elems = driver.find_elements_by_xpath("//a[@href]") for elem in elems:     print(elem.get_attribute("href")) 

find_elements_by_* returns a list of elements (note the spelling of 'elements'). Loop through the list, take each element and fetch the required attribute value you want from it (in this case href).

vote vote

85

I have checked and tested that there is a function named find_elements_by_tag_name() you can use. This example works fine for me.

elems = driver.find_elements_by_tag_name('a')     for elem in elems:         href = elem.get_attribute('href')         if href is not None:             print(href) 
vote vote

78

You can try something like:

    links = driver.find_elements_by_partial_link_text('') 
vote vote

63

driver.get(URL) time.sleep(7) elems = driver.find_elements_by_xpath("//a[@href]") for elem in elems:     print(elem.get_attribute("href")) driver.close() 

Note: Adding delay is very important. First run it in debug mode and Make sure your URL page is getting loaded. If the page is loading slowly, increase delay (sleep time) and then extract.

If you still face any issues, please refer below link (explained with an example) or comment

Extract links from webpage using selenium webdriver

vote vote

57

You can import the HTML dom using html dom library in python. You can find it over here and install it using PIP:

https://pypi.python.org/pypi/htmldom/2.0

from htmldom import htmldom dom = htmldom.HtmlDom("https://www.github.com/")   dom = dom.createDom() 

The above code creates a HtmlDom object.The HtmlDom takes a default parameter, the url of the page. Once the dom object is created, you need to call "createDom" method of HtmlDom. This will parse the html data and constructs the parse tree which then can be used for searching and manipulating the html data. The only restriction the library imposes is that the data whether it is html or xml must have a root element.

You can query the elements using the "find" method of HtmlDom object:

p_links = dom.find("a")   for link in p_links:   print ("URL: " +link.attr("href")) 

The above code will print all the links/urls present on the web page

Top 3 video Explaining Fetch all href link using selenium in python

Related QUESTION?