what is the best html parser for java?

Question

Assuming we have to use java, what is the best html parser that is flexible to parse lots of different html content, and also requires not a whole lot of code to do complex types of parses?

score 11 · Accepted Answer · edited May 23 '17 at 10:30

I would recommend Jsoup for this. It has a very nice API with support for jQuery like CSS selectors and non-verbose element iteration. To take a copy of this answer as an example, this prints your own question and the name of all answerers here:

URL url = new URL("https://stackoverflow.com/questions/3121136");
Document document = Jsoup.parse(url, 3000);

String question = document.select("#question .post-text").text();
System.out.println("Question: " + question);

Elements answerers = document.select("#answers .user-details a");
for (Element answerer : answerers) {
    System.out.println("Answerer: " + answerer.text());
}

An alternative would be XPath, but JSoup is more useful for webdevelopers who already have a good grasp on CSS selectors.

score 2 · Answer 2 · answered Jun 25 '10 at 20:19

2

The best would be the one that gets the job done right.

There is a opensource one called tagsoup, and also jTidy

answered Jun 25 '10 at 20:19

VoodooChild

9,776
8
66
99

what is the best html parser for java?

2 Answers2