Processing HTML document with C#

Question

I have a few hundred static HTML files that need to be processed.

They contain links like this

 <a href="http://www.mysite.com/">Link</a>

I need to add ?ref=self to any url that begins with http://www.mysite.com and becomes

<a href="http://www.mysite.com/?ref=self">Link</a>

however, I do not know whether it's going to be http://www.mysite.com or http://www.mysite.com/ also it could be linked to a sub directory.

What's the most efficient way to do this? in C#

I asked myself the same question and gave your question a upvote. — jgauffin, Aug 22 '10 at 06:20

score 1 · Answer 1 · edited May 23 '17 at 12:07

1

Parsing HTML can be tricky as HTML often contains poorly formed tags and attributes. I suggest looking into an existing HTML parsing library to do your heavy lifting, or, using XSLT to transform valid (x)HTML to your desired output.

This question What is the best way to parse html in C#? has some good links to HTML parsing libraries for C#.

edited May 23 '17 at 12:07

Community

1
1

answered Aug 22 '10 at 05:39

jscharf

5,829
3
24
16

A html parsing library is like taking a cannon to a duck hunt in this case. – jgauffin Aug 22 '10 at 06:20
@jgauffin, I don't see how. It's definitely an appropriate solution. – strager Aug 22 '10 at 08:28
Because the URI's are quite easy to find and replace in this case. – jgauffin Aug 22 '10 at 09:55

strager · Accepted Answer · 2010-08-22T08:30:21.110

1

What's the most efficient way to do this? in C#

Look for the string http://www.mysite.com.
If it doesn't exist, go to 7.
Look for the next ".
If it doesn't exist, error.
Insert ?ref=self before the ".
Go to 1.
Return.

This can be accomplished with the following regular expression substitution:

s#http://www.mysite.com[^"]*#&?ref=self#g

A nicer (more expressive) way would be to use an HTML parser and XPath.

edited Aug 22 '10 at 08:30

answered Aug 22 '10 at 05:43

strager

88,763
26
134
176

Bug: The `href` attribute could be in single quotes ☺ – Timwi Aug 22 '10 at 05:46
@Timwi, That's not a bug. The OP clearly stated what the expected input is (which didn't include `'`), and that efficiency was a factor (so they say...). – strager Aug 22 '10 at 05:58
I don’t see where he stated that. The OP clearly stated that the expected input is **HTML**. He neither stated that it is a special subset of HTML, nor did he state that his examples are exhaustive. If I hadn’t commented, he might not have realised that any of his HTML files could contain href attributes with single quotes and that your algorithm would silently skip them. – Timwi Aug 22 '10 at 12:58

score 0 · Answer 3 · answered Aug 22 '10 at 08:55

0

You could use Page.Request.UrlReferrer to detect where the request came from.

answered Aug 22 '10 at 08:55

bjhamltn

420
5
6

Processing HTML document with C#

3 Answers3