# parsing javascript with xpath

**URL:** <https://forums.plex.tv/t/parsing-javascript-with-xpath/3272>\
**Category:** Dev/API Corner\
**Tags:** plugin-dev\
**Created:** [December 31, 2009, 12:23am UTC](https://forums.plex.tv/t/parsing-javascript-with-xpath/3272 "2009-12-31T00:23:13Z")\
**Posts on this page:** 13\
**Page:** 1

<div class="post-metadata">

**Author:** ![David\_Veld](https://avatars.discourse-cdn.com/v4/letter/d/858c86/32.png) [@David\_Veld](https://forums.plex.tv/u/David_Veld)\
**Post date:** [December 31, 2009, 12:23am UTC](https://forums.plex.tv/t/parsing-javascript-with-xpath/3272/1 "2009-12-31T00:23:13Z")

</div>

Hi,  
  
  
  
Is it possible in plex parsing a code that was generated with javascript using xpath.  
  
  
  
For example, in the script below I want only to extract the string value of sGlobalFileName=‘what-i-think-of-tv-news’, so as a result I want to have ‘what-i-think-of-tv-news’ in the end. Is this possible using xpath? If yes, How?

---

<div class="post-metadata">

**Author:** ![jwray](https://sea1.discourse-cdn.com/plex/user_avatar/forums.plex.tv/jwray/32/99832_2.png) [@jwray](https://forums.plex.tv/u/jwray)\
**Post date:** [December 31, 2009, 12:48am UTC](https://forums.plex.tv/t/parsing-javascript-with-xpath/3272/2 "2009-12-31T00:48:10Z")

</div>

Short answer. No. Xpath is for parsing XML documents or fragments, no general text.  
  
  
  
For that, extract the text of the script tag (that you can do using xpath) then use either regular expression or python substring to extract out the part you need. Messy, but no other way around it.

> [@](#):
>
> Hi,  
>   
> Is it possible in plex parsing a code that was generated with javascript using xpath.  
>   
> For example, in the script below I want only to extract the string value of sGlobalFileName='what-i-think-of-tv-news', so as a result I want to have 'what-i-think-of-tv-news' in the end. Is this possible using xpath? If yes, How?

---

<div class="post-metadata">

**Author:** ![David\_Veld](https://avatars.discourse-cdn.com/v4/letter/d/858c86/32.png) [@David\_Veld](https://forums.plex.tv/u/David_Veld)\
**Post date:** [December 31, 2009, 12:58am UTC](https://forums.plex.tv/t/parsing-javascript-with-xpath/3272/3 "2009-12-31T00:58:24Z")

</div>

mmm could anyone help me how that would like for a bit, I am also searching the internet for regex but it is not very clear to me…

---

<div class="post-metadata">

**Author:** ![jwray](https://sea1.discourse-cdn.com/plex/user_avatar/forums.plex.tv/jwray/32/99832_2.png) [@jwray](https://forums.plex.tv/u/jwray)\
**Post date:** [December 31, 2009, 1:23am UTC](https://forums.plex.tv/t/parsing-javascript-with-xpath/3272/4 "2009-12-31T01:23:47Z")

</div>

If you’ve never used regular expressions before now is probably not the time to learn. They are the spawn of the devil (but very powerful).  
  
  
  
You need a good python reference. String objects (and all others) have a number of very useful methods on them. You need to use find. Something like (pseudocode)

```auto

<br />
start = text.find('sGlobalFileName=') + 17<br />
end = text.find(";", start)<br />
substring = text[start:end]<br />

```

  
  
  

> [@](#):
>
> mmm could anyone help me how that would like for a bit, I am also searching the internet for regex but it is not very clear to me..

---

<div class="post-metadata">

**Author:** ![sansnipple](https://avatars.discourse-cdn.com/v4/letter/s/9d8465/32.png) [@sansnipple](https://forums.plex.tv/u/sansnipple)\
**Post date:** [December 31, 2009, 5:17am UTC](https://forums.plex.tv/t/parsing-javascript-with-xpath/3272/5 "2009-12-31T05:17:39Z")

</div>

honestly, that js example looks like a primo candidate for rolling into a nice neat JSON object. as for exactly _how_ to do that, someone with more braincells than me will have to take a look.

---

<div class="post-metadata">

**Author:** ![sander1](https://avatars.discourse-cdn.com/v4/letter/s/4bbf92/32.png) [@sander1](https://forums.plex.tv/u/sander1)\
**Post date:** [December 31, 2009, 7:27pm UTC](https://forums.plex.tv/t/parsing-javascript-with-xpath/3272/6 "2009-12-31T19:27:15Z")

</div>

Although regexes are a bit more difficult, I think they are the best way to extract the data you want (in this case).

```auto

<br />
import re<br />
<br />
webpage_content = HTTP.Request('http://www.example.com/pagecontainingthejavascript.html')<br />
title = re.search("sGlobalFileName='(.+?)';", webpage_content).group(1)<br />

```

  
  
A little explanation about the regex:  
(...) = a group within a regex, you can have multiple groups within one regex, you retrieve them with the group function. The above expression could also be written like this:  

```auto

result = re.search("sGlobalFileName='(.+?)';", webpage_content)<br />
title = result.group(1)

```

  
. = matches any character except a newline  
+ = match 1 or more repetitions of the preceding expression  
? = make the expression ungreedy (= grab as few characters as possible)  
  
Without the "?" the result of this regular expression would be:  
what-i-think-of-tv-news';EmbedSEOLinkURL='http://www.break.com/';EmbedSEOLinkKeywords='Funny Videos';sGlobalContentID='355403';sGlobalContentTitle=document.getElementById("vid\_title").getAttribute("content");sGlobalCategoryID='7';sGlobalContentFilePath='2007/8';sGlobalContentUrl='http://www.break.com/index/what-i-think-of-tv-news.html';sGlobalContentIDEncoded='boEGsDRmcc5fL4GHC%2bhfyA%3d%3d';sSubmittedBY='rbbtcee';sKeywordTitle='Flashes,man,News,reporter,What I Think of TV News';sKeywordString='Flashes,man,News,reporter

---

<div class="post-metadata">

**Author:** ![David\_Veld](https://avatars.discourse-cdn.com/v4/letter/d/858c86/32.png) [@David\_Veld](https://forums.plex.tv/u/David_Veld)\
**Post date:** [January 1, 2010, 12:35pm UTC](https://forums.plex.tv/t/parsing-javascript-with-xpath/3272/7 "2010-01-01T12:35:19Z")

</div>

Wow thanks, This really helped me! Thanks for the explanation.

---

<div class="post-metadata">

**Author:** ![David\_Veld](https://avatars.discourse-cdn.com/v4/letter/d/858c86/32.png) [@David\_Veld](https://forums.plex.tv/u/David_Veld)\
**Post date:** [January 1, 2010, 1:26pm UTC](https://forums.plex.tv/t/parsing-javascript-with-xpath/3272/8 "2010-01-01T13:26:13Z")

</div>

I created the following code using regex but there must be something that I forgot because it does not work.  
  
  
  
Does anyone have any suggestions?

```auto

def Video(sender, url):<br />
<br />
  dir = MediaContainer(title3=sender.itemTitle, art=R(ART), viewGroup="InfoList")<br />
<br />
  videos = XML.ElementFromURL(url, isHTML=True, errors='ignore').xpath(XPATH_VIDEOS)<br />
  for content in videos:<br />
    title = content.xpath("./a/span")[0].text<br />
    thumb = content.xpath("./a/img")[0].get('src')<br />
    summary = content.xpath("./a")[0].get('title')<br />
    url = content.xpath("./a")[0].get('href')<br />
    dir.Append(Function(VideoItem(PlayVideo, title=title, summary=summary, thumb=thumb), url=url))<br />
<br />
  return dir<br />
<br />
####################################################################################################<br />
<br />
def PlayVideo(sender, url):<br />
	<br />
	video_link = HTTP.Request(url)<br />
	file_name_link = re.search("sGlobalFileName='(.+?)';", video_link)<br />
	file_path_link = re.search("sGlobalContentFilePath='(.+?)';", video_link)<br />
	file_path = file_path_link.group(1)<br />
	file_name = file_name_link.group(1)<br />
	<br />
	total_video_link = 'http://media1.break.com/dnet/media/' + file_path '/' + file_name '.flv'<br />
	<br />
	dir.Append(VideoItem(total_video_link))

```

---

<div class="post-metadata">

**Author:** ![sansnipple](https://avatars.discourse-cdn.com/v4/letter/s/9d8465/32.png) [@sansnipple](https://forums.plex.tv/u/sansnipple)\
**Post date:** [January 1, 2010, 5:23pm UTC](https://forums.plex.tv/t/parsing-javascript-with-xpath/3272/9 "2010-01-01T17:23:46Z")

</div>

see my response to the other topic you started.  
  
  
  
[http://forums.plexapp.com/index.php?/topic/11941-play-video-using-regex/page\_\_view\_\_findpost\_\_p\_\_70717](http://forums.plexapp.com/index.php?/topic/11941-play-video-using-regex/page __view__ findpost __p__ 70717)

---

<div class="post-metadata">

**Author:** ![sander1](https://avatars.discourse-cdn.com/v4/letter/s/4bbf92/32.png) [@sander1](https://forums.plex.tv/u/sander1)\
**Post date:** [January 1, 2010, 5:44pm UTC](https://forums.plex.tv/t/parsing-javascript-with-xpath/3272/10 "2010-01-01T17:44:33Z")

</div>

Your indentation is maybe wrong, but you also need to change the PlayVideo function.  
  
This:

```auto

dir.Append(VideoItem(total_video_link))

```

  
needs to be replaced by this:  

```auto

<br />
return Redirect(total_video_link)<br />

```

  
  
You're also missing two plusses here:  
total\_video\_link = 'http://media1.break.com/dnet/media/' + file\_path + '/' + file\_name + '.flv'

---

<div class="post-metadata">

**Author:** ![sander1](https://avatars.discourse-cdn.com/v4/letter/s/4bbf92/32.png) [@sander1](https://forums.plex.tv/u/sander1)\
**Post date:** [January 1, 2010, 6:20pm UTC](https://forums.plex.tv/t/parsing-javascript-with-xpath/3272/11 "2010-01-01T18:20:43Z")

</div>

You can also do the regex stuff once and grab everything you need. Here is an “optimized” (shorter) PlayVideo function:

```auto

<br />
def PlayVideo(sender, url):<br />
<br />
  video_link = HTTP.Request(url)<br />
  link = re.search("sGlobalFileName='(.+?)';.+sGlobalContentFilePath='(.+?)';", video_link, re.DOTALL)<br />
  total_video_link = 'http://media1.break.com/dnet/media/' + link.group(1) + '/' + link.group(2) + '.flv'<br />
<br />
  return Redirect(total_video_link)<br />

```

---

<div class="post-metadata">

**Author:** ![David\_Veld](https://avatars.discourse-cdn.com/v4/letter/d/858c86/32.png) [@David\_Veld](https://forums.plex.tv/u/David_Veld)\
**Post date:** [January 1, 2010, 11:16pm UTC](https://forums.plex.tv/t/parsing-javascript-with-xpath/3272/12 "2010-01-01T23:16:59Z")

</div>

Wow… I feel so stupid. I just left the computer for a couple of hours alone and tried to get my concentration back and now I returned to it and read your messages it all sees so clear.  
  
  
  
Thank you for your support!

---

<div class="post-metadata">

**Author:** ![system](https://global.discourse-cdn.com/plex/original/3X/2/a/2acb9765406f63293d357b4ec509ec39aa28f2ad.png) [@system](https://forums.plex.tv/u/system)\
**Post date:** [December 20, 2019, 8:34pm UTC](https://forums.plex.tv/t/parsing-javascript-with-xpath/3272/13 "2019-12-20T20:34:43Z")

</div>

This topic was automatically closed 90 days after the last reply. New replies are no longer allowed.
