# Parsing HTML with xpath

**URL:** <https://forums.plex.tv/t/parsing-html-with-xpath/5413>\
**Category:** Dev/API Corner\
**Tags:** scanner-agent-dev\
**Created:** [September 18, 2010, 7:47pm UTC](https://forums.plex.tv/t/parsing-html-with-xpath/5413 "2010-09-18T19:47:53Z")\
**Posts on this page:** 8\
**Page:** 1

<div class="post-metadata">

**Author:** ![ptath](https://avatars.discourse-cdn.com/v4/letter/p/df788c/32.png) [@ptath](https://forums.plex.tv/u/ptath)\
**Post date:** [September 18, 2010, 7:47pm UTC](https://forums.plex.tv/t/parsing-html-with-xpath/5413/1 "2010-09-18T19:47:53Z")

</div>

noob questions
Finished [this](http://forums.plexapp.com/index.php?/topic/18002-plugin-for-non-imdb-site/), and now trying to parse html page for metadata.  
  
This ((http://www.kinopoisk.ru/level/1/film/251733/) for example) works good  
  

```auto

kinopoiskHtml = HTML.ElementFromURL(kinopoiskUrl)<br />
metadata.summary = str(kinopoiskHtml.xpath("//span[@class='_reachbanner_']")[0].text)<br />
metadata.tagline = str(kinopoiskHtml.xpath("//td[@style='color: #555']")[0].text)<br />

```

  
  
But I'm completely noob with other cases, especially when regex needed. Cannot find good examples how to parse something like this:  
  

```auto

<br />
<tr><td class="type">год</td><td class=""><a href="/level/10/m_act%5Byear%5D/2009/">2009</a></td></tr><br />
<tr><td class="type">жанр</td><td><a href="/level/10/m_act%5Bgenre%5D/2/">фантастика</a>, <a href="/level/10/m_act%5Bgenre%5D/3/">боевик</a>, <a href="/level/10/m_act%5Bgenre%5D/8/">драма</a>, <a href="/level/10/m_act%5Bgenre%5D/10/">приключения</a>, <a href="/level/92/film/251733/">...</a></td></tr><br />

```

  
  
Year is inside \*\* tag and is a part of \*href\* parameter. Please help me with xpath.

---

<div class="post-metadata">

**Author:** ![sander1](https://avatars.discourse-cdn.com/v4/letter/s/4bbf92/32.png) [@sander1](https://forums.plex.tv/u/sander1)\
**Post date:** [September 18, 2010, 7:59pm UTC](https://forums.plex.tv/t/parsing-html-with-xpath/5413/2 "2010-09-18T19:59:01Z")

</div>

Hi! You probably don’t need regex in this case (pfew ;)). Using the [_contains_ function](http://www.w3schools.com/Xpath/xpath_functions.asp) with your xpath can help you find the right a tag, like so:

```auto

kinopoiskHtml = HTML.ElementFromURL(kinopoiskUrl)<br />
...<br />
...<br />
year = int(kinopoiskHtml.xpath('//a[contains(@href, "year")]')[0].text)<br />

```

  
This searches for the string "year" inside all href attributes of a tags.

---

<div class="post-metadata">

**Author:** ![ptath](https://avatars.discourse-cdn.com/v4/letter/p/df788c/32.png) [@ptath](https://forums.plex.tv/u/ptath)\
**Post date:** [September 18, 2010, 8:05pm UTC](https://forums.plex.tv/t/parsing-html-with-xpath/5413/3 "2010-09-18T20:05:34Z")

</div>

> [@](#):
>
> Hi! You probably don't need regex in this case (pfew ;)). Using the [\*contains\* function](http://www.w3schools.com/Xpath/xpath\_functions.asp) with your xpath can help you find the right a tag, like so:  
> 
> ```auto
> 
> kinopoiskHtml = HTML.ElementFromURL(kinopoiskUrl)<br />
> ...<br />
> ...<br />
> year = int(kinopoiskHtml.xpath('//a[contains(@href, "year")]')[0].text)<br />
> 
> ```
> 
>   
> This searches for the string "year" inside all href attributes of a tags.

  
  
Oh thank you, it works =)  
  
Any hint how to deal with lists (genres, actors etc like in html code above)?

---

<div class="post-metadata">

**Author:** ![sander1](https://avatars.discourse-cdn.com/v4/letter/s/4bbf92/32.png) [@sander1](https://forums.plex.tv/u/sander1)\
**Post date:** [September 18, 2010, 8:16pm UTC](https://forums.plex.tv/t/parsing-html-with-xpath/5413/4 "2010-09-18T20:16:15Z")

</div>

> [@](#):
>
> Any hint how to deal with lists (genres, actors etc like in html code above)?

  
  
I haven't worked with "genres" yet, but by looking at the Cine-Passion agent, it should be something like this:  

```auto

<br />
metadata.genres.clear()<br />
genres = kinopoiskHtml.xpath('//a[contains(@href, "genre")]')<br />
<br />
for genre in genres:<br />
  metadata.genres.add( genre.text.strip() )<br />

```

---

<div class="post-metadata">

**Author:** ![ptath](https://avatars.discourse-cdn.com/v4/letter/p/df788c/32.png) [@ptath](https://forums.plex.tv/u/ptath)\
**Post date:** [September 19, 2010, 10:20am UTC](https://forums.plex.tv/t/parsing-html-with-xpath/5413/5 "2010-09-19T10:20:47Z")

</div>

> [@](#):
>
> I haven't worked with "genres" yet, but by looking at the Cine-Passion agent, it should be something like this:  
> 
> ```auto
> 
> <br />
> metadata.genres.clear()<br />
> genres = kinopoiskHtml.xpath('//a[contains(@href, "genre")]')<br />
> <br />
> for genre in genres:<br />
> metadata.genres.add( genre.text.strip() )<br />
> 
> ```

  
  
Thank you, it works, but show only first genre. Cinepassion agent is good for this examples.  
  
Still problem with directors, actors and so on, there no info inside tag:  
   
 

```auto

<tr><td class="type">режиссер</td><td><a href="/level/4/people/27977/">Джеймс Кэмерон</a></td></tr><br />
<tr><td class="type">DIRECTOR</td><td><a href="/level/4/people/27977/">JAMES CAMERON</a></td></tr>

```

   
   
Any ideas please?

---

<div class="post-metadata">

**Author:** ![hrcolb0](https://sea1.discourse-cdn.com/plex/user_avatar/forums.plex.tv/hrcolb0/32/380_2.png) [@hrcolb0](https://forums.plex.tv/u/hrcolb0)\
**Post date:** [September 19, 2010, 12:15pm UTC](https://forums.plex.tv/t/parsing-html-with-xpath/5413/6 "2010-09-19T12:15:50Z")

</div>

> [@](#):
>
> Thank you, it works, but show only first genre. Cinepassion agent is good for this examples.  
>   
> Still problem with directors, actors and so on, there no info inside tag:  
>    
>  
> 
> ```auto
> 
> <tr><td class="type">режиссер</td><td><a href="/level/4/people/27977/">Джеймс Кэмерон</a></td></tr><br />
> <tr><td class="type">DIRECTOR</td><td><a href="/level/4/people/27977/">JAMES CAMERON</a></td></tr>
> 
> ```
> 
>    
>    
> Any ideas please?

   
Use .get('href') for the link and .text for the Russian. Is thatxbmc nfo file?

---

<div class="post-metadata">

**Author:** ![sander1](https://avatars.discourse-cdn.com/v4/letter/s/4bbf92/32.png) [@sander1](https://forums.plex.tv/u/sander1)\
**Post date:** [September 19, 2010, 2:46pm UTC](https://forums.plex.tv/t/parsing-html-with-xpath/5413/7 "2010-09-19T14:46:23Z")

</div>

> [@](#):
>
> Still problem with directors, actors and so on, there no info inside tag:  
>    
>  
> 
> ```auto
> 
> <tr><td class="type">режиссер</td><td><a href="/level/4/people/27977/">Джеймс Кэмерон</a></td></tr><br />
> <tr><td class="type">DIRECTOR</td><td><a href="/level/4/people/27977/">JAMES CAMERON</a></td></tr>
> 
> ```
> 
>    
>    
> Any ideas please?

   
   
Find the \*td\* tag that contains a text node with value "DIRECTOR", get its parent node \*tr\* and find all \*a\* tags that are descendants of this node that contain the string "people" inside their href attribute:  
 

```auto

//td[text()="DIRECTOR"]/parent::tr//a[contains(@href,"people")]

```

---

<div class="post-metadata">

**Author:** ![system](https://global.discourse-cdn.com/plex/original/3X/2/a/2acb9765406f63293d357b4ec509ec39aa28f2ad.png) [@system](https://forums.plex.tv/u/system)\
**Post date:** [December 20, 2019, 8:46pm UTC](https://forums.plex.tv/t/parsing-html-with-xpath/5413/8 "2019-12-20T20:46:31Z")

</div>

This topic was automatically closed 90 days after the last reply. New replies are no longer allowed.
