# xpath / CDATA

**URL:** <https://forums.plex.tv/t/xpath-cdata/22945>\
**Category:** Dev/API Corner\
**Tags:** plugin-dev\
**Created:** [November 4, 2012, 7:58pm UTC](https://forums.plex.tv/t/xpath-cdata/22945 "2012-11-04T19:58:41Z")\
**Posts on this page:** 7\
**Page:** 1

<div class="post-metadata">

**Author:** ![dominiqueD](https://sea1.discourse-cdn.com/plex/user_avatar/forums.plex.tv/dominiqued/32/231897_2.png) [@dominiqueD](https://forums.plex.tv/u/dominiqueD)\
**Post date:** [November 4, 2012, 7:58pm UTC](https://forums.plex.tv/t/xpath-cdata/22945/1 "2012-11-04T19:58:41Z")

</div>

trouble with cdata
Hi all,  
  
I'm trying to setup a simple plugins and I've got trouble with CDATA.  
Information is provided by an XMLfile, some comments are not available.  
  

```auto

	URL="http://www.europe1.fr/podcasts/revue-de-presque.xml"<br />
	data=HTML.ElementFromURL(URL,encoding="utf-8")<br />
	item=0<br />
	for item in data.xpath('//item'):<br />
		title	= item.xpath('title')[0].text<br />
		pudDate	= item.xpath('pubdate')[0].text<br />
		url = item.xpath('enclosure')[0].get('url')<br />
		summary = str(item.xpath('summary')[0].text)<br />
		print summary<br />
	return dir

```

  
  
Sometime item.xpath('summary')[0].text \> return "None" instead of content.  
I'm not sure, but i think that is a trouble with UTF-8 (with french éàè ).  
  
You can find the full code here: https://github.com/whoo/Cantelou.bundle.git  
  
Have you got some clue to get the full summary inside CDATA ?  
Thanks ;)

---

<div class="post-metadata">

**Author:** ![mikedm139](https://sea1.discourse-cdn.com/plex/user_avatar/forums.plex.tv/mikedm139/32/96909_2.png) [@mikedm139](https://forums.plex.tv/u/mikedm139)\
**Post date:** [November 4, 2012, 8:15pm UTC](https://forums.plex.tv/t/xpath-cdata/22945/2 "2012-11-04T20:15:25Z")

</div>

It’s less likely that the CDATA is the problem, than the HTML tags inside the CDATA. Try using:

```auto

<br />
summary = item.xpath('./summary/text()')<br />

```

  
  
  
It should return a list of strings which are separated by tags in the XML. You can put the string back together again using pythons .join(), like so:  
  
  

```auto

<br />
summary = ''.join(item.xpath('./summary/text()'))<br />

```

---

<div class="post-metadata">

**Author:** ![dominiqueD](https://sea1.discourse-cdn.com/plex/user_avatar/forums.plex.tv/dominiqued/32/231897_2.png) [@dominiqueD](https://forums.plex.tv/u/dominiqueD)\
**Post date:** [November 4, 2012, 8:42pm UTC](https://forums.plex.tv/t/xpath-cdata/22945/3 "2012-11-04T20:42:01Z")

</div>

Thank’s for your quick answer;  
  
  
  
I’ve try both solution ☹  
  
but item.xpath(’./summary/text()’) still return nothing when there is some special char “éaè”.

---

<div class="post-metadata">

**Author:** ![sander1](https://avatars.discourse-cdn.com/v4/letter/s/4bbf92/32.png) [@sander1](https://forums.plex.tv/u/sander1)\
**Post date:** [November 4, 2012, 9:47pm UTC](https://forums.plex.tv/t/xpath-cdata/22945/4 "2012-11-04T21:47:09Z")

</div>

Hi!  
  
You could parse the XML file as XML and use ‘itunes’ namespace to get to the summary element. It’ll look something like this:

```auto

url = "http://www.europe1.fr/podcasts/revue-de-presque.xml"<br />
data = XML.ElementFromURL(url) # Parse as xml<br />
<br />
for item in data.xpath('//item'):<br />
    title = item.xpath('./title')[0].text.strip()<br />
    pubDate = item.xpath('./pubDate')[0].text<br />
    url = item.xpath('./enclosure')[0].get('url')<br />
<br />
    # Use the itunes namespace to grab the summary element<br />
    summary = item.xpath('./itunes:summary', namespaces={'itunes':'http://www.itunes.com/dtds/podcast-1.0.dtd'})[0].text<br />
    # Strip out HTML tags<br />
    summary = String.StripTags(summary).strip()

```

---

<div class="post-metadata">

**Author:** ![dominiqueD](https://sea1.discourse-cdn.com/plex/user_avatar/forums.plex.tv/dominiqued/32/231897_2.png) [@dominiqueD](https://forums.plex.tv/u/dominiqueD)\
**Post date:** [November 4, 2012, 10:10pm UTC](https://forums.plex.tv/t/xpath-cdata/22945/5 "2012-11-04T22:10:01Z")

</div>

Thank’s for your answer.  
  
  
  
It’s working fine now.  
  
I’ve changed:

```auto

	<<data=HTML.ElementFromURL(URL,encoding=None)<br />
	>>data=XML.ElementFromURL(URL,encoding=None)

```

  
  
And namespace to use specials tags.  
  

```auto

	for item in data.xpath('//item'):<br />
		summary= item.xpath('t:summary',namespaces={'t':'http://www.itunes.com/dtds/podcast-1.0.dtd'})[0].text<br />
		keyword= item.xpath('t:keywords',namespaces={'t':'http://www.itunes.com/dtds/podcast-1.0.dtd'})[0].text<br />
		pubDate=item.xpath('pubDate')[0].text<br />
		url= item.xpath('enclosure')[0].get('url')<br />
		title=item.xpath('title')[0].text.strip()<br />
		summary="[%s]

%s 
%s
keywords:%s "%(title,summary.strip(),pubDate,keyword)<br />
		title=title.strip()<br />
		dir.Append(TrackItem(url,title,"info","Rubrique",summary=summary,art=R(ICON)))<br />
	return dir

```

  
  
My Code is available on github:  
https://github.com/whoo/Cantelou.bundle.git

---

<div class="post-metadata">

**Author:** ![mikedm139](https://sea1.discourse-cdn.com/plex/user_avatar/forums.plex.tv/mikedm139/32/96909_2.png) [@mikedm139](https://forums.plex.tv/u/mikedm139)\
**Post date:** [November 14, 2012, 10:53pm UTC](https://forums.plex.tv/t/xpath-cdata/22945/6 "2012-11-14T22:53:20Z")

</div>

> [@](#):
>
> Thank's for your answer.  
>   
> It's working fine now.  
> I've changed:   
> 
> ```auto
> 
> <<data=HTML.ElementFromURL(URL,encoding=None)<br />
> >>data=XML.ElementFromURL(URL,encoding=None)
> 
> ```
> 
>   
>   
> And namespace to use specials tags.  
>   
> 
> ```auto
> 
> for item in data.xpath('//item'):<br />
> summary= item.xpath('t:summary',namespaces={'t':'http://www.itunes.com/dtds/podcast-1.0.dtd'})[0].text<br />
> keyword= item.xpath('t:keywords',namespaces={'t':'http://www.itunes.com/dtds/podcast-1.0.dtd'})[0].text<br />
> pubDate=item.xpath('pubDate')[0].text<br />
> url= item.xpath('enclosure')[0].get('url')<br />
> title=item.xpath('title')[0].text.strip()<br />
> summary="[%s]
> 
> %s 
> %s
> keywords:%s "%(title,summary.strip(),pubDate,keyword)<br />
> title=title.strip()<br />
> dir.Append(TrackItem(url,title,"info","Rubrique",summary=summary,art=R(ICON)))<br />
> return dir
> 
> ```
> 
>   
>   
> My Code is available on github:  
> [https://github.com/w...elou.bundle.git](https://github.com/whoo/Cantelou.bundle.git)

  
  
If you're interested in making your channel available in the Plex Channel Directory, there are instructions [here](http://wiki.plexapp.com/index.php/App\_Store\_Submission) and you can file a ticket for review on the [lighthouse project.](https://plexapp.lighthouseapp.com/projects/31804-plex-plug-ins/overview)

---

<div class="post-metadata">

**Author:** ![system](https://global.discourse-cdn.com/plex/original/3X/2/a/2acb9765406f63293d357b4ec509ec39aa28f2ad.png) [@system](https://forums.plex.tv/u/system)\
**Post date:** [December 20, 2019, 10:23pm UTC](https://forums.plex.tv/t/xpath-cdata/22945/7 "2019-12-20T22:23:32Z")

</div>

This topic was automatically closed 90 days after the last reply. New replies are no longer allowed.
