[ Skip to the content ]

Institute of Formal and Applied Linguistics Wiki


[ Back to the navigation ]

This is an old revision of the document!


Table of Contents

Hindi (hi)

Hyderabad Dependency Treebank (HyDT-Hindi)

Versions

There has been no official release of the treebank yet. There have been two as-is sample releases for the purposes of the NLP tools contests in parsing Indian languages, attached to the ICON 2009 and 2010 conferences.

Obtaining and License

There is no standard distribution channel for the treebank after the ICON 2010 evaluation period. Inquire at the LTRC (ltrc (at) iiit (dot) ac (dot) in) about the possibility of getting the data. The ICON 2010 license in short:

HyDT-Hindi is being created by members of the Language Technologies Research Centre, International Institute of Information Technology, Gachibowli, Hyderabad, 500032, India.

References

Domain

Unknown.

Size

HyDT-Bangla shows dependencies between chunks, not words. The node/tree ratio is thus much lower than in other treebanks. The ICON 2009 version came with a data split into three parts: training, development and test:

Part Sentences Chunks Ratio
Training 980 6449 6.58
Development 150 811 5.41
Test 150 961 6.41
TOTAL 1280 8221 6.42

The ICON 2010 version came with a data split into three parts: training, development and test:

Part Sentences Chunks Ratio Words Ratio
Training 979 6440 6.58 10305 10.52
Development 150 812 5.41 1196 7.97
Test 150 961 6.41 1350 9.00
TOTAL 1279 8213 6.42 12851 10.04

I have counted the sentences and chunks. The number of words comes from (Husain et al., 2010). Note that the paper gives the number of training sentences as 980 (instead of 979), which is a mistake. The last training sentence has the id 980 but there is no sentence with id 418.

Apparently the training-development-test data split was more or less identical in both years, except for the minor discrepancies (number of training sentences and development chunks).

Inside

The text uses the WX encoding of Indian letters. If we know what the original script is (Bengali in this case) we can map the WX encoding to the original characters in UTF-8. WX uses English letters so if there was embedded English (or other string using Latin letters) it will probably get lost during the conversion.

The CoNLL format contains only the chunk heads. The native SSF format shows the other words in the chunk, too, but it does not capture intra-chunk dependency relations. This is an example of a multi-word chunk:

3       ((      NP      <fs af='rumAla,n,,sg,,d,0,0' head="rumAla" drel=k2:VGF name=NP3>
3.1     ekatA   QC      <fs af='eka,num,,,,,,'>
3.2     ledisa  JJ      <fs af='ledisa,unk,,,,,,'>
3.3     rumAla  NN      <fs af='rumAla,n,,sg,,d,0,0' name="rumAla">
        ))

In the CoNLL format, the CPOS column contains the chunk label (e.g. NP = noun phrase) and the POS column contains the part of speech of the chunk head.

Occasionally there are NULL nodes that do not correspond to any surface chunk or token. They represent ellided participants.

The syntactic tags (dependency relation labels) are karaka relations, i.e. deep syntactic roles according to the Pāṇinian grammar. There are separate versions of the treebank with fine-grained and coarse-grained syntactic tags.

According to (Husain et al., 2010), in the ICON 2010 version, the chunk tags, POS tags and inter-chunk dependencies (topology + tags) were annotated manually. The rest (lemma, morphosyntactic features, headword of chunk) was marked automatically.

Note: There have been cycles in the Hindi part of HyDT but no such problem occurs in the Bengali part.

Sample

The first two sentences of the ICON 2010 training data (with fine-grained syntactic tags) in the Shakti format:

<document docid="hi">
<head>
<title>  </title>			
<author>			
<firstname>  </firstname>			
<middlename>    </middlename>			
<lastname></lastname>			
</author>			
<availability format="electronic" />			
<bibl>			
</bibl>			
<bytecount>8.0K</bytecount>			
<domain name="general" />			
<creation creationdate="19/06/2007" institutename="IIIT Hyderabad">			
<creatorname>			
<lastname>Dipti</lastname>			
<middlename>			
</middlename>			
<firstname>Sharma</firstname>			
</creatorname>			
</creation>			
<distributor>CLIA Consortia, DIT</distributor>			
<edition number="1.0" />			
<encodingdesc>			
<newencoding>Unicode(UTF-8)</newencoding>			
<originalencoding>UTF-8</originalencoding>			
</encodingdesc>			
<sentencemarker marker=".">Specify Marker</sentencemarker>			
<language name="hi" writingsystem="LTR" script="Devanagari" />			
<normalization normalized="no">			
<utilityname>xxx.exe</utilityname>			
</normalization>			
<projectdesc name="ILMT" />			
<pubaddress addresstype="web">			
</pubaddress>			
<pubdate>			
<dateofpublication></dateofpublication>			
</pubdate>			
<publicationstmt type="copyrightfree">			
</publicationstmt>			
<publisher>			
<name></name>			
<url>xxx.com</url>			
</publisher>			
<pubplace place="books" />			
<wordcount>2  </wordcount>			
<caption>xuvryavahAra se biParIM bipASA Pilma mahowsava se vApasa lOta gaI bipASA govA. </caption>			
</caption>			
 
<annotated-resource name="HyDT-Hindi" version="2.0" type="dep-words" layers="morph,pos,chunk,dep-word" language="hin" date-of-release="20100823">
    <annotation-standard>
        <morph-standard name="Anncorra-morph" version="1.31" date="20080920" />
        <pos-standard name="Anncorra-pos" version="" date="20061215" />
        <chunk-standard name="Anncorra-chunk" version="" date="20061215" />
        <intrachunk-dependency-standard name="Anncorra-intrachunk-dep" version="1.0" date="" dep-tagset-granularity="5" />
        <dependency-standard name="Anncorra-dep" version="2.0" date="" dep-tagset-granularity="6" />
    </annotation-standard>
</annotated-resource>
</head>
<body>
<tb number="1" segment="no" bullet="no">
<foreign language="select" writingsystem="LTR"></foreign>
<text>
<Sentence id="1">
1	bAwa	NN	<fs af='bAwa,n,f,sg,3,d,0,0' drel='k1:ho' posn='10' name='bAwa' chunkId='NP' chunkType='head:NP'>
2	galawa	JJ	<fs af='galawa,adj,any,any,,any,,' drel='k1s:ho' posn='20' name='galawa' chunkId='JJP' chunkType='head:JJP'>
3	ho	VM	<fs af='ho,v,any,any,any,,0,0' drel='vmod:hE' stype='declarative' posn='30' voicetype='active' name='ho' chunkId='VGF' chunkType='head:VGF'>
4	wo	CC	<fs af='wo,avy,,,,,,' posn='40' name='wo' chunkId='CCP' chunkType='head:CCP'>
5	gussA	NN	<fs af='gussA,n,m,sg,3,d,0,0' drel='pof:AnA' posn='50' name='gussA' chunkId='NP2' chunkType='head:NP2'>
6	selebritija	NN	<fs af='selebritija,unk,,,,,0_ko,' drel='k4a:AnA' posn='60' vpos='vib_2_RP' name='selebritija' chunkId='NP3' chunkType='head:NP3'>
7	ko	PSP	<fs af='ko,psp,,,,,,' posn='70' drel='lwg__psp:selebritija' chunkType='child:NP3' name='ko'>
8	BI	RP	<fs af='BI,avy,,,,,,' posn='80' drel='lwg__rp:selebritija' chunkType='child:NP3' name='BI'>
9	AnA	VM	<fs af='A,v,any,any,any,d,nA,nA' drel='k1:hE' posn='90' name='AnA' chunkId='VGNN' chunkType='head:VGNN'>
10	lAjamI	JJ	<fs af='lAjamI,adj,any,any,,,,' drel='pof:hE' posn='100' name='lAjamI' chunkId='JJP2' chunkType='head:JJP2'>
11	hE	VM	<fs af='hE,v,any,sg,3,,hE,hE' drel='ccof:wo' stype='declarative' posn='110' voicetype='active' name='hE' chunkId='VGF2' chunkType='head:VGF2'>
12	.	SYM	<fs af='.,punc,,,,,,' posn='120' drel='rsym:hE' chunkType='child:VGF2' name='.'>
</Sentence>
 
 
<Sentence id="2">
1	bqhaspawivAra	NNP	<fs af='bqhaspawivAra,n,m,sg,3,o,0_ko,0' drel='k7t:hue' posn='10' vpos='vib_2' name='bqhaspawivAra' chunkId='NP' chunkType='head:NP'>
2	ko	PSP	<fs af='ko,psp,,,,,,' posn='20' drel='lwg__psp:bqhaspawivAra' chunkType='child:NP' name='ko'>
3	jZI	NNP	<fs af='jI,n,m,sg,3,o,0_meM,0' drel='k7:hue' posn='30' vpos='vib_2' name='jZI' chunkId='NP2' chunkType='head:NP2'>
4	meM	PSP	<fs af='meM,psp,,,,,,' posn='40' drel='lwg__psp:jZI' chunkType='child:NP2' name='meM'>
5	SurU	NN	<fs af='SurU,n,m,sg,3,d,0,0' drel='pof:hue' posn='50' name='SurU' chunkId='NP3' chunkType='head:NP3'>
6	hue	VM	<fs af='ho,v,m,sg,any,,eM,eM' drel='nmod__k1inv:mahowsava' posn='60' name='hue' chunkId='VGNF' chunkType='head:VGNF'>
7	��veM	XC	<fs af='��veM,n,m,sg,3,d,0,0' posn='70' drel='mod:mahowsava' chunkType='child:NP4' name='��veM'>
8	aMwarrARtrIya	XC	<fs af='aMwarrARtrIya,n,m,sg,3,d,0,0' posn='80' drel='mod:mahowsava' chunkType='child:NP4' name='aMwarrARtrIya'>
9	Pilma	XC	<fs af='Pilma,n,f,sg,3,d,0,0' posn='90' drel='mod:mahowsava' chunkType='child:NP4' name='Pilma'>
10	mahowsava	NNP	<fs af='mahowsava,n,m,sg,,o,0_kA,0' drel='r6:raMga' posn='100' vpos='vib_5' name='mahowsava' chunkId='NP4' chunkType='head:NP4'>
11	ke	PSP	<fs af='kA,psp,m,sg,,o,,' posn='110' drel='lwg__psp:mahowsava' chunkType='child:NP4' name='ke'>
12	raMga	NN	<fs af='raMga,n,m,sg,3,o,0_meM,0' drel='k7:padZA' posn='120' vpos='vib_2' name='raMga' chunkId='NP5' chunkType='head:NP5'>
13	meM	PSP	<fs af='meM,psp,,,,,,' posn='130' drel='lwg__psp:raMga' chunkType='child:NP5' name='meM2'>
14	BaMga	JJ	<fs af='BaMga,adj,any,any,,any,,' drel='pof:padZA' posn='140' name='BaMga' chunkId='JJP' chunkType='head:JJP'>
15	usa	DEM	<fs af='vaha,pn,any,sg,3,o,,' posn='150' drel='nmod__adj:samaya' chunkType='child:NP6' name='usa'>
16	samaya	NN	<fs af='samaya,n,any,sg,3,d,0,0' drel='k7t:padZA' posn='160' name='samaya' chunkId='NP6' chunkType='head:NP6'>
17	padZA	VM	<fs af='pada,v,any,any,any,,yA,yA' stype='declarative' posn='170' voicetype='active' name='padZA' chunkId='VGF' chunkType='head:VGF'>
18	jaba	PRP	<fs af='jaba,pn,,,,,,' drel='k7t:kiyA' posn='180' coref='samaya' name='jaba' chunkId='NP7' chunkType='head:NP7'>
19	vahAM	PRP	<fs af='vahAz,pn,,,,,0_para,' drel='jjmod:wEnAwa' posn='190' vpos='vib_2' name='vahAM' chunkId='NP8' chunkType='head:NP8'>
20	para	PSP	<fs af='para,psp,,,,,,' posn='200' drel='lwg__psp:vahAM' chunkType='child:NP8' name='para'>
21	wEnAwa	JJ	<fs af='wEnAwa,adj,any,any,,o,,' drel='nmod:surakRAkarmiyoM' posn='210' name='wEnAwa' chunkId='JJP2' chunkType='head:JJP2'>
22	surakRAkarmiyoM	NN	<fs af='surakRAkarmI,n,m,pl,3,o,0_ne,0' drel='k1:kiyA' posn='220' vpos='vib_2' name='surakRAkarmiyoM' chunkId='NP9' chunkType='head:NP9'>
23	ne	PSP	<fs af='ne,psp,,,,,,' posn='230' drel='lwg__psp:surakRAkarmiyoM' chunkType='child:NP9' name='ne'>
24	bOYlIvuda	NN	<fs af='bOYlIvuda,n,m,sg,3,o,0_kA,0' drel='r6:basu' posn='240' vpos='vib_2' name='bOYlIvuda' chunkId='NP10' chunkType='head:NP10'>
25	kI	PSP	<fs af='kA,psp,f,sg,,o,,' posn='250' drel='lwg__psp:bOYlIvuda' chunkType='child:NP10' name='kI'>
26	aBinewrI	NN	<fs af='aBinewrI,n,f,sg,3,o,0,0' posn='260' drel='nmod:bipASA' chunkType='child:NP11' name='aBinewrI'>
27	bipASA	NN	<fs af='bipASA,n,f,sg,3,d,0,0' posn='270' drel='nmod:basu' chunkType='child:NP11' name='bipASA'>
28	basu	NNP	<fs af='basu,n,f,sg,3,o,0_ke_sAWa,0' drel='k2:kiyA' posn='280' vpos='vib_vib_vib_4_5' name='basu' chunkId='NP11' chunkType='head:NP11'>
29	ke	PSP	<fs af='ke,psp,,,,,,' posn='290' drel='lwg__psp:basu' chunkType='child:NP11' name='ke2'>
30	sAWa	NST	<fs af='sAWa,nst,m,sg,3,d,,' posn='300' drel='lwg__psp:basu' chunkType='child:NP11' name='sAWa'>
31	xuvyarvahAra	NN	<fs af='xuvyarvahAra,n,m,sg,3,d,0,0' drel='pof:kiyA' posn='310' name='xuvyarvahAra' chunkId='NP12' chunkType='head:NP12'>
32	kiyA	VM	<fs af='kara,v,m,sg,any,,yA,yA' drel='nmod__relc:samaya' stype='declarative' posn='320' voicetype='active' name='kiyA' chunkId='VGF2' chunkType='head:VGF2'>
33	.	SYM	<fs af='.,punc,,,,,,' posn='330' drel='rsym:kiyA' chunkType='child:VGF2' name='.'>
</Sentence>

The same two sentences converted to the CoNLL format, WX characters decoded back to Devanagari in UTF-8:

1 बात बात NN n lex-bAwa|cat-n|gend-f|num-sg|pers-3|case-d|vib-0|tam-0|posn-10|name-bAwa|chunkId-NP|chunkType-head:NP 3 k1 _ _
2 गलत गलत JJ adj lex-galawa|cat-adj|gend-any|num-any|pers-|case-any|vib-|tam-|posn-20|name-galawa|chunkId-JJP|chunkType-head:JJP 3 k1s _ _
3 हो हो VM v lex-ho|cat-v|gend-any|num-any|pers-any|case-|vib-0|tam-0|stype-declarative|posn-30|voicetype-active|name-ho|chunkId-VGF|chunkType-head:VGF 11 vmod _ _
4 तो तो CC avy lex-wo|cat-avy|gend-|num-|pers-|case-|vib-|tam-|posn-40|name-wo|chunkId-CCP|chunkType-head:CCP 0 main _ _
5 गुस्सा गुस्सा NN n lex-gussA|cat-n|gend-m|num-sg|pers-3|case-d|vib-0|tam-0|posn-50|name-gussA|chunkId-NP2|chunkType-head:NP2 9 pof _ _
6 सेलेब्रिटिज सेलेब्रिटिज NN unk lex-selebritija|cat-unk|gend-|num-|pers-|case-|vib-0_ko|tam-|posn-60|vpos-vib_2_RP|name-selebritija|chunkId-NP3|chunkType-head:NP3 9 k4a _ _
7 को को PSP psp lex-ko|cat-psp|gend-|num-|pers-|case-|vib-|tam-|posn-70|chunkType-child:NP3|name-ko 6 lwg__psp _ _
8 भी भी RP avy lex-BI|cat-avy|gend-|num-|pers-|case-|vib-|tam-|posn-80|chunkType-child:NP3|name-BI 6 lwg__rp _ _
9 आना VM v lex-A|cat-v|gend-any|num-any|pers-any|case-d|vib-nA|tam-nA|posn-90|name-AnA|chunkId-VGNN|chunkType-head:VGNN 11 k1 _ _
10 लाजमी लाजमी JJ adj lex-lAjamI|cat-adj|gend-any|num-any|pers-|case-|vib-|tam-|posn-100|name-lAjamI|chunkId-JJP2|chunkType-head:JJP2 11 pof _ _
11 है है VM v lex-hE|cat-v|gend-any|num-sg|pers-3|case-|vib-hE|tam-hE|stype-declarative|posn-110|voicetype-active|name-hE|chunkId-VGF2|chunkType-head:VGF2 4 ccof _ _
12 . . SYM punc lex-.|cat-punc|gend-|num-|pers-|case-|vib-|tam-|posn-120|chunkType-child:VGF2|name-. 11 rsym _ _
1 बृहस्पतिवार बृहस्पतिवार NNP n lex-bqhaspawivAra|cat-n|gend-m|num-sg|pers-3|case-o|vib-0_ko|tam-0|posn-10|vpos-vib_2|name-bqhaspawivAra|chunkId-NP|chunkType-head:NP 6 k7t _ _
2 को को PSP psp lex-ko|cat-psp|gend-|num-|pers-|case-|vib-|tam-|posn-20|chunkType-child:NP|name-ko 1 lwg__psp _ _
3 ज़ी जी NNP n lex-jI|cat-n|gend-m|num-sg|pers-3|case-o|vib-0_meM|tam-0|posn-30|vpos-vib_2|name-jZI|chunkId-NP2|chunkType-head:NP2 6 k7 _ _
4 में में PSP psp lex-meM|cat-psp|gend-|num-|pers-|case-|vib-|tam-|posn-40|chunkType-child:NP2|name-meM 3 lwg__psp _ _
5 शुरू शुरू NN n lex-SurU|cat-n|gend-m|num-sg|pers-3|case-d|vib-0|tam-0|posn-50|name-SurU|chunkId-NP3|chunkType-head:NP3 6 pof _ _
6 हुए हो VM v lex-ho|cat-v|gend-m|num-sg|pers-any|case-|vib-eM|tam-eM|posn-60|name-hue|chunkId-VGNF|chunkType-head:VGNF 10 nmod__k1inv _ _
7 ��वें ��वें XC n lex-��veM|cat-n|gend-m|num-sg|pers-3|case-d|vib-0|tam-0|posn-70|chunkType-child:NP4|name-��veM 10 mod _ _
8 अंतर्राष्ट्रीय अंतर्राष्ट्रीय XC n lex-aMwarrARtrIya|cat-n|gend-m|num-sg|pers-3|case-d|vib-0|tam-0|posn-80|chunkType-child:NP4|name-aMwarrARtrIya 10 mod _ _
9 फिल्म फिल्म XC n lex-Pilma|cat-n|gend-f|num-sg|pers-3|case-d|vib-0|tam-0|posn-90|chunkType-child:NP4|name-Pilma 10 mod _ _
10 महोत्सव महोत्सव NNP n lex-mahowsava|cat-n|gend-m|num-sg|pers-|case-o|vib-0_kA|tam-0|posn-100|vpos-vib_5|name-mahowsava|chunkId-NP4|chunkType-head:NP4 12 r6 _ _
11 के का PSP psp lex-kA|cat-psp|gend-m|num-sg|pers-|case-o|vib-|tam-|posn-110|chunkType-child:NP4|name-ke 10 lwg__psp _ _
12 रंग रंग NN n lex-raMga|cat-n|gend-m|num-sg|pers-3|case-o|vib-0_meM|tam-0|posn-120|vpos-vib_2|name-raMga|chunkId-NP5|chunkType-head:NP5 17 k7 _ _
13 में में PSP psp lex-meM|cat-psp|gend-|num-|pers-|case-|vib-|tam-|posn-130|chunkType-child:NP5|name-meM2 12 lwg__psp _ _
14 भंग भंग JJ adj lex-BaMga|cat-adj|gend-any|num-any|pers-|case-any|vib-|tam-|posn-140|name-BaMga|chunkId-JJP|chunkType-head:JJP 17 pof _ _
15 उस वह DEM pn lex-vaha|cat-pn|gend-any|num-sg|pers-3|case-o|vib-|tam-|posn-150|chunkType-child:NP6|name-usa 16 nmod__adj _ _
16 समय समय NN n lex-samaya|cat-n|gend-any|num-sg|pers-3|case-d|vib-0|tam-0|posn-160|name-samaya|chunkId-NP6|chunkType-head:NP6 17 k7t _ _
17 पड़ा पड VM v lex-pada|cat-v|gend-any|num-any|pers-any|case-|vib-yA|tam-yA|stype-declarative|posn-170|voicetype-active|name-padZA|chunkId-VGF|chunkType-head:VGF 0 main _ _
18 जब जब PRP pn lex-jaba|cat-pn|gend-|num-|pers-|case-|vib-|tam-|posn-180|coref-samaya|name-jaba|chunkId-NP7|chunkType-head:NP7 32 k7t _ _
19 वहां वहाँ PRP pn lex-vahAz|cat-pn|gend-|num-|pers-|case-|vib-0_para|tam-|posn-190|vpos-vib_2|name-vahAM|chunkId-NP8|chunkType-head:NP8 21 jjmod _ _
20 पर पर PSP psp lex-para|cat-psp|gend-|num-|pers-|case-|vib-|tam-|posn-200|chunkType-child:NP8|name-para 19 lwg__psp _ _
21 तैनात तैनात JJ adj lex-wEnAwa|cat-adj|gend-any|num-any|pers-|case-o|vib-|tam-|posn-210|name-wEnAwa|chunkId-JJP2|chunkType-head:JJP2 22 nmod _ _
22 सुरक्षाकर्मियों सुरक्षाकर्मी NN n lex-surakRAkarmI|cat-n|gend-m|num-pl|pers-3|case-o|vib-0_ne|tam-0|posn-220|vpos-vib_2|name-surakRAkarmiyoM|chunkId-NP9|chunkType-head:NP9 32 k1 _ _
23 ने ने PSP psp lex-ne|cat-psp|gend-|num-|pers-|case-|vib-|tam-|posn-230|chunkType-child:NP9|name-ne 22 lwg__psp _ _
24 बॉलीवुड बॉलीवुड NN n lex-bOYlIvuda|cat-n|gend-m|num-sg|pers-3|case-o|vib-0_kA|tam-0|posn-240|vpos-vib_2|name-bOYlIvuda|chunkId-NP10|chunkType-head:NP10 28 r6 _ _
25 की का PSP psp lex-kA|cat-psp|gend-f|num-sg|pers-|case-o|vib-|tam-|posn-250|chunkType-child:NP10|name-kI 24 lwg__psp _ _
26 अभिनेत्री अभिनेत्री NN n lex-aBinewrI|cat-n|gend-f|num-sg|pers-3|case-o|vib-0|tam-0|posn-260|chunkType-child:NP11|name-aBinewrI 27 nmod _ _
27 बिपाशा बिपाशा NN n lex-bipASA|cat-n|gend-f|num-sg|pers-3|case-d|vib-0|tam-0|posn-270|chunkType-child:NP11|name-bipASA 28 nmod _ _
28 बसु बसु NNP n lex-basu|cat-n|gend-f|num-sg|pers-3|case-o|vib-0_ke_sAWa|tam-0|posn-280|vpos-vib_vib_vib_4_5|name-basu|chunkId-NP11|chunkType-head:NP11 32 k2 _ _
29 के के PSP psp lex-ke|cat-psp|gend-|num-|pers-|case-|vib-|tam-|posn-290|chunkType-child:NP11|name-ke2 28 lwg__psp _ _
30 साथ साथ NST nst lex-sAWa|cat-nst|gend-m|num-sg|pers-3|case-d|vib-|tam-|posn-300|chunkType-child:NP11|name-sAWa 28 lwg__psp _ _
31 दुव्यर्वहार दुव्यर्वहार NN n lex-xuvyarvahAra|cat-n|gend-m|num-sg|pers-3|case-d|vib-0|tam-0|posn-310|name-xuvyarvahAra|chunkId-NP12|chunkType-head:NP12 32 pof _ _
32 किया कर VM v lex-kara|cat-v|gend-m|num-sg|pers-any|case-|vib-yA|tam-yA|stype-declarative|posn-320|voicetype-active|name-kiyA|chunkId-VGF2|chunkType-head:VGF2 16 nmod__relc _ _
33 . . SYM punc lex-.|cat-punc|gend-|num-|pers-|case-|vib-|tam-|posn-330|chunkType-child:VGF2|name-. 32 rsym _ _

The first sentence of the ICON 2010 development data (with fine-grained syntactic tags) in the Shakti format:

<document docid="fullnews_id_2489467">
<head>
	<caption>jela meM svasWa hE sarabajIwa xo BArawIya aXikAriyoM ne mulAkAwa kI pre isalAmAbAxa.</caption>
	<language>Hindi </language>
	<domain_name>News Articles </domain_name>
	<word_count>524</word_count>
	<byte_count>64554</byte_count>
	<availability>
		<format>CML/SSF</format>
	<sentence_marker>.</sentence_marker>
	<normalization>No</normalization>
	</availability>
	<encoding_description>
		<original_encoding>ISO 8859</format>
		<new_encoding>Unicode UTF8</new_encoding>
	</encoding_description>
	<distributor>LTRC, IIIT Hyderabad</distributor>
	<project_description>NSF Hindi/Urdu Dependency Treebanking Project</place>
	<creation>
		</raw_corpus creation_date="" institute_name="IIIT Hyderabad">
		</annotated_corpus creation_date="06/01/2009" institute_name="IIIT Hyderabad">
	<edition_number>1.0</edition_number>
	</creation>
	<publication>
		<place>New Delhi</place>
		<date>30/5/2004</date>
		<type>Newspaper</type>
		<publisher>
			<name>Amar Ujala</name>
			<url>http://www.amarujala.com</url>
		</publisher>
	</publication>
 
<annotated-resource name="HyDT-Hindi" version="2.0" type="dep-words" layers="morph,pos,chunk,dep-word" language="hin" date-of-release="20100831">
    <annotation-standard>
        <morph-standard name="Anncorra-morph" version="1.31" date="20080920" />
        <pos-standard name="Anncorra-pos" version="" date="20061215" />
        <chunk-standard name="Anncorra-chunk" version="" date="20061215" />
        <intrachunk-dependency-standard name="Anncorra-intrachunk-dep" version="1.0" date="" dep-tagset-granularity="5" />
        <dependency-standard name="Anncorra-dep" version="2.0" date="" dep-tagset-granularity="6" />
    </annotation-standard>
</annotated-resource>
</head>
<body>
<tb number="1" segment="no" bullet="no">
<foreign language="select" writingsystem="LTR"></foreign>
<text>
<Sentence id="1">
1	kota	XC	<fs af='kota,n,m,sg,3,d,0,0' posn='10' drel='mod:lAhOra' chunkType='child:NP' name='kota'>
2	laKapawa	XC	<fs af='laKapawa,n,m,sg,3,d,0,0' posn='20' drel='mod:lAhOra' chunkType='child:NP' name='laKapawa'>
3	jela	XC	<fs af='jela,n,m,sg,3,d,0,0' posn='30' drel='mod:lAhOra' chunkType='child:NP' name='jela'>
4	lAhOra	NNP	<fs af='lAhOra,n,m,sg,3,o,0_meM,0' drel='jjmod:baMxa' posn='40' vpos='vib_5' name='lAhOra' chunkId='NP' chunkType='head:NP'>
5	meM	PSP	<fs af='meM,psp,,,,,,' posn='50' drel='lwg__psp:lAhOra' chunkType='child:NP' name='meM'>
6	baMxa	JJ	<fs af='baMxa,adj,any,any,,o,,' drel='nmod:siMha' posn='60' name='baMxa' chunkId='JJP' chunkType='head:JJP'>
7	sarabajIwa	XC	<fs af='sarabajIwa,n,m,sg,3,d,0,0' posn='70' drel='mod:siMha' chunkType='child:NP2' name='sarabajIwa'>
8	siMha	NNP	<fs af='siMha,n,m,sg,3,o,0_ne,0' drel='k1:xIM' posn='80' vpos='vib_3' name='siMha' chunkId='NP2' chunkType='head:NP2'>
9	ne	PSP	<fs af='ne,psp,,,,,,' posn='90' drel='lwg__psp:siMha' chunkType='child:NP2' name='ne'>
10	maMgalavAra	NNP	<fs af='maMgalavAra,n,m,sg,3,o,0_ko,0' drel='k7t:xIM' posn='100' vpos='vib_2' name='maMgalavAra' chunkId='NP3' chunkType='head:NP3'>
11	ko	PSP	<fs af='ko,psp,,,,,,' posn='110' drel='lwg__psp:maMgalavAra' chunkType='child:NP3' name='ko'>
12	BArawIya	JJ	<fs af='BArawIya,adj,any,any,,o,,' posn='120' drel='nmod__adj:xUwAvAsa' chunkType='child:NP4' name='BArawIya'>
13	xUwAvAsa	NN	<fs af='xUwAvAsa,n,m,sg,3,o,0_kA,0' drel='r6:aXikAriyoM' posn='130' vpos='vib_3' name='xUwAvAsa' chunkId='NP4' chunkType='head:NP4'>
14	ke	PSP	<fs af='kA,psp,m,pl,,o,,' posn='140' drel='lwg__psp:xUwAvAsa' chunkType='child:NP4' name='ke'>
15	xo	QC	<fs af='xo,num,any,pl,,o,,' posn='150' drel='nmod__adj:aXikAriyoM' chunkType='child:NP5' name='xo'>
16	aXikAriyoM	NN	<fs af='aXikArI,n,m,pl,3,o,0_ko,0' drel='k4:xIM' posn='160' vpos='vib_3' name='aXikAriyoM' chunkId='NP5' chunkType='head:NP5'>
17	ko	PSP	<fs af='ko,psp,,,,,,' posn='170' drel='lwg__psp:aXikAriyoM' chunkType='child:NP5' name='ko2'>
18	apane	PRP	<fs af='apanA,pn,any,sg,1,o,0_bAre_meM,0' drel='k7:xIM' posn='180' vpos='vib_2_3' name='apane' chunkId='NP6' chunkType='head:NP6'>
19	bAre	PSP	<fs af='bAre,psp,,,,,,' posn='190' drel='lwg__psp:apane' chunkType='child:NP6' name='bAre'>
20	meM	PSP	<fs af='meM,psp,,,,,,' posn='200' drel='lwg__psp:apane' chunkType='child:NP6' name='meM2'>
21	wamAma	JJ	<fs af='wamAma,adj,any,any,,d,,' posn='210' drel='nmod__adj:jAnakAriyAM' chunkType='child:NP7' name='wamAma'>
22	vyakwigawa	JJ	<fs af='vyakwigawa,adj,any,any,,d,,' posn='220' drel='nmod__adj:jAnakAriyAM' chunkType='child:NP7' name='vyakwigawa'>
23	jAnakAriyAM	NN	<fs af='jAnakAriyAM,n,f,pl,3,d,0,0' drel='k2:xIM' posn='230' name='jAnakAriyAM' chunkId='NP7' chunkType='head:NP7'>
24	xIM	VM	<fs af='xe,v,f,pl,3,,yA,yA' stype='declarative' posn='240' voicetype='active' name='xIM' chunkId='VGF' chunkType='head:VGF'>
25	ki	CC	<fs af='ki,avy,,,,,,' drel='rs:jAnakAriyAM' posn='250' name='ki' chunkId='CCP' chunkType='head:CCP'>
26	kina	WQ	<fs af='kOna,pn,any,pl,3,o,,' posn='260' drel='mod__wq:parisWiwiyoM' chunkType='child:NP8' name='kina'>
27	parisWiwiyoM	NN	<fs af='parisWiwi,n,f,pl,3,o,0_meM,0' drel='k7:kiyA' posn='270' vpos='vib_3' name='parisWiwiyoM' chunkId='NP8' chunkType='head:NP8'>
28	meM	PSP	<fs af='meM,psp,,,,,,' posn='280' drel='lwg__psp:parisWiwiyoM' chunkType='child:NP8' name='meM3'>
29	use	PRP	<fs af='vaha,pn,any,sg,3,o,ko,ko' drel='k2:kiyA' posn='290' name='use' chunkId='NP9' chunkType='head:NP9'>
30	giraPwAra	JJ	<fs af='giraPwAra,adj,any,any,,,,' drel='pof:kiyA' posn='300' name='giraPwAra' chunkId='JJP2' chunkType='head:JJP2'>
31	kiyA	VM	<fs af='kara,v,m,sg,3,,yA_jA+yA�,yA' drel='ccof:Ora' stype='declarative' posn='310' voicetype='passive' vpos='tam_2' name='kiyA' chunkId='VGF2' chunkType='head:VGF2'>
32	gayA	VAUX	<fs af='jA,v,m,sg,3,,yA�,yA1' posn='320' drel='lwg__vaux:kiyA' chunkType='child:VGF2' name='gayA'>
33	,	SYM	<fs af=',s,punc,,,,,' posn='330' drel='rsym:kiyA' chunkType='child:VGF2' name=','>
34	mukaxamA	NN	<fs af='mukaxamA,n,m,sg,3,d,0,0' drel='k1:calA' posn='340' name='mukaxamA' chunkId='NP10' chunkType='head:NP10'>
35	calA	VM	<fs af='cala,v,m,sg,3,,yA,yA' hlt='true' drel='ccof:Ora' stype='declarative' posn='350' voicetype='active' name='calA' chunkId='VGF3' chunkType='head:VGF3'>
36	Ora	CC	<fs af='Ora,avy,,,,,,' drel='ccof:ki' posn='360' name='Ora' chunkId='CCP2' chunkType='head:CCP2'>
37	sajA	NN	<fs af='sajA,n,f,sg,3,d,0,0' drel='k1:huI' posn='370' name='sajA' chunkId='NP11' chunkType='head:NP11'>
38	huI	VM	<fs af='ho,v,f,sg,3,,yA,yA' drel='ccof:Ora' stype='declarative' posn='380' voicetype='active' name='huI' chunkId='VGF4' chunkType='head:VGF4'>
39	.	SYM	<fs af='.,punc,,,,,,' posn='390' drel='rsym:huI' chunkType='child:VGF4' name='.'>
</Sentence>

And in the CoNLL format:

1 kota kota XC n lex-kota|cat-n|gend-m|num-sg|pers-3|case-d|vib-0|tam-0|posn-10|chunkType-child:NP|name-kota 4 mod _ _
2 laKapawa laKapawa XC n lex-laKapawa|cat-n|gend-m|num-sg|pers-3|case-d|vib-0|tam-0|posn-20|chunkType-child:NP|name-laKapawa 4 mod _ _
3 jela jela XC n lex-jela|cat-n|gend-m|num-sg|pers-3|case-d|vib-0|tam-0|posn-30|chunkType-child:NP|name-jela 4 mod _ _
4 lAhOra lAhOra NNP n lex-lAhOra|cat-n|gend-m|num-sg|pers-3|case-o|vib-0_meM|tam-0|posn-40|vpos-vib_5|name-lAhOra|chunkId-NP|chunkType-head:NP 6 jjmod _ _
5 meM meM PSP psp lex-meM|cat-psp|gend-|num-|pers-|case-|vib-|tam-|posn-50|chunkType-child:NP|name-meM 4 lwg__psp _ _
6 baMxa baMxa JJ adj lex-baMxa|cat-adj|gend-any|num-any|pers-|case-o|vib-|tam-|posn-60|name-baMxa|chunkId-JJP|chunkType-head:JJP 8 nmod _ _
7 sarabajIwa sarabajIwa XC n lex-sarabajIwa|cat-n|gend-m|num-sg|pers-3|case-d|vib-0|tam-0|posn-70|chunkType-child:NP2|name-sarabajIwa 8 mod _ _
8 siMha siMha NNP n lex-siMha|cat-n|gend-m|num-sg|pers-3|case-o|vib-0_ne|tam-0|posn-80|vpos-vib_3|name-siMha|chunkId-NP2|chunkType-head:NP2 24 k1 _ _
9 ne ne PSP psp lex-ne|cat-psp|gend-|num-|pers-|case-|vib-|tam-|posn-90|chunkType-child:NP2|name-ne 8 lwg__psp _ _
10 maMgalavAra maMgalavAra NNP n lex-maMgalavAra|cat-n|gend-m|num-sg|pers-3|case-o|vib-0_ko|tam-0|posn-100|vpos-vib_2|name-maMgalavAra|chunkId-NP3|chunkType-head:NP3 24 k7t _ _
11 ko ko PSP psp lex-ko|cat-psp|gend-|num-|pers-|case-|vib-|tam-|posn-110|chunkType-child:NP3|name-ko 10 lwg__psp _ _
12 BArawIya BArawIya JJ adj lex-BArawIya|cat-adj|gend-any|num-any|pers-|case-o|vib-|tam-|posn-120|chunkType-child:NP4|name-BArawIya 13 nmod__adj _ _
13 xUwAvAsa xUwAvAsa NN n lex-xUwAvAsa|cat-n|gend-m|num-sg|pers-3|case-o|vib-0_kA|tam-0|posn-130|vpos-vib_3|name-xUwAvAsa|chunkId-NP4|chunkType-head:NP4 16 r6 _ _
14 ke kA PSP psp lex-kA|cat-psp|gend-m|num-pl|pers-|case-o|vib-|tam-|posn-140|chunkType-child:NP4|name-ke 13 lwg__psp _ _
15 xo xo QC num lex-xo|cat-num|gend-any|num-pl|pers-|case-o|vib-|tam-|posn-150|chunkType-child:NP5|name-xo 16 nmod__adj _ _
16 aXikAriyoM aXikArI NN n lex-aXikArI|cat-n|gend-m|num-pl|pers-3|case-o|vib-0_ko|tam-0|posn-160|vpos-vib_3|name-aXikAriyoM|chunkId-NP5|chunkType-head:NP5 24 k4 _ _
17 ko ko PSP psp lex-ko|cat-psp|gend-|num-|pers-|case-|vib-|tam-|posn-170|chunkType-child:NP5|name-ko2 16 lwg__psp _ _
18 apane apanA PRP pn lex-apanA|cat-pn|gend-any|num-sg|pers-1|case-o|vib-0_bAre_meM|tam-0|posn-180|vpos-vib_2_3|name-apane|chunkId-NP6|chunkType-head:NP6 24 k7 _ _
19 bAre bAre PSP psp lex-bAre|cat-psp|gend-|num-|pers-|case-|vib-|tam-|posn-190|chunkType-child:NP6|name-bAre 18 lwg__psp _ _
20 meM meM PSP psp lex-meM|cat-psp|gend-|num-|pers-|case-|vib-|tam-|posn-200|chunkType-child:NP6|name-meM2 18 lwg__psp _ _
21 wamAma wamAma JJ adj lex-wamAma|cat-adj|gend-any|num-any|pers-|case-d|vib-|tam-|posn-210|chunkType-child:NP7|name-wamAma 23 nmod__adj _ _
22 vyakwigawa vyakwigawa JJ adj lex-vyakwigawa|cat-adj|gend-any|num-any|pers-|case-d|vib-|tam-|posn-220|chunkType-child:NP7|name-vyakwigawa 23 nmod__adj _ _
23 jAnakAriyAM jAnakAriyAM NN n lex-jAnakAriyAM|cat-n|gend-f|num-pl|pers-3|case-d|vib-0|tam-0|posn-230|name-jAnakAriyAM|chunkId-NP7|chunkType-head:NP7 24 k2 _ _
24 xIM xe VM v lex-xe|cat-v|gend-f|num-pl|pers-3|case-|vib-yA|tam-yA|stype-declarative|posn-240|voicetype-active|name-xIM|chunkId-VGF|chunkType-head:VGF 0 main _ _
25 ki ki CC avy lex-ki|cat-avy|gend-|num-|pers-|case-|vib-|tam-|posn-250|name-ki|chunkId-CCP|chunkType-head:CCP 23 rs _ _
26 kina kOna WQ pn lex-kOna|cat-pn|gend-any|num-pl|pers-3|case-o|vib-|tam-|posn-260|chunkType-child:NP8|name-kina 27 mod__wq _ _
27 parisWiwiyoM parisWiwi NN n lex-parisWiwi|cat-n|gend-f|num-pl|pers-3|case-o|vib-0_meM|tam-0|posn-270|vpos-vib_3|name-parisWiwiyoM|chunkId-NP8|chunkType-head:NP8 31 k7 _ _
28 meM meM PSP psp lex-meM|cat-psp|gend-|num-|pers-|case-|vib-|tam-|posn-280|chunkType-child:NP8|name-meM3 27 lwg__psp _ _
29 use vaha PRP pn lex-vaha|cat-pn|gend-any|num-sg|pers-3|case-o|vib-ko|tam-ko|posn-290|name-use|chunkId-NP9|chunkType-head:NP9 31 k2 _ _
30 giraPwAra giraPwAra JJ adj lex-giraPwAra|cat-adj|gend-any|num-any|pers-|case-|vib-|tam-|posn-300|name-giraPwAra|chunkId-JJP2|chunkType-head:JJP2 31 pof _ _
31 kiyA kara VM v lex-kara|cat-v|gend-m|num-sg|pers-3|case-|vib-yA_jA+yA�|tam-yA|stype-declarative|posn-310|voicetype-passive|vpos-tam_2|name-kiyA|chunkId-VGF2|chunkType-head:VGF2 36 ccof _ _
32 gayA jA VAUX v lex-jA|cat-v|gend-m|num-sg|pers-3|case-|vib-yA�|tam-yA1|posn-320|chunkType-child:VGF2|name-gayA 31 lwg__vaux _ _
33 , , SYM s lex-|cat-s|gend-punc|num-|pers-|case-|vib-|tam-|posn-330|chunkType-child:VGF2|name-, 31 rsym _ _
34 mukaxamA mukaxamA NN n lex-mukaxamA|cat-n|gend-m|num-sg|pers-3|case-d|vib-0|tam-0|posn-340|name-mukaxamA|chunkId-NP10|chunkType-head:NP10 35 k1 _ _
35 calA cala VM v lex-cala|cat-v|gend-m|num-sg|pers-3|case-|vib-yA|tam-yA|hlt-true|stype-declarative|posn-350|voicetype-active|name-calA|chunkId-VGF3|chunkType-head:VGF3 36 ccof _ _
36 Ora Ora CC avy lex-Ora|cat-avy|gend-|num-|pers-|case-|vib-|tam-|posn-360|name-Ora|chunkId-CCP2|chunkType-head:CCP2 25 ccof _ _
37 sajA sajA NN n lex-sajA|cat-n|gend-f|num-sg|pers-3|case-d|vib-0|tam-0|posn-370|name-sajA|chunkId-NP11|chunkType-head:NP11 38 k1 _ _
38 huI ho VM v lex-ho|cat-v|gend-f|num-sg|pers-3|case-|vib-yA|tam-yA|stype-declarative|posn-380|voicetype-active|name-huI|chunkId-VGF4|chunkType-head:VGF4 36 ccof _ _
39 . . SYM punc lex-.|cat-punc|gend-|num-|pers-|case-|vib-|tam-|posn-390|chunkType-child:VGF4|name-. 38 rsym _ _

And after conversion of the WX encoding to the Devanagari script in UTF-8:

1 कोट कोट XC n lex-kota|cat-n|gend-m|num-sg|pers-3|case-d|vib-0|tam-0|posn-10|chunkType-child:NP|name-kota 4 mod _ _
2 लखपत लखपत XC n lex-laKapawa|cat-n|gend-m|num-sg|pers-3|case-d|vib-0|tam-0|posn-20|chunkType-child:NP|name-laKapawa 4 mod _ _
3 जेल जेल XC n lex-jela|cat-n|gend-m|num-sg|pers-3|case-d|vib-0|tam-0|posn-30|chunkType-child:NP|name-jela 4 mod _ _
4 लाहौर लाहौर NNP n lex-lAhOra|cat-n|gend-m|num-sg|pers-3|case-o|vib-0_meM|tam-0|posn-40|vpos-vib_5|name-lAhOra|chunkId-NP|chunkType-head:NP 6 jjmod _ _
5 में में PSP psp lex-meM|cat-psp|gend-|num-|pers-|case-|vib-|tam-|posn-50|chunkType-child:NP|name-meM 4 lwg__psp _ _
6 बंद बंद JJ adj lex-baMxa|cat-adj|gend-any|num-any|pers-|case-o|vib-|tam-|posn-60|name-baMxa|chunkId-JJP|chunkType-head:JJP 8 nmod _ _
7 सरबजीत सरबजीत XC n lex-sarabajIwa|cat-n|gend-m|num-sg|pers-3|case-d|vib-0|tam-0|posn-70|chunkType-child:NP2|name-sarabajIwa 8 mod _ _
8 सिंह सिंह NNP n lex-siMha|cat-n|gend-m|num-sg|pers-3|case-o|vib-0_ne|tam-0|posn-80|vpos-vib_3|name-siMha|chunkId-NP2|chunkType-head:NP2 24 k1 _ _
9 ने ने PSP psp lex-ne|cat-psp|gend-|num-|pers-|case-|vib-|tam-|posn-90|chunkType-child:NP2|name-ne 8 lwg__psp _ _
10 मंगलवार मंगलवार NNP n lex-maMgalavAra|cat-n|gend-m|num-sg|pers-3|case-o|vib-0_ko|tam-0|posn-100|vpos-vib_2|name-maMgalavAra|chunkId-NP3|chunkType-head:NP3 24 k7t _ _
11 को को PSP psp lex-ko|cat-psp|gend-|num-|pers-|case-|vib-|tam-|posn-110|chunkType-child:NP3|name-ko 10 lwg__psp _ _
12 भारतीय भारतीय JJ adj lex-BArawIya|cat-adj|gend-any|num-any|pers-|case-o|vib-|tam-|posn-120|chunkType-child:NP4|name-BArawIya 13 nmod__adj _ _
13 दूतावास दूतावास NN n lex-xUwAvAsa|cat-n|gend-m|num-sg|pers-3|case-o|vib-0_kA|tam-0|posn-130|vpos-vib_3|name-xUwAvAsa|chunkId-NP4|chunkType-head:NP4 16 r6 _ _
14 के का PSP psp lex-kA|cat-psp|gend-m|num-pl|pers-|case-o|vib-|tam-|posn-140|chunkType-child:NP4|name-ke 13 lwg__psp _ _
15 दो दो QC num lex-xo|cat-num|gend-any|num-pl|pers-|case-o|vib-|tam-|posn-150|chunkType-child:NP5|name-xo 16 nmod__adj _ _
16 अधिकारियों अधिकारी NN n lex-aXikArI|cat-n|gend-m|num-pl|pers-3|case-o|vib-0_ko|tam-0|posn-160|vpos-vib_3|name-aXikAriyoM|chunkId-NP5|chunkType-head:NP5 24 k4 _ _
17 को को PSP psp lex-ko|cat-psp|gend-|num-|pers-|case-|vib-|tam-|posn-170|chunkType-child:NP5|name-ko2 16 lwg__psp _ _
18 अपने अपना PRP pn lex-apanA|cat-pn|gend-any|num-sg|pers-1|case-o|vib-0_bAre_meM|tam-0|posn-180|vpos-vib_2_3|name-apane|chunkId-NP6|chunkType-head:NP6 24 k7 _ _
19 बारे बारे PSP psp lex-bAre|cat-psp|gend-|num-|pers-|case-|vib-|tam-|posn-190|chunkType-child:NP6|name-bAre 18 lwg__psp _ _
20 में में PSP psp lex-meM|cat-psp|gend-|num-|pers-|case-|vib-|tam-|posn-200|chunkType-child:NP6|name-meM2 18 lwg__psp _ _
21 तमाम तमाम JJ adj lex-wamAma|cat-adj|gend-any|num-any|pers-|case-d|vib-|tam-|posn-210|chunkType-child:NP7|name-wamAma 23 nmod__adj _ _
22 व्यक्तिगत व्यक्तिगत JJ adj lex-vyakwigawa|cat-adj|gend-any|num-any|pers-|case-d|vib-|tam-|posn-220|chunkType-child:NP7|name-vyakwigawa 23 nmod__adj _ _
23 जानकारियां जानकारियां NN n lex-jAnakAriyAM|cat-n|gend-f|num-pl|pers-3|case-d|vib-0|tam-0|posn-230|name-jAnakAriyAM|chunkId-NP7|chunkType-head:NP7 24 k2 _ _
24 दीं दे VM v lex-xe|cat-v|gend-f|num-pl|pers-3|case-|vib-yA|tam-yA|stype-declarative|posn-240|voicetype-active|name-xIM|chunkId-VGF|chunkType-head:VGF 0 main _ _
25 कि कि CC avy lex-ki|cat-avy|gend-|num-|pers-|case-|vib-|tam-|posn-250|name-ki|chunkId-CCP|chunkType-head:CCP 23 rs _ _
26 किन कौन WQ pn lex-kOna|cat-pn|gend-any|num-pl|pers-3|case-o|vib-|tam-|posn-260|chunkType-child:NP8|name-kina 27 mod__wq _ _
27 परिस्थितियों परिस्थिति NN n lex-parisWiwi|cat-n|gend-f|num-pl|pers-3|case-o|vib-0_meM|tam-0|posn-270|vpos-vib_3|name-parisWiwiyoM|chunkId-NP8|chunkType-head:NP8 31 k7 _ _
28 में में PSP psp lex-meM|cat-psp|gend-|num-|pers-|case-|vib-|tam-|posn-280|chunkType-child:NP8|name-meM3 27 lwg__psp _ _
29 उसे वह PRP pn lex-vaha|cat-pn|gend-any|num-sg|pers-3|case-o|vib-ko|tam-ko|posn-290|name-use|chunkId-NP9|chunkType-head:NP9 31 k2 _ _
30 गिरफ्तार गिरफ्तार JJ adj lex-giraPwAra|cat-adj|gend-any|num-any|pers-|case-|vib-|tam-|posn-300|name-giraPwAra|chunkId-JJP2|chunkType-head:JJP2 31 pof _ _
31 किया कर VM v lex-kara|cat-v|gend-m|num-sg|pers-3|case-|vib-yA_jA+yA�|tam-yA|stype-declarative|posn-310|voicetype-passive|vpos-tam_2|name-kiyA|chunkId-VGF2|chunkType-head:VGF2 36 ccof _ _
32 गया जा VAUX v lex-jA|cat-v|gend-m|num-sg|pers-3|case-|vib-yA�|tam-yA1|posn-320|chunkType-child:VGF2|name-gayA 31 lwg__vaux _ _
33 , , SYM s lex-|cat-s|gend-punc|num-|pers-|case-|vib-|tam-|posn-330|chunkType-child:VGF2|name-, 31 rsym _ _
34 मुकदमा मुकदमा NN n lex-mukaxamA|cat-n|gend-m|num-sg|pers-3|case-d|vib-0|tam-0|posn-340|name-mukaxamA|chunkId-NP10|chunkType-head:NP10 35 k1 _ _
35 चला चल VM v lex-cala|cat-v|gend-m|num-sg|pers-3|case-|vib-yA|tam-yA|hlt-true|stype-declarative|posn-350|voicetype-active|name-calA|chunkId-VGF3|chunkType-head:VGF3 36 ccof _ _
36 और और CC avy lex-Ora|cat-avy|gend-|num-|pers-|case-|vib-|tam-|posn-360|name-Ora|chunkId-CCP2|chunkType-head:CCP2 25 ccof _ _
37 सजा सजा NN n lex-sajA|cat-n|gend-f|num-sg|pers-3|case-d|vib-0|tam-0|posn-370|name-sajA|chunkId-NP11|chunkType-head:NP11 38 k1 _ _
38 हुई हो VM v lex-ho|cat-v|gend-f|num-sg|pers-3|case-|vib-yA|tam-yA|stype-declarative|posn-380|voicetype-active|name-huI|chunkId-VGF4|chunkType-head:VGF4 36 ccof _ _
39 . . SYM punc lex-.|cat-punc|gend-|num-|pers-|case-|vib-|tam-|posn-390|chunkType-child:VGF4|name-. 38 rsym _ _

The first sentence of the ICON 2010 test data (with fine-grained syntactic tags) in the Shakti format:

<document id="">
<head>
<annotated-resource name="HyDT-Bangla" version="0.5" type="dep-interchunk-only" layers="morph,pos,chunk,dep-interchunk-only" language="ben" date-of-release="20101013">
    <annotation-standard>
        <morph-standard name="Anncorra-morph" version="1.31" date="20080920" />
	<pos-standard name="Anncorra-pos" version="" date="20061215" />
	<chunk-standard name="Anncorra-chunk" version="" date="20061215" />
	<dependency-standard name="Anncorra-dep" version="2.0" date="" dep-tagset-granularity="6" />
    </annotation-standard>
<annotated-resource>
</head>
<Sentence id="1">
1	((	NP	<fs af='mAXabIlawA,n,,sg,,d,0,0' head="mAXabIlawA" drel=k1:VGF name=NP>
1.1	mAXabIlawA	NNP	<fs af='mAXabIlawA,n,,sg,,d,0,0' name="mAXabIlawA">
	))		
2	((	NP	<fs af='waKana,pn,,,,d,0,0' head="waKana" drel=k7t:VGF name=NP2>
2.1	waKana	PRP	<fs af='waKana,pn,,,,d,0,0' name="waKana">
	))		
3	((	NP	<fs af='hAwa,n,,sg,,o,era,era' head="hAwera" drel=r6:NP4 name=NP3>
3.1	hAwera	NN	<fs af='hAwa,n,,sg,,o,era,era' name="hAwera">
	))		
4	((	NP	<fs af='GadZi,unk,,,,,,' head="GadZi" drel=k2:VGNF name=NP4>
4.1	GadZi	NN	<fs af='GadZi,unk,,,,,,' name="GadZi">
	))		
5	((	VGNF	<fs af='Kul,v,,,5,,ne,ne' head="Kule" drel=vmod:VGF name=VGNF>
5.1	Kule	VM	<fs af='Kul,v,,,5,,ne,ne' name="Kule">
	))		
6	((	NP	<fs af='tebila,n,,sg,,d,me,me' head="tebile" drel=k7p:VGF name=NP5>
6.1	tebile	NN	<fs af='tebila,n,,sg,,d,me,me' name="tebile">
	))		
7	((	VGF	<fs af='rAK,v,,,5,,Cila,Cila' head="rAKaCila" name=VGF>
7.1	rAKaCila	VM	<fs af='rAK,v,,,5,,Cila,Cila' name="rAKaCila">
7.2	।	SYM	
	))		
</Sentence>

And in the CoNLL format:

1 mAXabIlawA mAXabIlawA NP NNP lex-mAXabIlawA|cat-n|gend-|num-sg|pers-|case-d|vib-0|tam-0|head-mAXabIlawA|name-NP 7 k1 _ _
2 waKana waKana NP PRP lex-waKana|cat-pn|gend-|num-|pers-|case-d|vib-0|tam-0|head-waKana|name-NP2 7 k7t _ _
3 hAwera hAwa NP NN lex-hAwa|cat-n|gend-|num-sg|pers-|case-o|vib-era|tam-era|head-hAwera|name-NP3 4 r6 _ _
4 GadZi GadZi NP NN lex-GadZi|cat-unk|gend-|num-|pers-|case-|vib-|tam-|head-GadZi|name-NP4 5 k2 _ _
5 Kule Kul VGNF VM lex-Kul|cat-v|gend-|num-|pers-5|case-|vib-ne|tam-ne|head-Kule|name-VGNF 7 vmod _ _
6 tebile tebila NP NN lex-tebila|cat-n|gend-|num-sg|pers-|case-d|vib-me|tam-me|head-tebile|name-NP5 7 k7p _ _
7 rAKaCila rAK VGF VM lex-rAK|cat-v|gend-|num-|pers-5|case-|vib-Cila|tam-Cila|head-rAKaCila|name-VGF 0 main _ _

And after conversion of the WX encoding to the Bengali script in UTF-8:

1 মাধবীলতা মাধবীলতা NP NNP lex-mAXabIlawA|cat-n|gend-|num-sg|pers-|case-d|vib-0|tam-0|head-mAXabIlawA|name-NP 7 k1 _ _
2 তখন তখন NP PRP lex-waKana|cat-pn|gend-|num-|pers-|case-d|vib-0|tam-0|head-waKana|name-NP2 7 k7t _ _
3 হাতের হাত NP NN lex-hAwa|cat-n|gend-|num-sg|pers-|case-o|vib-era|tam-era|head-hAwera|name-NP3 4 r6 _ _
4 ঘড়ি ঘড়ি NP NN lex-GadZi|cat-unk|gend-|num-|pers-|case-|vib-|tam-|head-GadZi|name-NP4 5 k2 _ _
5 খুলে খুল্ VGNF VM lex-Kul|cat-v|gend-|num-|pers-5|case-|vib-ne|tam-ne|head-Kule|name-VGNF 7 vmod _ _
6 টেবিলে টেবিল NP NN lex-tebila|cat-n|gend-|num-sg|pers-|case-d|vib-me|tam-me|head-tebile|name-NP5 7 k7p _ _
7 রাখছিল রাখ্ VGF VM lex-rAK|cat-v|gend-|num-|pers-5|case-|vib-Cila|tam-Cila|head-rAKaCila|name-VGF 0 main _ _

Parsing

Nonprojectivities in HyDT-Bangla are not frequent. Only 78 of the 7252 chunks in the training+development ICON 2010 version are attached nonprojectively (1.08%).

The results of the ICON 2009 NLP tools contest have been published in (Husain, 2009). There were two evaluation rounds, the first with the coarse-grained syntactic tags, the second with the fine-grained syntactic tags. To reward language independence, only systems that parsed all three languages were officially ranked. The following table presents the Bengali/coarse-grained results of the four officially ranked systems, and the best Bengali-only* system.

Parser (Authors) LAS UAS
Kolkata (De et al.)* 84.29 90.32
Hyderabad (Ambati et al.) 78.25 90.22
Malt (Nivre) 76.07 88.97
Malt+MST (Zeman) 71.49 86.89
Mannem 70.34 83.56

The results of the ICON 2010 NLP tools contest have been published in (Husain et al., 2010), page 6. These are the best results for Bengali with fine-grained syntactic tags:

Parser (Authors) LAS UAS
Attardi et al. 70.66 87.41
Kosaraju et al. 70.55 86.16
Kolachina et al. 70.14 87.10

[ Back to the navigation ] [ Back to the content ]