[ Skip to the content ]

Institute of Formal and Applied Linguistics Wiki


[ Back to the navigation ]

Differences

This shows you the differences between two versions of the page.

Link to this comparison view

Both sides previous revision Previous revision
Next revision
Previous revision
Next revision Both sides next revision
user:zeman:treebanks:te [2012/03/22 11:30]
zeman Size.
user:zeman:treebanks:te [2012/03/22 11:46]
zeman ICON 2009 Telugu data size.
Line 46: Line 46:
  
 ^ Part ^ Sentences ^ Chunks ^ Ratio ^ ^ Part ^ Sentences ^ Chunks ^ Ratio ^
-| Training | 980 6449 6.58 +| Training     1456  5494 3.77 
-| Development | 150 | 811 5.41 +| Development |   150 |   675 4.50 
-| Test | 150 | 961 6.41 +| Test          150 |   583 3.89 
-| TOTAL | 1280 8221 6.42 |+| TOTAL        1756  6752 3.85 |
  
-The ICON 2010 version came with a data split into three parts: training, development and test:+The data distributed for ICON 2010 was slightly smaller, maybe it had been cleaned up? Note that the number of training words7602, is identical to the number published for ICON 2009. I cannot verify it because I only see chunks, not words in the CoNLL data format.
  
 ^ Part ^ Sentences ^ Chunks ^ Ratio ^ Words ^ Ratio ^ ^ Part ^ Sentences ^ Chunks ^ Ratio ^ Words ^ Ratio ^

[ Back to the navigation ] [ Back to the content ]