Differences

This shows you the differences between two versions of the page.

--- user:zeman:treebanks:te [2012/03/22 11:34]
zeman Training data size (both sentences and words) was identical in ICON 2009 and 2010.
+++ user:zeman:treebanks:te [2012/03/22 11:46]
zeman ICON 2009 Telugu data size.
@@ Line 43: / Line 43: @@
 ==== Size ====
-HyDT-Telugu shows dependencies between chunks, not words. The node/tree ratio is thus much lower than in other treebanks. The ICON 2009 version came with a data split into three parts: training, development and test; the same data was also distributed for ICON 2010:
+HyDT-Telugu shows dependencies between chunks, not words. The node/tree ratio is thus much lower than in other treebanks. The ICON 2009 version came with a data split into three parts: training, development and test:
+^ Part ^ Sentences ^ Chunks ^ Ratio ^
+| Training    |  1456 |  5494 | 3.77 |
+| Development |   150 |   675 | 4.50 |
+| Test        |   150 |   583 | 3.89 |
+| TOTAL       |  1756 |  6752 | 3.85 |
+The data distributed for ICON 2010 was slightly smaller, maybe it had been cleaned up? Note that the number of training words, 7602, is identical to the number published for ICON 2009. I cannot verify it because I only see chunks, not words in the CoNLL data format.
 ^ Part ^ Sentences ^ Chunks ^ Ratio ^ Words ^ Ratio ^

Institute of Formal and Applied Linguistics Wiki