[ Skip to the content ]

Institute of Formal and Applied Linguistics Wiki


[ Back to the navigation ]

Differences

This shows you the differences between two versions of the page.

Link to this comparison view

Next revision
Previous revision
Next revision Both sides next revision
user:zeman:treebanks:te [2012/03/22 11:14]
zeman vytvořeno
user:zeman:treebanks:te [2012/03/22 11:34]
zeman Training data size (both sentences and words) was identical in ICON 2009 and 2010.
Line 43: Line 43:
 ==== Size ==== ==== Size ====
  
-HyDT-Telugu shows dependencies between chunks, not words. The node/tree ratio is thus much lower than in other treebanks. The ICON 2009 version came with a data split into three parts: training, development and test+HyDT-Telugu shows dependencies between chunks, not words. The node/tree ratio is thus much lower than in other treebanks. The ICON 2009 version came with a data split into three parts: training, development and test; the same data was also distributed for ICON 2010:
- +
-^ Part ^ Sentences ^ Chunks ^ Ratio ^ +
-| Training | 980 | 6449 | 6.58 | +
-| Development | 150 | 811 | 5.41 | +
-| Test | 150 | 961 | 6.41 | +
-| TOTAL | 1280 | 8221 | 6.42 | +
- +
-The ICON 2010 version came with a data split into three parts: training, development and test:+
  
 ^ Part ^ Sentences ^ Chunks ^ Ratio ^ Words ^ Ratio ^ ^ Part ^ Sentences ^ Chunks ^ Ratio ^ Words ^ Ratio ^
-| Training | 979 6440 6.58 10305 10.52 +| Training | 1400 7602 5.43 
-| Development | 150 | 812 5.41 1196 7.97 +| Development | 150 | 839 5.59 
-| Test | 150 | 961 6.41 1350 9.00 +| Test | 150 | 836 5.57 
-| TOTAL | 1279 8213 6.42 12851 10.04 |+| TOTAL | 1700 9277 5.46 |
  
-I have counted the sentences and chunks. The number of words comes from (Husain et al., 2010). Note that the paper gives the number of training sentences as 980 (instead of 979), which is a mistake. The last training sentence has the id 980 but there is no sentence with id 418.+We drew our training and test data from the ICON 2010 datasets but we have fewer sentences – why?
  
-Apparently the training-development-test data split was more or less identical in both years, except for the minor discrepancies (number of training sentences and development chunks).+^ Part ^ Sentences ^ Chunks ^ Ratio ^ 
 +| Training |  1300 |  5125 | 3.94 | 
 +| Test       150 |   597 | 3.98 | 
 +| TOTAL    |  1450 |  5722 | 3.95 |
  
 ==== Inside ==== ==== Inside ====

[ Back to the navigation ] [ Back to the content ]