Differences
This shows you the differences between two versions of the page.
— |
user:zeman:treebanks:bg [2011/11/20 21:19] (current) zeman vytvořeno |
||
---|---|---|---|
Line 1: | Line 1: | ||
+ | ===== Bulgarian (bg) ===== | ||
+ | |||
+ | [[http:// | ||
+ | |||
+ | ==== Versions ==== | ||
+ | |||
+ | * Original BTB in native format | ||
+ | * CoNLL 2006 (BulTreeBank-DP) | ||
+ | |||
+ | The original BTB is based on HPSG (head-driven phrase-structure grammar). The CoNLL version contains only the dependency information encoded in HPSG BulTreeBank. | ||
+ | |||
+ | ==== Obtaining and License ==== | ||
+ | |||
+ | Only the CoNLL version seems to be distributed but you may ask the creators about the HPSG version. For the dependency version, print the [[http:// | ||
+ | |||
+ | * research usage | ||
+ | * no redistribution | ||
+ | * cite [[http:// | ||
+ | |||
+ | BTB was created by members of the [[http:// | ||
+ | |||
+ | ==== References ==== | ||
+ | |||
+ | * Website | ||
+ | * http:// | ||
+ | * Data | ||
+ | * //no separate citation// | ||
+ | * Principal publications | ||
+ | * Kiril Simov, Petya Osenova, Alexander Simov, Milen Kouylekov: //Design and Implementation of the Bulgarian HPSG-based Treebank.// In: Erhard Hinrichs, Kiril Simov (eds.): Journal of Research on Language and Computation, | ||
+ | * Documentation | ||
+ | * Kiril Simov, Petya Osenova, Milena Slavcheva: [[http:// | ||
+ | * Petya Osenova, Kiril Simov: [[http:// | ||
+ | * http:// | ||
+ | |||
+ | ==== Domain ==== | ||
+ | |||
+ | Unknown (“A set of Bulgarian sentences marked-up with detailed syntactic information. These sentences are mainly extracted from authentic Bulgarian texts. They are chosen with regards two criteria. First, they cover the variety of syntactic structures of Bulgarian. Second, they show the statistical distribution of these phenomena in real texts.”) At least part of it is probably news (Novinar, Sega, Standart). | ||
+ | |||
+ | ==== Size ==== | ||
+ | |||
+ | The CoNLL 2006 version contains 196,151 tokens in 13221 sentences, yielding 14.84 tokens per sentence on average (CoNLL 2006 data split: 190,217 tokens / 12823 sentences training, 5934 tokens / 398 sentences test). | ||
+ | |||
+ | ==== Inside ==== | ||
+ | |||
+ | The original morphosyntactic tags have been converted to fit into the three columns (CPOS, POS and FEAT) of the CoNLL format. There //should// be a 1-1 mapping between the [[http:// | ||
+ | |||
+ | The morphological analysis does not include lemmas. The morphosyntactic tags have been assigned (probably) manually. | ||
+ | |||
+ | The guidelines for syntactic annotation are documented in the other [[http:// | ||
+ | |||
+ | ==== Sample ==== | ||
+ | |||
+ | The first three sentences of the CoNLL 2006 training data: | ||
+ | |||
+ | | 1 | Глава | _ | N | Nc | _ | 0 | ROOT | 0 | ROOT | | ||
+ | | 2 | трета | _ | M | Mo | gen=f< | ||
+ | | |||||||||| | ||
+ | | 1 | НАРОДНО | _ | A | An | gen=n< | ||
+ | | 2 | СЪБРАНИЕ | _ | N | Nc | gen=n< | ||
+ | | |||||||||| | ||
+ | | 1 | Народното | _ | A | An | gen=n< | ||
+ | | 2 | събрание | _ | N | Nc | gen=n< | ||
+ | | 3 | осъществява | _ | V | Vpi | trans=t< | ||
+ | | 4 | законодателната | _ | A | Af | gen=f< | ||
+ | | 5 | власт | _ | N | Nc | _ | 3 | obj | 3 | obj | | ||
+ | | 6 | и | _ | C | Cp | _ | 3 | conj | 3 | conj | | ||
+ | | 7 | упражнява | _ | V | Vpi | trans=t< | ||
+ | | 8 | парламентарен | _ | A | Am | gen=m< | ||
+ | | 9 | контрол | _ | N | Nc | gen=m< | ||
+ | | 10 | . | _ | Punct | Punct | _ | 3 | punct | 3 | punct | | ||
+ | |||
+ | The first three sentences of the CoNLL 2006 test data: | ||
+ | |||
+ | | 1 | Единственото | _ | A | An | gen=n< | ||
+ | | 2 | решение | _ | N | Nc | gen=n< | ||
+ | | |||||||||| | ||
+ | | 1 | Ерик | _ | N | Np | gen=m< | ||
+ | | 2 | Франк | _ | N | Np | gen=m< | ||
+ | | 3 | Ръсел | _ | H | Hm | gen=m< | ||
+ | | |||||||||| | ||
+ | | 1 | Пълен | _ | A | Am | gen=m< | ||
+ | | 2 | мрак | _ | N | Nc | gen=m< | ||
+ | | 3 | и | _ | C | Cp | _ | 2 | conj | 2 | conj | | ||
+ | | 4 | пълна | _ | A | Af | gen=f< | ||
+ | | 5 | самота | _ | N | Nc | _ | 2 | conjarg | 2 | conjarg | | ||
+ | | 6 | . | _ | Punct | Punct | _ | 2 | punct | 2 | punct | | ||
+ | |||
+ | ==== Parsing ==== | ||
+ | |||
+ | Nonprojectivities in BTB are rare. Only 747 of the 196,151 tokens in the CoNLL 2006 version are attached nonprojectively (0.38%). | ||
+ | |||
+ | The results of the CoNLL 2006 shared task are [[http:// | ||
+ | |||
+ | ^ Parser (Authors) ^ LAS ^ UAS ^ | ||
+ | | MST (McDonald et al.) | 87.57 | 92.04 | | ||
+ | | Malt (Nivre et al.) | 87.41 | 91.72 | | ||
+ | | Nara (Yuchang Cheng) | 86.34 | 91.30 | | ||