Differences
This shows you the differences between two versions of the page.
Both sides previous revision Previous revision Next revision | Previous revision Next revision Both sides next revision | ||
user:hladka:playcoref [2009/02/26 12:40] hladka |
user:hladka:playcoref [2009/03/11 14:59] hladka |
||
---|---|---|---|
Line 77: | Line 77: | ||
====== Specification ====== | ====== Specification ====== | ||
+ | |||
Line 85: | Line 86: | ||
* A game of two players. Players are paired randomly. Computer as a player: automatic coreference resolution **???????** | * A game of two players. Players are paired randomly. Computer as a player: automatic coreference resolution **???????** | ||
* Session time up to **???????** minutes. | * Session time up to **???????** minutes. | ||
- | * At the beginning of the game, if there is no coreference pair in the first two sentences (as determined by the manual/ | + | * At the beginning of the game, if there is no coreference pair in the first two sentences (as determined by the manual/ |
* What my partner is doing? If (s)he hooks up the same pair of words as I hooked up then the pair of words starts **??????? | * What my partner is doing? If (s)he hooks up the same pair of words as I hooked up then the pair of words starts **??????? | ||
* The players can re-hook up any word any time in the session. | * The players can re-hook up any word any time in the session. | ||
Line 92: | Line 93: | ||
- POS tagger | - POS tagger | ||
- coreference resolution procedure | - coreference resolution procedure | ||
+ | |||
+ | |||
+ | |||
+ | |||
+ | |||
+ | |||
+ | |||
+ | |||
Line 106: | Line 115: | ||
* Anja's data ## // PDT data that are currently being annotated for the extended coreference // | * Anja's data ## // PDT data that are currently being annotated for the extended coreference // | ||
* **JM**: It would be nice if the players could choose a domain of the texts to play on (science-fiction, | * **JM**: It would be nice if the players could choose a domain of the texts to play on (science-fiction, | ||
- | | + | ***JM (6/3/09)**: Predelal jsem data pro playcoref, ted obsahuji jenom koreference mezi uzly s tagy N nebo P. Data jsou v adresari: ''/ |
- | vety/dokument; sipky_noun_noun-noun_pronoun-pronoun-pronoun/document; | + | ***BH (11/ |
* **EN** | * **EN** | ||
- | * search the data that are available | + | * search the data that are available; **BH (11/3/09)** Z dokumentace dat, ktera bychom meli mit, jsem nasla MUC6, ale nevidim tam data s koreferenci. Jirka zjisti, jestli jsou nekde jinde nebo jak jinak se k nim muzeme dostat. |
=== Coding === | === Coding === | ||
* utf-8 | * utf-8 | ||
Line 124: | Line 132: | ||
* sentence by sentence | * sentence by sentence | ||
* supervised selection of documents for a session | * supervised selection of documents for a session | ||
+ | |||
+ | |||
===== Scoring ===== | ===== Scoring ===== | ||
- | * '' | + | * '' |
**JM**: | **JM**: | ||
Line 135: | Line 145: | ||
**BH**: Jirka ma pravdu. Pocitani skore musi byt objektivni. Proto jsem vzorecek upravila tak, ze nebude pocitat shodu hrace vzhledem k rucni anotaci. | **BH**: Jirka ma pravdu. Pocitani skore musi byt objektivni. Proto jsem vzorecek upravila tak, ze nebude pocitat shodu hrace vzhledem k rucni anotaci. | ||
- | |||
Line 141: | Line 150: | ||
===== Output Data Needed ===== | ===== Output Data Needed ===== | ||
* score list ## // | * score list ## // | ||
- | * documents after the '' | + | * documents after the '' |
+ | (**JM**: Mluvil jsem kvůli měření mezianotátorské shody v anotování koreference se Zdeňkem a vyšlo z toho, že na měření shody na šipkách by použil prostě jen F-measure. Její smysl je jasný a je symetrická. Kappa je nevhodná kvůli tomu, že pravděpodobnost náhodné shody je poměrně nízká a těžko se určuje; kappa se hodí spíš pro klasifikační úlohy (proto ji použiju v Anjiině projektu na shodu v určování typu koreference, | ||
+ | - kappa measure | ||
+ | - G-theory - see [[http:// | ||
+ | Identifying Sources of Disagreement: | ||
+ | - the Pearson correlation - see (Snow et al., 2008) [[http:// | ||
* session | * session | ||
* player_A_id, | * player_A_id, | ||
* document(s) | * document(s) | ||
* number of corrections by player_A and by player_B (**JM**: I do not see the point in this) | * number of corrections by player_A and by player_B (**JM**: I do not see the point in this) | ||
- | * corrections by player_A and by player_B (**JM**: and maybe nor in this) (**BH**: I am interested in the manner of the players. Maybe the corrections will be total mess, but we have to see the data at least from the very first sessions. ) | + | * corrections by player_A and by player_B (**JM**: and maybe nor in this) (**BH**: I am interested in the players' behaviour. Maybe the corrections will be total mess, but we have to see the data at least from the very first sessions. ) |
===== Design ===== | ===== Design ===== | ||
Line 163: | Line 178: | ||
+ | ===== Tools needed ===== | ||
+ | * tagger ## tool_chain (CAC2.0) | ||
+ | * Linh's coreference resolution procedure - see TectoMT - **JM** | ||
+ | * vyzkouset - trenink a test - na datech Anji | ||
+ | * conversion: csts <-> pml m_coref scheme | ||
- | ===== Tools needed | + | |
- | | + | |
- | * Linh's coreference resolution procedure **---PS TO DO---** What type of input data the Linh's procedure works with? '' | + | |
- | * conversion: csts <-> pml m_coref scheme | + | ====== ACL - IJCNLP2009 ====== |
+ | | ||
+ | | ||
+ | |