Differences
This shows you the differences between two versions of the page.
Both sides previous revision Previous revision Next revision | Previous revision Next revision Both sides next revision | ||
external:tectomt:tutorial [2009/01/17 16:34] kravalova |
external:tectomt:tutorial [2009/01/20 17:43] popel |
||
---|---|---|---|
Line 2: | Line 2: | ||
Welcome at TectoMT Tutorial. This tutorial should take about 2 hours. | Welcome at TectoMT Tutorial. This tutorial should take about 2 hours. | ||
+ | |||
Line 7: | Line 8: | ||
TectoMT is a highly modular NLP (Natural Language Processing) software system implemented in Perl programming language under Linux. It is primarily aimed at Machine Translation, | TectoMT is a highly modular NLP (Natural Language Processing) software system implemented in Perl programming language under Linux. It is primarily aimed at Machine Translation, | ||
+ | |||
===== Prerequisities ===== | ===== Prerequisities ===== | ||
+ | |||
+ | In this tutorial, we assume | ||
+ | |||
+ | * Your system is Linux | ||
+ | * Your shell is bash | ||
+ | * You have basic experience bash and you can read Perl | ||
+ | |||
+ | |||
+ | |||
+ | |||
+ | |||
+ | |||
+ | |||
+ | |||
==== Installation and setup ==== | ==== Installation and setup ==== | ||
- | In this tutorial, we assume that TectoMT has been successfully installed on your machine. For installation | + | * Checkout SVN repository. If you are running |
- | Before running any experiments with TectoMT, you must set up your environment by running | + | <code bash> |
+ | cd ~/BIG | ||
+ | svn --username < | ||
+ | </ | ||
+ | |||
+ | * In '' | ||
<code bash> | <code bash> | ||
- | source devel/config/init_devel_environ.sh | + | cd tectomt/install |
+ | ./install.sh | ||
</ | </ | ||
- | ==== Theoretical background ==== | + | * In your '' |
+ | |||
+ | <code bash> | ||
+ | source ~/ | ||
+ | </ | ||
+ | |||
+ | |||
+ | |||
+ | |||
+ | |||
Line 29: | Line 61: | ||
- | ==== TrEd ==== | ||
Line 37: | Line 68: | ||
===== TectoMT Architecture ===== | ===== TectoMT Architecture ===== | ||
+ | |||
+ | |||
Line 43: | Line 76: | ||
In TectoMT, there is the following hierarchy of processing units (software components that process data): | In TectoMT, there is the following hierarchy of processing units (software components that process data): | ||
- | * The basic units are blocks. They serve for some very limited, well defined, and often linguistically interpretable tasks (e.g., tokenization, | + | * The basic units are blocks. They serve for some very limited, well defined, and often linguistically interpretable tasks (e.g., tokenization, |
* To solve a more complex task, selected blocks can be chained into a block sequence, called also a scenario. Technically, | * To solve a more complex task, selected blocks can be chained into a block sequence, called also a scenario. Technically, | ||
- | * The highest unit is called application. Applications correspond to end-to-end tasks, be they real end-user applications (such as machine translation), | + | * The highest unit is called application. Applications correspond to end-to-end tasks, be they real end-user applications (such as machine translation), |
+ | |||
+ | This tutorial itself has its blocks in '' | ||
+ | |||
+ | |||
- | This tutorial itself has its blocks in '' | ||
Line 53: | Line 90: | ||
==== Layers of Linguistic Structures ==== | ==== Layers of Linguistic Structures ==== | ||
- | TectoMT blocks repository is saved in '' | + | {{ external: |
+ | |||
+ | TectoMT blocks repository is saved in '' | ||
- | Thus, the set of TectoMT layers is Cartesian product {S,T} x {English, | + | Thus, the set of TectoMT layers is a Cartesian product {S,T} x {English, |
* {S,T} distinguishes whether the data was created by analysis or transfer/ | * {S,T} distinguishes whether the data was created by analysis or transfer/ | ||
Line 61: | Line 100: | ||
* {W, | * {W, | ||
- | // | + | // |
+ | |||
+ | There are also other directories for other purpose blocks, for example blocks which only print out some information go to '' | ||
- | There are also other directories for other purpose blocks, for example blocks which only print out some information go to '' | ||
Line 71: | Line 112: | ||
===== First application ===== | ===== First application ===== | ||
- | Once you have TectoMT installed on your machine, you can find this tutorial in '' | + | Once you have TectoMT installed on your machine, you can find this tutorial in '' |
- | Most applications are defined in Makefiles, which describe sequence of blocks to be applied on our data. In our particular '' | + | Most applications are defined in Makefiles, which describe sequence of blocks to be applied on our data. In our particular '' |
We can run the application: | We can run the application: | ||
Line 81: | Line 122: | ||
</ | </ | ||
- | Our plain text data '' | + | Our plain text data '' |
- | * One physical file corresponds to one document. | + | * One physical |
* A document consists of a sequence of bundles (''< | * A document consists of a sequence of bundles (''< | ||
* Each bundle contains tree shaped sentence representations on various linguistic layers. In our example '' | * Each bundle contains tree shaped sentence representations on various linguistic layers. In our example '' | ||
* Trees are formed by nodes and edges. Attributes can be attached only to nodes. Edge's attributes must be equivalently stored as the lower node's attributes. Tree's attributes must be stored as attributes of the root node. | * Trees are formed by nodes and edges. Attributes can be attached only to nodes. Edge's attributes must be equivalently stored as the lower node's attributes. Tree's attributes must be stored as attributes of the root node. | ||
+ | |||
+ | |||
Line 107: | Line 150: | ||
===== Changing the scenario ===== | ===== Changing the scenario ===== | ||
- | We'll now add syntax analysis to our scenario by adding four more blocks. Instead of | + | We'll now add a syntax analysis |
<code bash> | <code bash> | ||
Line 115: | Line 158: | ||
SEnglishW_to_SEnglishM:: | SEnglishW_to_SEnglishM:: | ||
SEnglishW_to_SEnglishM:: | SEnglishW_to_SEnglishM:: | ||
- | SEnglishW_to_SEnglishM:: | + | SEnglishW_to_SEnglishM:: |
+ | | ||
</ | </ | ||
Line 126: | Line 170: | ||
SEnglishW_to_SEnglishM:: | SEnglishW_to_SEnglishM:: | ||
SEnglishW_to_SEnglishM:: | SEnglishW_to_SEnglishM:: | ||
- | SEnglishW_to_SEnglishM:: | + | SEnglishW_to_SEnglishM:: |
SEnglishM_to_SEnglishA:: | SEnglishM_to_SEnglishA:: | ||
SEnglishM_to_SEnglishA:: | SEnglishM_to_SEnglishA:: | ||
- | SEnglishM_to_SEnglishA:: | + | SEnglishM_to_SEnglishA:: |
+ | | ||
</ | </ | ||
Line 141: | Line 186: | ||
we can examine our '' | we can examine our '' | ||
+ | |||
+ | You can view the trees in '' | ||
+ | |||
+ | <code bash> | ||
+ | tmttred sample.tmt | ||
+ | </ | ||
+ | |||
+ | |||
+ | |||
+ | |||
+ | |||
+ | |||
+ | |||
+ | |||
+ | |||
Line 168: | Line 228: | ||
* node - '' | * node - '' | ||
- | We'll now examine an example of a new block in file '' | + | You can get TectoMT automatically execute your block code on each document or bundle by defining the main block entry point: |
- | This block illustrates the most common methods for accessing objects: | + | * '' |
+ | * '' | ||
+ | |||
+ | Each block must have exactly one entry point. | ||
+ | |||
+ | We'll now examine an example of a new block in file '' | ||
+ | |||
+ | This block illustrates | ||
* '' | * '' | ||
Line 183: | Line 250: | ||
* '' | * '' | ||
- | Our tutorial block '' | + | Our tutorial block '' |
- | * Copy the block to the right place in blocks repository '' | + | <code bash> |
- | | + | print_info: |
- | cp libs/ | + | brunblocks -S -o Tutorial::Print_node_info |
- | </ | + | </ |
- | * in copied file '' | + | |
- | * Add this block to our scenario: | + | |
- | <code bash> | + | |
- | print_afun: | + | |
- | brunblocks -S -o Print::Tutorial | + | |
- | </ | + | |
We can observe our new block behaviour: | We can observe our new block behaviour: | ||
<code bash> | <code bash> | ||
- | make print_afun | + | make print_info |
</ | </ | ||
+ | Try to change the block so that it prints out the information only for verbs. (You need to look at attribute '' | ||
+ | |||
+ | |||
+ | ===== Advanced block: finite clauses ===== | ||
- | ===== Advanced block: finite clauses ===== | ||
==== Motivation ==== | ==== Motivation ==== | ||
+ | |||
+ | It is assumed that finite clauses can be translated independently, | ||
+ | |||
+ | |||
==== Task ==== | ==== Task ==== | ||
- | A block which, given an analytical tree ('' | + | A block which, given an analytical tree ('' |
+ | |||
+ | |||
+ | |||
+ | |||
+ | |||
+ | |||
+ | |||
+ | |||
+ | |||
+ | |||
+ | |||
+ | |||
+ | |||
+ | |||
+ | |||
+ | |||
+ | |||
+ | |||
- | ==== Algorithm ==== | ||
- | |||
==== Instructions ==== | ==== Instructions ==== | ||
- | There is a block template with hints in '' | + | There is a block template with hints in '' |
<code bash> | <code bash> | ||
finite_clauses: | finite_clauses: | ||
brunblocks -S -o \ | brunblocks -S -o \ | ||
- | | + | |
- | | + | |
</ | </ | ||
+ | |||
+ | You are going to need these methods: | ||
+ | |||
+ | * '' | ||
+ | * '' | ||
+ | * '' | ||
+ | * '' | ||
+ | |||
+ | //Note//: '' | ||
+ | |||
+ | |||
+ | |||
+ | //Advanced version//: The output of our block might still be incorrect in special cases - we don't solve coordination and subordinate conjunctions. | ||
+ | |||
+ | |||
+ | |||
+ | ===== Your turn: more tasks ===== | ||
+ | |||
+ | |||
+ | |||
+ | |||
+ | |||
+ | |||
+ | |||
+ | |||
+ | ==== SVO to SOV ==== | ||
+ | |||
+ | **Motivation**: | ||
+ | |||
+ | **Task**: Change the word order from SVO to SOV. | ||
+ | |||
+ | **Instructions**: | ||
+ | |||
+ | * To find an object to a verb, look for objects among effective children of a verb ('' | ||
+ | * Once you have node '' | ||
+ | * For debugging, a method returning word order of a node is useful: '' | ||
+ | |||
+ | |||
+ | |||
+ | |||
+ | |||
+ | |||
+ | |||
+ | |||
+ | |||
+ | |||
+ | |||
+ | |||
+ | |||
+ | |||
+ | |||
+ | |||
+ | |||
+ | |||
+ | |||
+ | |||
+ | ==== Prepositions ==== | ||
+ | |||
+ | **Motivation**: | ||
+ | |||
+ | TODO obrazek | ||
+ | |||
+ | **Task**: The task is to rehang all prepositions as indicated at the picture. You may assume that prepositions have at most 1 child. | ||
+ | |||
+ | ** Instructions**: | ||
+ | |||
+ | You are going to need these new methods: | ||
+ | * '' | ||
+ | * '' | ||
+ | * '' | ||
+ | |||
+ | // | ||
+ | * On analytical layer, you can use this test to recognize prepositions: | ||
+ | * You can use block template in '' | ||
+ | |||
+ | |||
+ | //Advanced version//: What happens in case of multiword prepositions? | ||
+ | |||
+ | |||
- | ===== Final task: ? ===== | ||
===== Further information ===== | ===== Further information ===== | ||
- | * [[http:// | + | * [[http:// |
+ | * Questions? Ask '' | ||
+ | * Solutions to this tutorial tasks are in '' | ||
+ | * [[http:// | ||