Differences
This shows you the differences between two versions of the page.
Both sides previous revision Previous revision Next revision | Previous revision Next revision Both sides next revision | ||
external:tectomt:tutorial [2009/01/19 11:57] kravalova |
external:tectomt:tutorial [2009/01/20 17:43] popel |
||
---|---|---|---|
Line 2: | Line 2: | ||
Welcome at TectoMT Tutorial. This tutorial should take about 2 hours. | Welcome at TectoMT Tutorial. This tutorial should take about 2 hours. | ||
+ | |||
Line 7: | Line 8: | ||
TectoMT is a highly modular NLP (Natural Language Processing) software system implemented in Perl programming language under Linux. It is primarily aimed at Machine Translation, | TectoMT is a highly modular NLP (Natural Language Processing) software system implemented in Perl programming language under Linux. It is primarily aimed at Machine Translation, | ||
+ | |||
===== Prerequisities ===== | ===== Prerequisities ===== | ||
+ | |||
+ | In this tutorial, we assume | ||
+ | |||
+ | * Your system is Linux | ||
+ | * Your shell is bash | ||
+ | * You have basic experience bash and you can read Perl | ||
+ | |||
+ | |||
+ | |||
+ | |||
+ | |||
+ | |||
+ | |||
+ | |||
==== Installation and setup ==== | ==== Installation and setup ==== | ||
- | In this tutorial, we assume that TectoMT has been successfully installed on your machine. For installation | + | * Checkout SVN repository. If you are running |
- | Before running any experiments with TectoMT, you must set up your environment by running | + | <code bash> |
+ | cd ~/BIG | ||
+ | svn --username < | ||
+ | </ | ||
+ | |||
+ | * In '' | ||
<code bash> | <code bash> | ||
- | source devel/config/init_devel_environ.sh | + | cd tectomt/install |
+ | ./install.sh | ||
</ | </ | ||
+ | |||
+ | * In your '' | ||
+ | |||
+ | <code bash> | ||
+ | source ~/ | ||
+ | </ | ||
+ | |||
+ | |||
+ | |||
+ | |||
- | ==== Theoretical background ==== | ||
- | TODO obrazek | ||
Line 31: | Line 61: | ||
- | ==== TrEd ==== | ||
Line 39: | Line 68: | ||
===== TectoMT Architecture ===== | ===== TectoMT Architecture ===== | ||
+ | |||
+ | |||
Line 45: | Line 76: | ||
In TectoMT, there is the following hierarchy of processing units (software components that process data): | In TectoMT, there is the following hierarchy of processing units (software components that process data): | ||
- | * The basic units are blocks. They serve for some very limited, well defined, and often linguistically interpretable tasks (e.g., tokenization, | + | * The basic units are blocks. They serve for some very limited, well defined, and often linguistically interpretable tasks (e.g., tokenization, |
* To solve a more complex task, selected blocks can be chained into a block sequence, called also a scenario. Technically, | * To solve a more complex task, selected blocks can be chained into a block sequence, called also a scenario. Technically, | ||
- | * The highest unit is called application. Applications correspond to end-to-end tasks, be they real end-user applications (such as machine translation), | + | * The highest unit is called application. Applications correspond to end-to-end tasks, be they real end-user applications (such as machine translation), |
+ | |||
+ | This tutorial itself has its blocks in '' | ||
+ | |||
+ | |||
- | This tutorial itself has its blocks in '' | ||
Line 55: | Line 90: | ||
==== Layers of Linguistic Structures ==== | ==== Layers of Linguistic Structures ==== | ||
- | TectoMT blocks repository is saved in '' | + | {{ external: |
- | Thus, the set of TectoMT layers is Cartesian product {S,T} x {English, | + | TectoMT blocks repository is saved in '' |
+ | |||
+ | Thus, the set of TectoMT layers is a Cartesian product {S,T} x {English, | ||
* {S,T} distinguishes whether the data was created by analysis or transfer/ | * {S,T} distinguishes whether the data was created by analysis or transfer/ | ||
Line 63: | Line 100: | ||
* {W, | * {W, | ||
- | // | + | // |
+ | |||
+ | There are also other directories for other purpose blocks, for example blocks which only print out some information go to '' | ||
- | There are also other directories for other purpose blocks, for example blocks which only print out some information go to '' | ||
Line 73: | Line 112: | ||
===== First application ===== | ===== First application ===== | ||
- | Once you have TectoMT installed on your machine, you can find this tutorial in '' | + | Once you have TectoMT installed on your machine, you can find this tutorial in '' |
- | Most applications are defined in Makefiles, which describe sequence of blocks to be applied on our data. In our particular '' | + | Most applications are defined in Makefiles, which describe sequence of blocks to be applied on our data. In our particular '' |
We can run the application: | We can run the application: | ||
Line 83: | Line 122: | ||
</ | </ | ||
- | Our plain text data '' | + | Our plain text data '' |
- | * One physical file corresponds to one document. | + | * One physical |
* A document consists of a sequence of bundles (''< | * A document consists of a sequence of bundles (''< | ||
* Each bundle contains tree shaped sentence representations on various linguistic layers. In our example '' | * Each bundle contains tree shaped sentence representations on various linguistic layers. In our example '' | ||
* Trees are formed by nodes and edges. Attributes can be attached only to nodes. Edge's attributes must be equivalently stored as the lower node's attributes. Tree's attributes must be stored as attributes of the root node. | * Trees are formed by nodes and edges. Attributes can be attached only to nodes. Edge's attributes must be equivalently stored as the lower node's attributes. Tree's attributes must be stored as attributes of the root node. | ||
+ | |||
Line 110: | Line 150: | ||
===== Changing the scenario ===== | ===== Changing the scenario ===== | ||
- | We'll now add syntax analysis to our scenario by adding four more blocks. Instead of | + | We'll now add a syntax analysis |
<code bash> | <code bash> | ||
Line 118: | Line 158: | ||
SEnglishW_to_SEnglishM:: | SEnglishW_to_SEnglishM:: | ||
SEnglishW_to_SEnglishM:: | SEnglishW_to_SEnglishM:: | ||
- | SEnglishW_to_SEnglishM:: | + | SEnglishW_to_SEnglishM:: |
+ | | ||
</ | </ | ||
Line 129: | Line 170: | ||
SEnglishW_to_SEnglishM:: | SEnglishW_to_SEnglishM:: | ||
SEnglishW_to_SEnglishM:: | SEnglishW_to_SEnglishM:: | ||
- | SEnglishW_to_SEnglishM:: | + | SEnglishW_to_SEnglishM:: |
SEnglishM_to_SEnglishA:: | SEnglishM_to_SEnglishA:: | ||
SEnglishM_to_SEnglishA:: | SEnglishM_to_SEnglishA:: | ||
- | SEnglishM_to_SEnglishA:: | + | SEnglishM_to_SEnglishA:: |
+ | | ||
</ | </ | ||
Line 150: | Line 192: | ||
tmttred sample.tmt | tmttred sample.tmt | ||
</ | </ | ||
+ | |||
+ | |||
+ | |||
+ | |||
+ | |||
+ | |||
+ | |||
Line 201: | Line 250: | ||
* '' | * '' | ||
- | Our tutorial block '' | + | Our tutorial block '' |
- | * Copy the block to the right place in blocks repository '' | + | <code bash> |
- | | + | print_info: |
- | cp libs/ | + | brunblocks -S -o Tutorial:: |
- | </ | + | </ |
- | * in copied file '' | + | |
- | * Add this block to our scenario: | + | |
- | <code bash> | + | |
- | print_afun: | + | |
- | | + | |
- | </ | + | |
We can observe our new block behaviour: | We can observe our new block behaviour: | ||
<code bash> | <code bash> | ||
- | make print_afun | + | make print_info |
</ | </ | ||
+ | Try to change the block so that it prints out the information only for verbs. (You need to look at attribute '' | ||
+ | |||
+ | |||
+ | ===== Advanced block: finite clauses ===== | ||
- | ===== Advanced block: finite clauses ===== | ||
==== Motivation ==== | ==== Motivation ==== | ||
+ | |||
+ | It is assumed that finite clauses can be translated independently, | ||
+ | |||
+ | |||
==== Task ==== | ==== Task ==== | ||
- | A block which, given an analytical tree ('' | + | A block which, given an analytical tree ('' |
+ | |||
+ | |||
+ | |||
+ | |||
+ | |||
+ | |||
+ | |||
+ | |||
+ | |||
+ | |||
+ | |||
+ | |||
+ | |||
+ | |||
+ | |||
- | ==== Algorithm ==== | ||
Line 247: | Line 312: | ||
==== Instructions ==== | ==== Instructions ==== | ||
- | There is a block template with hints in '' | + | There is a block template with hints in '' |
<code bash> | <code bash> | ||
finite_clauses: | finite_clauses: | ||
brunblocks -S -o \ | brunblocks -S -o \ | ||
- | | + | |
- | | + | |
</ | </ | ||
You are going to need these methods: | You are going to need these methods: | ||
- | * '' | + | * '' |
- | * '' | + | * '' |
* '' | * '' | ||
- | * '' | + | * '' |
- | * '' | + | |
+ | //Note//: '' | ||
+ | |||
+ | //Advanced version//: The output of our block might still be incorrect in special cases - we don't solve coordination and subordinate conjunctions. | ||
Line 274: | Line 341: | ||
- | ==== Coordination ==== | ||
- | This time ... | ||
- | You can use block template in '' | + | |
+ | |||
+ | ==== SVO to SOV ==== | ||
+ | |||
+ | **Motivation**: | ||
+ | |||
+ | **Task**: Change the word order from SVO to SOV. | ||
+ | |||
+ | **Instructions**: | ||
+ | |||
+ | * To find an object to a verb, look for objects among effective children of a verb ('' | ||
+ | * Once you have node '' | ||
+ | * For debugging, a method returning word order of a node is useful: '' | ||
+ | |||
+ | |||
+ | |||
+ | |||
+ | |||
+ | |||
+ | |||
+ | |||
+ | |||
+ | |||
+ | |||
+ | |||
+ | |||
+ | |||
+ | |||
+ | |||
+ | |||
+ | |||
+ | |||
+ | |||
+ | ==== Prepositions ==== | ||
+ | |||
+ | **Motivation**: | ||
+ | |||
+ | TODO obrazek | ||
+ | |||
+ | **Task**: The task is to rehang all prepositions as indicated at the picture. You may assume that prepositions have at most 1 child. | ||
+ | |||
+ | ** Instructions**: | ||
+ | |||
+ | You are going to need these new methods: | ||
+ | * '' | ||
+ | * '' | ||
+ | * '' | ||
+ | |||
+ | //Hint//: | ||
+ | * On analytical layer, you can use this test to recognize prepositions: | ||
+ | * You can use block template in '' | ||
+ | |||
+ | |||
+ | //Advanced version//: What happens in case of multiword prepositions? | ||
+ | |||
+ | |||
Line 286: | Line 407: | ||
* [[http:// | * [[http:// | ||
* Questions? Ask '' | * Questions? Ask '' | ||
- | * Solutions to | + | * Solutions to this tutorial tasks are in '' |
+ | * [[http:// | ||