Terminology recognition and auto-completer in OmegaT 3.0

Опубликовано: 28 Август 2026
на канале: CATguruEN
10,696
165

http://wordfast.fi/blog/?p=962
  / catguruen  

Terminology recognition in OmegaT is handled via tokenizers. Starting with version 3.0.0, tokenizers are included in the standard OmegaT distribution, whereas one had to download them separately in previous versions. They are also automatically selected during the project creation process, whereas one had to launch them via the command line in previous versions. Tokenizers are especially important for terminology recognition in heavily inflected languages. This video shows how the tokenizer works with Finnish as the source language.

Starting with OmegaT version 3.0.1, recognized terminology can be inserted in the target segment via a new auto-completer feature, which works entirely in the editor pane and with the keyboard (the shortcut is Ctrl+space, or Command+space in OS X). In previous versions, one had to right-click with the mouse in the glossary pane. This video shows how terminology can be inserted in the target segment, using a sample Finnish-English project.

Related links:
Finnish tokenizer (Lucene): http://grepcode.com/file/repo1.maven....
Wikipedia article on tokenization: http://en.wikipedia.org/wiki/Tokeniza...

Related videos:
First steps with OmegaT:    • First steps with OmegaT (2012 edition)  
Machine translation in OmegaT for Mac:    • Machine translation in OmegaT for Mac  
Machine translation in OmegaT for Windows:    • Machine translation in OmegaT for Windows  
First steps with OmegaT:    • First steps with OmegaT (2012 edition)