Process units¶
The core of LIMA is the execution of process units in a pipeline. This section is the technical documentation of the process units of the standard LIMA distribution. Thanks to its plugin mechanism, LIMA can be extended with new units, which have their own documentation.
For each process unit, we describe:
- class: the identifier used to instantiate the corresponding C++ class
(the
classattribute of its configuration group); - role: what the unit does;
- inputs: the state of the LIMA data structures needed to run the unit, and the parameters that change its behavior;
- outputs: the data written to files or to the standard output;
- preconditions: the state the data structures must have reached before running the unit;
- effects: the changes made to the LIMA data structures by the unit.
| Family | Units |
|---|---|
| Neural units | RnnTokenizer, ConlluReader, RnnTokensAnalyzer, RnnDependencyParser, RnnNER |
| Tokenization | FlatTokenizer, SentenceBoundariesFinder |
| Morphology | SimpleWord, HyphenWordAlternatives, EnchantSpellingAlternatives, RegexMatcher, DefaultProperties, SimpleDefaultProperties |
| Entities and rules | ApplyRecognizer, GeoEntitiesTagger |
| PoS tagging | ViterbiPosTagger, SvmToolPosTagger, DynamicSvmToolPosTagger |
| Syntactic analysis | SyntacticAnalyzerChains, SyntacticAnalyzerDeps and related units |
| Semantics | CoreferencesSolving, WordSenseDisambiguation |
| Loggers and debugging | StatusLogger, XML loggers, graph writers… |
| Dumpers | ConllDumper, BowDumper, TextDumper, XML dumpers… |
Run analyzeText --availableUnits to list the units known to your
installation.