Skip to content

Process units

The core of LIMA is the execution of process units in a pipeline. This section is the technical documentation of the process units of the standard LIMA distribution. Thanks to its plugin mechanism, LIMA can be extended with new units, which have their own documentation.

For each process unit, we describe:

  • class: the identifier used to instantiate the corresponding C++ class (the class attribute of its configuration group);
  • role: what the unit does;
  • inputs: the state of the LIMA data structures needed to run the unit, and the parameters that change its behavior;
  • outputs: the data written to files or to the standard output;
  • preconditions: the state the data structures must have reached before running the unit;
  • effects: the changes made to the LIMA data structures by the unit.
Family Units
Neural units RnnTokenizer, ConlluReader, RnnTokensAnalyzer, RnnDependencyParser, RnnNER
Tokenization FlatTokenizer, SentenceBoundariesFinder
Morphology SimpleWord, HyphenWordAlternatives, EnchantSpellingAlternatives, RegexMatcher, DefaultProperties, SimpleDefaultProperties
Entities and rules ApplyRecognizer, GeoEntitiesTagger
PoS tagging ViterbiPosTagger, SvmToolPosTagger, DynamicSvmToolPosTagger
Syntactic analysis SyntacticAnalyzerChains, SyntacticAnalyzerDeps and related units
Semantics CoreferencesSolving, WordSenseDisambiguation
Loggers and debugging StatusLogger, XML loggers, graph writers…
Dumpers ConllDumper, BowDumper, TextDumper, XML dumpers…

Run analyzeText --availableUnits to list the units known to your installation.