Skip to content

Pipelines

A pipeline is an ordered list of process units, each performing one analysis step on the result of the previous ones. Pipelines are defined in the configuration files and selected with the -p option of analyzeText or the pipes/pipeline arguments of the Python API.

Neural pipelines (all languages)

These pipelines are defined in lima-lp-ud.xml and work for every language for which models are installed.

Pipeline Input Steps Output
deepud Plain text Tokenization and sentence splitting, PoS tagging, morphological features and lemmatization, dependency parsing CoNLL-U
deepud-pretok CoNLL-U PoS tagging, morphological features and lemmatization, dependency parsing CoNLL-U
deeplima Plain text Tokenization and sentence splitting, PoS tagging, morphological features and lemmatization CoNLL-U

They use the neural process units RnnTokenizer, RnnTokensAnalyzer and RnnDependencyParser (see Neural units). Steps without a model for the selected treebank are skipped. See Universal Dependencies analysis for details.

Legacy pipelines (English and French)

The legacy configurations lima-lp-eng.xml and lima-lp-fre.xml (languages eng and fre) use dictionaries, rules and statistical models. They are useful for resource development and rule-based extraction.

Pipeline Role
main Full legacy analysis: tokenization, morphological analysis, idioms, named entities, PoS tagging, rule-based syntactic analysis
ner-rules, ner-rules-pretok Rule-based named entity recognition on raw text or CoNLL-U input
ner-deep, ner-fusion (and -pretok) Kept for compatibility; currently identical to ner-rules (see NER)

The linguistic processing steps page describes what each step of the legacy analysis does.

Defining your own pipeline

A pipeline is a ProcessUnitPipeline group in the Processors module of a lima-lp-<lang>.xml file:

<group name="my-pipeline" class="ProcessUnitPipeline">
  <list name="processUnitSequence">
    <item value="RnnTokenizer"/>
    <item value="RnnTokensAnalyzer"/>
    <item value="conllDumper"/>
  </list>
</group>

Copy the configuration file to your own configuration directory before editing it (see Configuring LIMA). Dependencies between process units are not checked automatically: removing a unit that a later one needs can make LIMA fail.