Neural units¶
The neural process units are built on the in-tree deeplima library and
libtorch. They are used by the deepud, deepud-pretok and deeplima
pipelines and defined in
lima-lp-ud.xml.
Their model file names are built from a prefix in which $udlang is replaced
by the treebank identifier selected at analysis time (e.g.
eng-UD_English-EWT, see Language models). Models
are searched in the ud subdirectory of the unit's directory in the
resources, e.g. RnnTokenizer/ud/tokenizer-eng-UD_English-EWT.pt.
All these units accept a data parameter naming the sentence boundaries data
they use or produce (default: SentenceBoundaries).
RnnTokenizer¶
Class: RnnTokenizer
Role: splits the text into tokens and sentences with a neural model. It is
the first unit of the deepud and deeplima pipelines.
Inputs: an AnalysisContent containing the initial text.
| Parameter | Description |
|---|---|
model_prefix |
Model file name without extension, in RnnTokenizer/ud/. Default configuration: tokenizer-$udlang |
Effects: creates the AnalysisGraph (one vertex per token) and the
sentence boundaries.
ConlluReader¶
Class: ConlluReader
Role: reads already tokenized text in CoNLL-U format instead of tokenizing
raw text. It is the first unit of the deepud-pretok pipeline.
| Parameter | Description |
|---|---|
boundaryMicro |
Category used for sentence boundaries (SENT in the default configuration) |
Effects: creates the AnalysisGraph and the sentence boundaries from the
CoNLL-U tokens and sentences.
RnnTokensAnalyzer¶
Class: RnnTokensAnalyzer
Role: assigns to each token its universal part of speech and morphological features and, when a lemmatizer model is available, its lemma.
Preconditions: the AnalysisGraph and the sentence boundaries exist.
| Parameter | Description |
|---|---|
tagger_model_prefix |
Tagger model name in RnnTagger/ud/. Default configuration: tagger-$udlang |
lemmatizer_model_prefix |
Lemmatizer model name in RnnLemmatizer/ud/. Default configuration: lemmatizer-$udlang. An optional .dic file with the same name is used as a lemma cache |
Effects: creates the PosGraph and the annotation data; makes the
tagging results available to the dependency parser.
If no lemmatizer model is installed for the treebank, the lemmas are left empty.
RnnDependencyParser¶
Class: RnnDependencyParser
Role: computes the labeled dependency tree of each sentence (heads and relations), using the features predicted by the tagger.
Preconditions: RnnTokensAnalyzer has been run.
| Parameter | Description |
|---|---|
dependency_parser_model_prefix |
Parser model name in RnnDependencyParser/ud/. Default configuration: dependencyParser-$udlang |
tagger_model_prefix |
Name of the tagger model whose outputs feed the parser. Default configuration: tagger-$udlang |
Effects: creates the SyntacticData holding the dependency relations,
written in the HEAD and DEPREL columns by conllDumper.
If no parser model is installed for the treebank, the unit does nothing.
RnnNER¶
Class: RnnNER
Role: neural named entity recognition.
| Parameter | Description |
|---|---|
model_prefix |
Model name. Default configuration: ner-$udlang |
Note
This unit is defined in the configuration but not used by the default pipelines: no NER models are published yet. See Named entity recognition.