Universal Dependencies analysis¶
LIMA's neural pipelines analyze text following the
Universal Dependencies (UD) guidelines,
with models trained on UD treebanks for more than 60 languages. They are built
on the in-tree deeplima library and libtorch.
| Pipeline | Description | Input | Output |
|---|---|---|---|
deepud |
Full analysis, including tokenization and sentence splitting | Plain text | CoNLL-U |
deepud-pretok |
Analysis of already tokenized text | CoNLL-U | CoNLL-U |
deeplima |
Like deepud, without dependency parsing |
Plain text | CoNLL-U |
Installation¶
Install the models of the treebanks you need:
lima_models.py -i # available languages and treebanks
lima_models.py -l fra # best French treebank
lima_models.py -l fra-UD_French-GSD # or a specific one
Usage¶
The examples below use fra-UD_French-GSD; use the treebank identifier
printed by lima_models.py for the one you installed.
# Raw text
analyzeText -l fra-UD_French-GSD -p deepud my-text.txt
# Pre-tokenized text
analyzeText -l fra-UD_French-GSD -p deepud-pretok my-corpus.conllu
See Language models for the other ways to select the treebank.
Processing units¶
deepud |
deepud-pretok |
deeplima |
|
|---|---|---|---|
| Input | |||
RnnTokenizer: tokenization and sentence splitting |
✓ | ✓ | |
conllureader: reads CoNLL-U |
✓ | ||
| Analysis | |||
RnnTokensAnalyzer: UPOS, features, lemmas |
✓ | ✓ | ✓ |
RnnDependencyParser: heads and relations |
✓ | ✓ | |
| Output | |||
conllDumper |
✓ | ✓ | ✓ |
The up-to-date definitions of these pipelines are in
lima_linguisticprocessing/conf/lima-lp-ud.xml.
The units are described in the neural units reference.
Each unit uses the model of the selected treebank. Lemmatizer and dependency parser models are not available for every treebank yet: when one is missing, the corresponding columns are left empty and the rest of the analysis runs normally.
Performance¶
The scores of the models for every treebank are reported in the lima-models evaluation page.
Limitations¶
- Memory: depending on the language model and the size of its word embeddings, analysis can require several GB of RAM.
- XPOS: the XPOS column and the
TypoandAbbrfeatures are not produced. - Quantized embeddings: the distributed fastText embeddings are quantized to reduce their size from about 7 GB to about 600 MB per language, which slightly lowers accuracy.