Using LIMA from Python¶
The aymara.lima module gives access to LIMA's main features with an API
inspired by spaCy. Install it with
pip.
The PyPI release lags behind the source code
The current PyPI release, aymara 0.5.0b6 (April 2024), was built before
LIMA's neural modules moved from TensorFlow to libtorch. It ships its own
model installers (lima_models for the legacy models and
deeplima_models for the libtorch ones) and its own ud-eng/ud-fra
pipelines. A new release built from the current code is in preparation;
until then, the Docker image or a
source build give you the latest version.
Analyzing text¶
Create a Lima analyzer once (initialization loads the configuration and the
models, which takes some time) and call it on texts:
import aymara.lima
nlp = aymara.lima.Lima("ud-eng", pipes="deepud",
meta={"udlang": "eng-UD_English-EWT"})
doc = nlp("The author wrote a novel. It was a success.")
With the 0.5.0b6 PyPI release, aymara.lima.Lima("ud-eng") is enough: that
version selects the English models by itself.
The constructor takes:
langs: comma-separated list of languages to initialize (see Language models);pipes: comma-separated list of pipelines to initialize (see Pipelines);meta: metadata, e.g.{"udlang": "eng-UD_English-EWT"}to select the treebank of the models;user_config_pathanduser_resources_path: directories searched before the configuration and resources of the package (see Configuring LIMA).
When several languages or pipelines are initialized, choose them at analysis
time with the lang and pipeline arguments, e.g.
nlp(text, lang="eng", pipeline="main"). Without them, the first initialized
language and pipeline are used. The models of the neural pipelines are loaded
when the analyzer is created, so create one analyzer per treebank.
Documents, sentences and tokens¶
A Doc is a sequence of Tokens. Sentences and named entities are Spans,
contiguous sequences of tokens.
for token in doc:
print(token.i, token.text, token.lemma, token.pos, token.dep, token.head)
for sentence in doc.sents:
print(sentence)
for entity in doc.ents:
print(entity.text, entity.label)
Printing a Doc with repr() gives its CoNLL-U representation:
>>> print(repr(nlp("Hello, World!")))
1 Hello hello INTJ _ _ 0 root _ Pos=0|Len=5
2 , , PUNCT _ _ 1 punct _ Pos=5|Len=1
3 World World PROPN _ Number:Sing 1 vocative _ Pos=7|Len=5
4 ! ! PUNCT _ _ 1 punct _ Pos=12|Len=1
analyzeText() returns the raw CoNLL-U string instead of a Doc:
Customizing the configuration¶
Lima.export_system_conf() copies the configuration files of the package to a
directory, so that you can edit them and pass that directory as
user_config_path. Lima.add_pipeline_unit() adds a process unit to a
pipeline at run time.
The full API is described in the Python API reference.
The lima command¶
The package also installs a lima command analyzing files and writing CoNLL-U
on the standard output:
Run lima --help for its options.
Note
Some messages may be displayed on the console when the analyzer is created. If you get a valid result, you can ignore them.