Skip to content

Using LIMA from Python

The aymara.lima module gives access to LIMA's main features with an API inspired by spaCy. Install it with pip.

The PyPI release lags behind the source code

The current PyPI release, aymara 0.5.0b6 (April 2024), was built before LIMA's neural modules moved from TensorFlow to libtorch. It ships its own model installers (lima_models for the legacy models and deeplima_models for the libtorch ones) and its own ud-eng/ud-fra pipelines. A new release built from the current code is in preparation; until then, the Docker image or a source build give you the latest version.

Analyzing text

Create a Lima analyzer once (initialization loads the configuration and the models, which takes some time) and call it on texts:

import aymara.lima

nlp = aymara.lima.Lima("ud-eng", pipes="deepud",
                       meta={"udlang": "eng-UD_English-EWT"})
doc = nlp("The author wrote a novel. It was a success.")

With the 0.5.0b6 PyPI release, aymara.lima.Lima("ud-eng") is enough: that version selects the English models by itself.

The constructor takes:

  • langs: comma-separated list of languages to initialize (see Language models);
  • pipes: comma-separated list of pipelines to initialize (see Pipelines);
  • meta: metadata, e.g. {"udlang": "eng-UD_English-EWT"} to select the treebank of the models;
  • user_config_path and user_resources_path: directories searched before the configuration and resources of the package (see Configuring LIMA).

When several languages or pipelines are initialized, choose them at analysis time with the lang and pipeline arguments, e.g. nlp(text, lang="eng", pipeline="main"). Without them, the first initialized language and pipeline are used. The models of the neural pipelines are loaded when the analyzer is created, so create one analyzer per treebank.

Documents, sentences and tokens

A Doc is a sequence of Tokens. Sentences and named entities are Spans, contiguous sequences of tokens.

for token in doc:
    print(token.i, token.text, token.lemma, token.pos, token.dep, token.head)

for sentence in doc.sents:
    print(sentence)

for entity in doc.ents:
    print(entity.text, entity.label)

Printing a Doc with repr() gives its CoNLL-U representation:

>>> print(repr(nlp("Hello, World!")))
1       Hello   hello   INTJ    _       _               0       root    _       Pos=0|Len=5
2       ,       ,       PUNCT   _       _               1       punct   _       Pos=5|Len=1
3       World   World   PROPN   _       Number:Sing     1       vocative        _       Pos=7|Len=5
4       !       !       PUNCT   _       _               1       punct   _       Pos=12|Len=1

analyzeText() returns the raw CoNLL-U string instead of a Doc:

print(nlp.analyzeText("The author wrote a novel.", lang="ud-eng"))

Customizing the configuration

Lima.export_system_conf() copies the configuration files of the package to a directory, so that you can edit them and pass that directory as user_config_path. Lima.add_pipeline_unit() adds a process unit to a pipeline at run time.

The full API is described in the Python API reference.

The lima command

The package also installs a lima command analyzing files and writing CoNLL-U on the standard output:

lima my-text.txt

Run lima --help for its options.

Note

Some messages may be displayed on the console when the analyzer is created. If you get a valid result, you can ignore them.