Skip to content

Python API

This reference is generated from the docstrings of the aymara.lima module. For an introduction, see Using LIMA from Python.

The LIMA python bindings.

This python API gives access to the major features of the LIMA linguistic analyzer. To make it easier to handle, it largely reproduces that of spaCy, including parts of the documentation. See the GitHub project for spaCy's copyright notice.

Example::

import aymara.lima
nlp = aymara.lima.Lima()
doc = nlp("Mr. Best flew to New York on Saturday morning.")
print(doc)

Classes:

Doc
Lima
Span
Token

Lima

Lima(langs: str = 'fre,eng', pipes: str = 'main,deepud', user_config_path: str = '', user_resources_path: str = '', meta: Dict[str, str] = {})

A text-processing pipeline

Usually you’ll load this once per process as nlp and pass the instance around your application. The Lima class is a wrapper around the LimaAnalyzer class which is itself a binding around the C++ classes necessary to analyze text.

Example::

        import aymara.lima
        nlp = aymara.lima.Lima()
        doc = nlp("Give it back! He pleaded.")
        print(doc)

Initialize the Lima analyzer

Parameters:

Name Type Description Default
langs str

a comma-separated list of language trigrams to initialize (Default value = "fre,eng")

'fre,eng'
pipes str

a comma-separated list of Lima pipelines to analyze (Default value = "main,deepud")

'main,deepud'
user_config_path str

a path where Lima configuration files will be searched for. This allows to override default configurations. (Default value = an empty string)

''
user_resources_path str

a path where Lima resource files will be searched for. This allows to override default configurations (Default value = an empty string)

''
meta Dict[str, str]

a list of named metadata values that will be used for each analysis.They can be completed or overriden at analysis time (Default value = an empty dictionary)

{}

langs instance-attribute

langs = langs.split(',')

pipes instance-attribute

pipes = pipes.split(',')

__call__

__call__(text: str, lang: str = None, pipeline: str = None, meta: Dict[str, str] = {}) -> Doc

Just 'call' your Lima instance to analyze the given text in the given language. The lang language must have been initialized when instantiating this object.

Example::

        import aymara.lima
        nlp = aymara.lima.Lima()
        doc = nlp("Give it back! He pleaded.")
        print(doc)

Parameters:

Name Type Description Default
text str

the text to analyze

required
lang str

the language of the text. If none, will backup to the first element of the langs member or to eng if empty (Default value = None). Its value can be one of the three historic pre-Universal Dependencies languages ("eng", "fre" and "por") or the value "ud". In the latter case, the meta parameter must include a pair "udlang":"", e.g.: "udlang":"fra".

None
pipeline str

the Lima pipeline to use for analysis. If none, will backup to the first element of the pipelines member or to main if empty (Default value = None).

None
meta Dict[str, str]

a dict of named metadata values (Default value = an empty dictionary).

{}

Returns:

Type Description
Doc

a Doc object representing the result of the analysis.

analyzeText

analyzeText(text: str, lang: str = None, pipeline: str = None, meta: Dict[str, str] = {}) -> str

Analyze the given text in the given language. The lang language must have been initialized when instantiating this object.

Example::

        import aymara.lima
        nlp = aymara.lima.Lima()
        result = nlp.analyzeText("Give it back! He pleaded.")
        print(result)

Parameters:

Name Type Description Default
text str

the text to analyze

required
lang str

the language of the text. If none, will backup to the first element of the langs member or to eng if empty (Default value = None).

None
pipeline str

the Lima pipeline to use for analysis. If none, will backup to the first element of the pipelines member or to main if empty (Default value = None).

None
meta Dict[str, str]

a dict of named metadata values (Default value = an empty dictionary).

{}

Returns:

Type Description
str

the content of the text written by the text dumper of Lima if any. An empty string otherwise

add_pipeline_unit

add_pipeline_unit(pipeline: str, language: str, group: str)

Add a pipeline unit defined by the group json string to the given pipeline of the given language. Media and pipeline must be already instantiated (defined in the LIMA confiuration files). See LIMA documentation for the content of the group. The new pipeline unit is added at the end of the pipeline.

Example::

        import aymara.lima
        nlp = aymara.lima.Lima(
            "ud-eng", "empty",
            meta={"udlang": "eng-UD_English-EWT"})
        group = {
            "name": "rnntokenizer",
            "class": "RnnTokenizer",
            "model_prefix": "tokenizer-$udlang"
        }
        nlp.add_pipeline_unit("empty", "ud-eng", json.dumps(group)
        group = {
            "name": "rnntokensanalyzer",
            "class": "RnnTokensAnalyzer",
            "tagger_model_prefix": "tagger-$udlang",
            "lemmatizer_model_prefix": "lemmatizer-$udlang"
        }
        nlp.add_pipeline_unit("empty", "ud-eng", json.dumps(group)
        group = {
            "name": "conllDumper",
            "class": "ConllDumper",
            "handler": "simpleStreamHandler"
            "fakeDependencyGraph": "false"
        }
        nlp.add_pipeline_unit("empty", "ud-eng", json.dumps(group)

        result = nlp.analyzeText("Give it back! He pleaded.")
        print(result)

Parameters:

Name Type Description Default
pipeline str

the pipeline to edit. Use "empty" to change the initialy empty pipeline. It is currently not possible to create a pipeline from scratch. You must add to a pipeline defined in confiuration files

required
language str

the language to edit

required
group str

a string dump of a json object representing the configuration of the pipeline unit, as defined in XML in LIMA's configuration files.

required

Returns:

Type Description
bool

True if successful and False otherwise

export_system_conf staticmethod

export_system_conf(dir: Path = None, lang: str = None) -> bool

Export LIMA configuration files from the module system path to the given dir in order to be able to easily change configuration files.

If lang is given, only the configuration files concerning this language are exported (NOT IMPLEMENTED).

Use this function to initiate a user configuration. For LIMA to take into account the configuration in the new path, you will have to add it in front of the LIMA_CONF environment variable (or define it if it does not exist).

Please refer to the LIMA documentation <https://aymara.github.io/lima/usage/configuration/>_ for how to configure the analysis:

Example::

import aymara.lima
aymara.lima.Lima.export_system_conf("~/MyLima")

Parameters:

Name Type Description Default
dir Path

the directory were to export the configuration (Default value = None)

None
lang str

the language whose configuration must be exported. If None, the whole configuration is exported (Default value = None)

None

Returns:

Type Description
bool

True if the configuration is correctly exported and False otherwise.

get_system_paths staticmethod

get_system_paths() -> Tuple[str, str]

Get the system configuration and resoures paths.

Example::

import aymara.lima
aymara.lima.Lima.get_system_paths()

Returns:

Type Description
Tuple[str, str]

the colon (; under Windows) -separated list of the paths that are searched by LIMA to load its configuration files and linguistic resources. This function is useful to understand from which dirs data are loaded to debug configuration errors. It can also be used to know where to put or edit files.

Doc

Doc(doc: Doc)

A document.

This is mainly an iterable of tokens.

Example::

import aymara.lima
nlp = aymara.lima.Lima()
doc = nlp("Give it back! He pleaded.")

TODO Some parts of the API are still not implemented:

compounds   The compounds found into the document text by the
    CompoundsBuilderFromSyntacticData LIMA pipeline unit
    List[Compound]

text class-attribute instance-attribute

text = property(fget=lambda self: self.limadoc.text(), doc='The original text.\n:type: str\n')

sents class-attribute instance-attribute

sents = property(fget=lambda self: _SentencesIterator(self), doc='    Iterate over the sentences in the document.\n        This property is only available when sentence boundaries have been set on the\n        document by the pipeline. It will raise an error otherwise.\n        Example::\nsents = list(doc.sents)\n          import aymara.lima\n          nlp = aymara.lima.Lima()\n          doc = nlp("This is a sentence. Here\'s another...")\n          sents = list(doc.sents)\n          assert len(sents) == 2\n          assert [s.root.text for s in sents] == ["is", "\'s"]\n\n        :yields:\tSentences in the document.\n        :type: Span\n')

lang class-attribute instance-attribute

lang = property(fget=lambda self: self.limadoc.language(), doc='Language of the document.')

ents class-attribute instance-attribute

ents = property(fget=lambda self: _DocEntitiesIterator(self), doc='Iterate over the entites in the document. Returns an iterator yieldingnamed entity Span objects.\n        Example::\n\n          import aymara.lima\n          nlp = aymara.lima.Lima()\n          doc = nlp("John Doe lives in New York")\n          ents = list(doc.ents)\n          assert ents[0].label == "Person.PERSON"\n          assert ents[0].text == "John Doe"\n\n        :yields:\tEntities in the document.\n        :type: Span\n')

__len__

__len__() -> int

'Returns the number of tokens of this document

Returns:

Type Description
int

the number of tokens of this document.

__getitem__

__getitem__(i: Union[int, slice]) -> Union[Token, Span]

Returns the token at position i or a contiguous slice of tokens.

Example::

doc = nlp("Give it back! He pleaded.")
assert doc[0].text == "Give"
assert doc[-1].text == "."
span = doc[1:3]
assert span.text == "it back"

Parameters:

Name Type Description Default
i Union[int, slice]

a position i or a contiguous slice of token to retrieve

required

Returns:

Type Description
Union[int, slice]

the token at position i or a contiguous slice of tokens.

Span

Span(doc, start: int, end: int, label: str = '')

Represents a continuous span of tokens in a Doc.

TODO Some parts of the API are still not implemented

ents    The named entities that fall completely within the span. Returns a tuple of
    Span objects.
    Example::

        import aymara.lima
        nlp = aymara.lima.Lima()
        doc = nlp("Mr. Best flew to New York on Saturday morning.")
        span = doc[0:6]
        ents = list(span.ents)
        assert ents[0].label == 346
        assert ents[0].label_ == "PERSON"
        assert ents[0].text == "Mr. Best"

    Name        Description
    RETURNS     Entities in the span, one Span per entity.
    Tuple[Span, …]

sent    The sentence span that this span is a part of.
    This property is only available when sentence boundaries have been set on the
    document by the pipeline. It will raise an error otherwise.

    If the span happens to cross sentence boundaries, only the first sentence will be returned. If it is required that the sentence always includes the full span, the result can be adjusted as such:

    sent = span.sent
    sent = doc[sent.start : max(sent.end, span.end)]

    Example::

        import aymara.lima
        nlp = aymara.lima.Lima()
        doc = nlp("Give it back! He pleaded.")
        span = doc[1:3]
        assert span.sent.text == "Give it back!"

Span

sents   Returns a generator over the sentences the span belongs to.
    This property is only available when sentence boundaries have been set on the
    document by the pipeline. It will raise an error otherwise.

    If the span happens to cross sentence boundaries, all sentences the span overlaps with will be returned.
    Example::

        import aymara.lima
        nlp = aymara.lima.Lima()
        doc = nlp("Give it back! He pleaded.")
        span = doc[2:4]
        assert len(span.sents) == 2

Iterable[Span]

Constructor of a Span

Parameters:

Name Type Description Default
doc Doc

The document on which is built the span.

required
start int

The id of the fist token of the span.

required
label str

A label to attach to the span, e.g. for named entities.

''

text class-attribute instance-attribute

text = property(fget=lambda self: self._doc.text[self._doc[self._start].idx:self._doc[self._end - 1].idx + len(self._doc[self._end - 1])], doc='A string representation of the span text.')

doc class-attribute instance-attribute

doc = property(fget=lambda self: self._doc, doc='The parent document.')

start class-attribute instance-attribute

start = property(fget=lambda self: self._start, doc='The token offset for the start of the span.')

end class-attribute instance-attribute

end = property(fget=lambda self: self._end, doc='The token offset for the end of the span.')

start_char class-attribute instance-attribute

start_char = property(fget=lambda self: self[0].idx, doc='The character offset for the start of the span.')

end_char class-attribute instance-attribute

end_char = property(fget=lambda self: self[-1].idx + len(self[-1]), doc='The character offset for the end of the span.')

label class-attribute instance-attribute

label = property(fget=lambda self: self._label, doc='A label to attach to the span, e.g. for named entities.')

__len__

__len__() -> int

Returns the number of tokens of this span

Example::

import aymara.lima
nlp = aymara.lima.Lima()
doc = nlp("Give it back! He pleaded.")
span = doc[1:4]
assert len(span) == 3

Returns:

Type Description
int

the number of tokens in this span

__getitem__

__getitem__(i: Union[int, slice])

Returns either the Token at position i in the span or the subspan defined by the slice i.

Example::

import aymara.lima
nlp = aymara.lima.Lima()
doc = nlp("Give it back! He pleaded.")
span = doc[1:4]
assert span[1].text == "back"
assert span[1:3].text == "back!"

Parameters:

Name Type Description Default
i Union[int, slice]

the position in the span of the item to retrieve or a slice defining the subspan to retriev.

required

Returns:

Type Description
Union[int, slice]

either the Token at position i in the span or the subspan defined by the slice i.

Token

Token(token: Token)

A token

TODO Some parts of the API are still not implemented

sent     The sentence span that this token is a part of.
Span

lang    Language of the parent document’s vocabulary.
str

Token's constructor

Parameters:

Name Type Description Default
token Token

the C++ binding Token class

required

text class-attribute instance-attribute

text = property(fget=lambda self: self.token.text, doc='The original text of the token.')

i class-attribute instance-attribute

i = property(fget=lambda self: self.token.i + 1, doc='The index of this token in its parent document.')

lemma class-attribute instance-attribute

lemma = property(fget=lambda self: self.token.lemma, doc='The token lemma.')

pos class-attribute instance-attribute

pos = property(fget=lambda self: self.token.tag, doc='Coarse-grained part-of-speech from the Universal POS tag set.')

head class-attribute instance-attribute

head = property(fget=lambda self: self.token.head, doc='The syntactic parent, or “governor”, of this token.')

dep class-attribute instance-attribute

dep = property(fget=lambda self: self.token.dep, doc='Syntactic dependency relation.')

idx class-attribute instance-attribute

idx = property(fget=lambda self: self.token.pos - 1, doc='Position of this token in its document text.')

features class-attribute instance-attribute

features = property(fget=lambda self: {} if self.token.features == '_' else dict(x.split('=') for x in self.token.features.split('|')), doc='Morphlogical features of this token .')

ent_type class-attribute instance-attribute

ent_type = property(fget=lambda self: self.token.neType, doc='Named entity type.')

ent_iob class-attribute instance-attribute

ent_iob = property(fget=lambda self: self.token.neIOB, doc='IOB code of named entity tag. “B” means the token begins an entity, “I” means it is inside an entity, “O” means it is outside an entity, and "" means no entity tag is set.')

t_status class-attribute instance-attribute

t_status = property(fget=lambda self: self.token.tStatus, doc='The tokenization status of this token. Can also be explored with the is_* properties. The possible values are::\n\n  t_alphanumeric\n  t_abbrev\n  t_acronym\n  t_capital\n  t_capital_1st\n  t_capital_small\n  t_cardinal_roman\n  t_comma_number\n  t_dot_number\n  t_fraction\n  t_integer\n  t_ordinal_integer\n  t_ordinal_roman\n  t_sentence_brk\n  t_small\n  t_word_brk\n\n')

is_alpha class-attribute instance-attribute

is_alpha = property(fget=lambda self: self.token.tStatus in ['t_alphanumeric', 't_capital', 't_capital_1st', 't_capital_small', 't_small'], doc='Does the token consist of alphabetic characters? Equivalent to token.text.isalpha().')

is_digit class-attribute instance-attribute

is_digit = property(fget=lambda self: self.token.tStatus == 't_integer', doc='Does the token consist of digits? Equivalent to token.text.isdigit().')

is_lower class-attribute instance-attribute

is_lower = property(fget=lambda self: self.token.text.islower(), doc='Is the token in lowercase? Equivalent to token.text.islower().')

is_upper class-attribute instance-attribute

is_upper = property(fget=lambda self: self.token.text.isupper(), doc='Is the token in lowercase? Equivalent to token.text.isupper().')

is_punct class-attribute instance-attribute

is_punct = property(fget=lambda self: self.token.tStatus in ['t_sentence_brk', 't_word_brk'], doc='Is the token punctuation?')

is_sent_start class-attribute instance-attribute

is_sent_start = property(fget=lambda self: self.token.i == 0, doc='Does the token start a sentence? bool or None if unknown. Default value = True for the first token in the Doc.\nTODO: implement for sentences other than the first one.')

is_sent_end class-attribute instance-attribute

is_sent_end = property(fget=lambda self: self.token.tStatus == 't_sentence_brk', doc='Does the token end a sentence? bool or None if unknown.')

is_space class-attribute instance-attribute

is_space = property(fget=lambda self: self.token.text.isspace(), doc='Does the token consist of whitespace characters? Equivalent to token.text.isspace(). Should always be False in LIMA as there is no space tokens')

is_bracket class-attribute instance-attribute

is_bracket = property(fget=lambda self: self.token.text in '()[]{}', doc='Is the token a bracket?')

is_quote class-attribute instance-attribute

is_quote = property(fget=lambda self: self.token.text in '"\'«»`', doc='Is the token a quotation mark?')

__len__

__len__() -> int

Return the length of the token in UTF-8 code points

Returns:

Type Description
int

the length of the token

LimaInternalError

Bases: Exception