Python API¶
This reference is generated from the docstrings of the
aymara.lima
module. For an introduction, see Using LIMA from Python.
The LIMA python bindings.
This python API gives access to the major features of the LIMA linguistic analyzer. To make it easier to handle, it largely reproduces that of spaCy, including parts of the documentation. See the GitHub project for spaCy's copyright notice.
Example::
import aymara.lima
nlp = aymara.lima.Lima()
doc = nlp("Mr. Best flew to New York on Saturday morning.")
print(doc)
Classes:
Doc
Lima
Span
Token
Lima ¶
Lima(langs: str = 'fre,eng', pipes: str = 'main,deepud', user_config_path: str = '', user_resources_path: str = '', meta: Dict[str, str] = {})
A text-processing pipeline
Usually you’ll load this once per process as nlp and pass the instance around your application. The Lima class is a wrapper around the LimaAnalyzer class which is itself a binding around the C++ classes necessary to analyze text.
Example::
import aymara.lima
nlp = aymara.lima.Lima()
doc = nlp("Give it back! He pleaded.")
print(doc)
Initialize the Lima analyzer
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
langs
|
str
|
a comma-separated list of language trigrams to initialize (Default value = "fre,eng") |
'fre,eng'
|
pipes
|
str
|
a comma-separated list of Lima pipelines to analyze (Default value = "main,deepud") |
'main,deepud'
|
user_config_path
|
str
|
a path where Lima configuration files will be searched for. This allows to override default configurations. (Default value = an empty string) |
''
|
user_resources_path
|
str
|
a path where Lima resource files will be searched for. This allows to override default configurations (Default value = an empty string) |
''
|
meta
|
Dict[str, str]
|
a list of named metadata values that will be used for each analysis.They can be completed or overriden at analysis time (Default value = an empty dictionary) |
{}
|
__call__ ¶
Just 'call' your Lima instance to analyze the given text in the given language. The lang language must have been initialized when instantiating this object.
Example::
import aymara.lima
nlp = aymara.lima.Lima()
doc = nlp("Give it back! He pleaded.")
print(doc)
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
text
|
str
|
the text to analyze |
required |
lang
|
str
|
the language of the text. If none, will backup to the first element
of the langs member or to eng if empty (Default value = |
None
|
pipeline
|
str
|
the Lima pipeline to use for analysis. If none, will backup to
the first element of the pipelines member or to main if empty (Default
value = |
None
|
meta
|
Dict[str, str]
|
a dict of named metadata values (Default value = an empty dictionary). |
{}
|
Returns:
| Type | Description |
|---|---|
Doc
|
a Doc object representing the result of the analysis. |
analyzeText ¶
Analyze the given text in the given language. The lang language must have been initialized when instantiating this object.
Example::
import aymara.lima
nlp = aymara.lima.Lima()
result = nlp.analyzeText("Give it back! He pleaded.")
print(result)
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
text
|
str
|
the text to analyze |
required |
lang
|
str
|
the language of the text. If none, will backup to the first element
of the langs member or to eng if empty (Default value = |
None
|
pipeline
|
str
|
the Lima pipeline to use for analysis. If none, will backup to
the first element of the pipelines member or to main if empty (Default
value = |
None
|
meta
|
Dict[str, str]
|
a dict of named metadata values (Default value = an empty dictionary). |
{}
|
Returns:
| Type | Description |
|---|---|
str
|
the content of the text written by the text dumper of Lima if any. An empty string otherwise |
add_pipeline_unit ¶
Add a pipeline unit defined by the group json string to the given pipeline of the given language. Media and pipeline must be already instantiated (defined in the LIMA confiuration files). See LIMA documentation for the content of the group. The new pipeline unit is added at the end of the pipeline.
Example::
import aymara.lima
nlp = aymara.lima.Lima(
"ud-eng", "empty",
meta={"udlang": "eng-UD_English-EWT"})
group = {
"name": "rnntokenizer",
"class": "RnnTokenizer",
"model_prefix": "tokenizer-$udlang"
}
nlp.add_pipeline_unit("empty", "ud-eng", json.dumps(group)
group = {
"name": "rnntokensanalyzer",
"class": "RnnTokensAnalyzer",
"tagger_model_prefix": "tagger-$udlang",
"lemmatizer_model_prefix": "lemmatizer-$udlang"
}
nlp.add_pipeline_unit("empty", "ud-eng", json.dumps(group)
group = {
"name": "conllDumper",
"class": "ConllDumper",
"handler": "simpleStreamHandler"
"fakeDependencyGraph": "false"
}
nlp.add_pipeline_unit("empty", "ud-eng", json.dumps(group)
result = nlp.analyzeText("Give it back! He pleaded.")
print(result)
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
pipeline
|
str
|
the pipeline to edit. Use "empty" to change the initialy empty pipeline. It is currently not possible to create a pipeline from scratch. You must add to a pipeline defined in confiuration files |
required |
language
|
str
|
the language to edit |
required |
group
|
str
|
a string dump of a json object representing the configuration of the pipeline unit, as defined in XML in LIMA's configuration files. |
required |
Returns:
| Type | Description |
|---|---|
bool
|
True if successful and False otherwise |
export_system_conf
staticmethod
¶
Export LIMA configuration files from the module system path to the given dir in order to be able to easily change configuration files.
If lang is given, only the configuration files concerning this language are exported (NOT IMPLEMENTED).
Use this function to initiate a user configuration. For LIMA to take into account the configuration in the new path, you will have to add it in front of the LIMA_CONF environment variable (or define it if it does not exist).
Please refer to the
LIMA documentation <https://aymara.github.io/lima/usage/configuration/>_
for how to configure the analysis:
Example::
import aymara.lima
aymara.lima.Lima.export_system_conf("~/MyLima")
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
dir
|
Path
|
the directory were to export the configuration (Default value = None) |
None
|
lang
|
str
|
the language whose configuration must be exported. If |
None
|
Returns:
| Type | Description |
|---|---|
bool
|
True if the configuration is correctly exported and False otherwise. |
get_system_paths
staticmethod
¶
Get the system configuration and resoures paths.
Example::
import aymara.lima
aymara.lima.Lima.get_system_paths()
Returns:
| Type | Description |
|---|---|
Tuple[str, str]
|
the colon (; under Windows) -separated list of the paths that are searched by LIMA to load its configuration files and linguistic resources. This function is useful to understand from which dirs data are loaded to debug configuration errors. It can also be used to know where to put or edit files. |
Doc ¶
A document.
This is mainly an iterable of tokens.
Example::
import aymara.lima
nlp = aymara.lima.Lima()
doc = nlp("Give it back! He pleaded.")
TODO Some parts of the API are still not implemented:
compounds The compounds found into the document text by the
CompoundsBuilderFromSyntacticData LIMA pipeline unit
List[Compound]
text
class-attribute
instance-attribute
¶
sents
class-attribute
instance-attribute
¶
sents = property(fget=lambda self: _SentencesIterator(self), doc=' Iterate over the sentences in the document.\n This property is only available when sentence boundaries have been set on the\n document by the pipeline. It will raise an error otherwise.\n Example::\nsents = list(doc.sents)\n import aymara.lima\n nlp = aymara.lima.Lima()\n doc = nlp("This is a sentence. Here\'s another...")\n sents = list(doc.sents)\n assert len(sents) == 2\n assert [s.root.text for s in sents] == ["is", "\'s"]\n\n :yields:\tSentences in the document.\n :type: Span\n')
lang
class-attribute
instance-attribute
¶
ents
class-attribute
instance-attribute
¶
ents = property(fget=lambda self: _DocEntitiesIterator(self), doc='Iterate over the entites in the document. Returns an iterator yieldingnamed entity Span objects.\n Example::\n\n import aymara.lima\n nlp = aymara.lima.Lima()\n doc = nlp("John Doe lives in New York")\n ents = list(doc.ents)\n assert ents[0].label == "Person.PERSON"\n assert ents[0].text == "John Doe"\n\n :yields:\tEntities in the document.\n :type: Span\n')
__len__ ¶
'Returns the number of tokens of this document
Returns:
| Type | Description |
|---|---|
int
|
the number of tokens of this document. |
__getitem__ ¶
Returns the token at position i or a contiguous slice of tokens.
Example::
doc = nlp("Give it back! He pleaded.")
assert doc[0].text == "Give"
assert doc[-1].text == "."
span = doc[1:3]
assert span.text == "it back"
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
i
|
Union[int, slice]
|
a position i or a contiguous slice of token to retrieve |
required |
Returns:
| Type | Description |
|---|---|
Union[int, slice]
|
the token at position i or a contiguous slice of tokens. |
Span ¶
Represents a continuous span of tokens in a Doc.
TODO Some parts of the API are still not implemented
ents The named entities that fall completely within the span. Returns a tuple of
Span objects.
Example::
import aymara.lima
nlp = aymara.lima.Lima()
doc = nlp("Mr. Best flew to New York on Saturday morning.")
span = doc[0:6]
ents = list(span.ents)
assert ents[0].label == 346
assert ents[0].label_ == "PERSON"
assert ents[0].text == "Mr. Best"
Name Description
RETURNS Entities in the span, one Span per entity.
Tuple[Span, …]
sent The sentence span that this span is a part of.
This property is only available when sentence boundaries have been set on the
document by the pipeline. It will raise an error otherwise.
If the span happens to cross sentence boundaries, only the first sentence will be returned. If it is required that the sentence always includes the full span, the result can be adjusted as such:
sent = span.sent
sent = doc[sent.start : max(sent.end, span.end)]
Example::
import aymara.lima
nlp = aymara.lima.Lima()
doc = nlp("Give it back! He pleaded.")
span = doc[1:3]
assert span.sent.text == "Give it back!"
Span
sents Returns a generator over the sentences the span belongs to.
This property is only available when sentence boundaries have been set on the
document by the pipeline. It will raise an error otherwise.
If the span happens to cross sentence boundaries, all sentences the span overlaps with will be returned.
Example::
import aymara.lima
nlp = aymara.lima.Lima()
doc = nlp("Give it back! He pleaded.")
span = doc[2:4]
assert len(span.sents) == 2
Iterable[Span]
Constructor of a Span
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
doc
|
Doc
|
The document on which is built the span. |
required |
start
|
int
|
The id of the fist token of the span. |
required |
label
|
str
|
A label to attach to the span, e.g. for named entities. |
''
|
text
class-attribute
instance-attribute
¶
text = property(fget=lambda self: self._doc.text[self._doc[self._start].idx:self._doc[self._end - 1].idx + len(self._doc[self._end - 1])], doc='A string representation of the span text.')
doc
class-attribute
instance-attribute
¶
start
class-attribute
instance-attribute
¶
end
class-attribute
instance-attribute
¶
start_char
class-attribute
instance-attribute
¶
start_char = property(fget=lambda self: self[0].idx, doc='The character offset for the start of the span.')
end_char
class-attribute
instance-attribute
¶
end_char = property(fget=lambda self: self[-1].idx + len(self[-1]), doc='The character offset for the end of the span.')
label
class-attribute
instance-attribute
¶
label = property(fget=lambda self: self._label, doc='A label to attach to the span, e.g. for named entities.')
__len__ ¶
Returns the number of tokens of this span
Example::
import aymara.lima
nlp = aymara.lima.Lima()
doc = nlp("Give it back! He pleaded.")
span = doc[1:4]
assert len(span) == 3
Returns:
| Type | Description |
|---|---|
int
|
the number of tokens in this span |
__getitem__ ¶
Returns either the Token at position i in the span or the subspan defined by the slice i.
Example::
import aymara.lima
nlp = aymara.lima.Lima()
doc = nlp("Give it back! He pleaded.")
span = doc[1:4]
assert span[1].text == "back"
assert span[1:3].text == "back!"
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
i
|
Union[int, slice]
|
the position in the span of the item to retrieve or a slice defining the subspan to retriev. |
required |
Returns:
| Type | Description |
|---|---|
Union[int, slice]
|
either the Token at position i in the span or the subspan defined by the slice i. |
Token ¶
A token
TODO Some parts of the API are still not implemented
sent The sentence span that this token is a part of.
Span
lang Language of the parent document’s vocabulary.
str
Token's constructor
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
token
|
Token
|
the C++ binding Token class |
required |
text
class-attribute
instance-attribute
¶
i
class-attribute
instance-attribute
¶
i = property(fget=lambda self: self.token.i + 1, doc='The index of this token in its parent document.')
lemma
class-attribute
instance-attribute
¶
pos
class-attribute
instance-attribute
¶
pos = property(fget=lambda self: self.token.tag, doc='Coarse-grained part-of-speech from the Universal POS tag set.')
head
class-attribute
instance-attribute
¶
head = property(fget=lambda self: self.token.head, doc='The syntactic parent, or “governor”, of this token.')
dep
class-attribute
instance-attribute
¶
idx
class-attribute
instance-attribute
¶
idx = property(fget=lambda self: self.token.pos - 1, doc='Position of this token in its document text.')
features
class-attribute
instance-attribute
¶
features = property(fget=lambda self: {} if self.token.features == '_' else dict(x.split('=') for x in self.token.features.split('|')), doc='Morphlogical features of this token .')
ent_type
class-attribute
instance-attribute
¶
ent_iob
class-attribute
instance-attribute
¶
ent_iob = property(fget=lambda self: self.token.neIOB, doc='IOB code of named entity tag. “B” means the token begins an entity, “I” means it is inside an entity, “O” means it is outside an entity, and "" means no entity tag is set.')
t_status
class-attribute
instance-attribute
¶
t_status = property(fget=lambda self: self.token.tStatus, doc='The tokenization status of this token. Can also be explored with the is_* properties. The possible values are::\n\n t_alphanumeric\n t_abbrev\n t_acronym\n t_capital\n t_capital_1st\n t_capital_small\n t_cardinal_roman\n t_comma_number\n t_dot_number\n t_fraction\n t_integer\n t_ordinal_integer\n t_ordinal_roman\n t_sentence_brk\n t_small\n t_word_brk\n\n')
is_alpha
class-attribute
instance-attribute
¶
is_alpha = property(fget=lambda self: self.token.tStatus in ['t_alphanumeric', 't_capital', 't_capital_1st', 't_capital_small', 't_small'], doc='Does the token consist of alphabetic characters? Equivalent to token.text.isalpha().')
is_digit
class-attribute
instance-attribute
¶
is_digit = property(fget=lambda self: self.token.tStatus == 't_integer', doc='Does the token consist of digits? Equivalent to token.text.isdigit().')
is_lower
class-attribute
instance-attribute
¶
is_lower = property(fget=lambda self: self.token.text.islower(), doc='Is the token in lowercase? Equivalent to token.text.islower().')
is_upper
class-attribute
instance-attribute
¶
is_upper = property(fget=lambda self: self.token.text.isupper(), doc='Is the token in lowercase? Equivalent to token.text.isupper().')
is_punct
class-attribute
instance-attribute
¶
is_punct = property(fget=lambda self: self.token.tStatus in ['t_sentence_brk', 't_word_brk'], doc='Is the token punctuation?')
is_sent_start
class-attribute
instance-attribute
¶
is_sent_start = property(fget=lambda self: self.token.i == 0, doc='Does the token start a sentence? bool or None if unknown. Default value = True for the first token in the Doc.\nTODO: implement for sentences other than the first one.')
is_sent_end
class-attribute
instance-attribute
¶
is_sent_end = property(fget=lambda self: self.token.tStatus == 't_sentence_brk', doc='Does the token end a sentence? bool or None if unknown.')
is_space
class-attribute
instance-attribute
¶
is_space = property(fget=lambda self: self.token.text.isspace(), doc='Does the token consist of whitespace characters? Equivalent to token.text.isspace(). Should always be False in LIMA as there is no space tokens')
is_bracket
class-attribute
instance-attribute
¶
is_quote
class-attribute
instance-attribute
¶
is_quote = property(fget=lambda self: self.token.text in '"\'«»`', doc='Is the token a quotation mark?')
__len__ ¶
Return the length of the token in UTF-8 code points
Returns:
| Type | Description |
|---|---|
int
|
the length of the token |
LimaInternalError ¶
Bases: Exception