Skip to content

Output formats

CoNLL-U

The default output of the pipelines is the CoNLL-U Plus format: one token per line, sentences separated by blank lines, ten tab-separated columns.

# Field Content in LIMA
1 ID Word index, starting at 1 in each sentence; a range for multiword tokens
2 FORM Word form or punctuation symbol
3 LEMMA Lemma, or _ if not available
4 UPOS Universal part-of-speech tag (legacy pipelines use LIMA's own tagset)
5 XPOS Language-specific part-of-speech tag; _ with the neural pipelines
6 FEATS Morphological features from the universal feature inventory
7 HEAD Head of the word (an ID, or 0 for the root)
8 DEPREL Dependency relation to the head
9 DEPS Enhanced dependencies; always _ (not produced by LIMA)
10 MISC Pipe-separated key=value annotations

LIMA always writes Pos (character offset in the text) and Len (length) in the MISC column, plus SpaceAfter=No when the token is not followed by a space. Named entities are annotated with the NE key, whose value is the entity type.

# sent_id = 1
# text = The author wrote a novel.
1  The     the     DET    _  Definite=Def|PronType=Art         2  det    _  Len=3|Pos=1
2  author  author  NOUN   _  Number=Sing                       3  nsubj  _  Len=6|Pos=5
3  wrote   write   VERB   _  Mood=Ind|Tense=Past|VerbForm=Fin  0  root   _  Len=5|Pos=12
4  a       a       DET    _  Definite=Ind|PronType=Art         5  det    _  Len=1|Pos=18
5  novel   novel   NOUN   _  Number=Sing                       3  obj    _  Len=5|Pos=20|SpaceAfter=No
6  .       .       PUNCT  _  _                                 3  punct  _  Len=1|Pos=25

The NER pipelines can produce CoNLL-03 instead (see the conllDumperNer configuration and the ConllDumper reference).

Bag of words

The legacy pipelines can produce a binary "bag of words" (BoW) representation of the text, containing the lemmas of the content words (nouns, adjectives, verbs) and the named entities. Replace conllDumper by bowDumper in your copy of lima-lp-<lang>.xml and pass -d bow to analyzeText. Print the result with readBowFile:

readBowFile file.txt.bin

For the text "On 4th April 2011, Bill Williams looked at Paris from his home window.", the output is:

(4th_April_2011-8192-4)->[*(4th-8192-4)(April-16384-8)(2011-8192-14)]:DateTime.DATE:date=2011-04-04;value=4th April 2011
(Bill_Williams-16384-20)->[*(Bill-16384-20)(Williams-16384-25)]:Person.PERSON:firstname=Bill;lastname=Williams;value=Bill Williams
(look-49152-34)
(Paris-16384-44)->[*(Paris-16384-44)]:Location.LOCATION:value=Paris
(home-8192-59)
(window-8192-64)

Each line is a term: simple (the verb "to look") or complex (the named entities "Paris", "Bill Williams" and "4th April 2011"). Each term is described by its normalized form, the numerical value of its category and its position in the text, followed by its structure. This text form is meant for inspection only; use the C++ BoW API to process bags of words in programs.

Other formats

LIMA has many other dumpers (XML, annotation graphs, Brat, dependency triples…) and loggers producing intermediate results for debugging.