Named entity recognition¶
LIMA includes predefined pipelines for named entity recognition (NER) in
English and French, available out of the box in the legacy eng and fre
configurations.
| Pipeline | Input | Output |
|---|---|---|
ner-rules |
Plain text | CoNLL-03 |
ner-rules-pretok |
CoNLL-U | CoNLL-03 |
The output can be switched to CoNLL-U in the configuration of the
conllDumperNer process unit (in
lima-lp-eng.xml and lima-lp-fre.xml).
Rule-based NER is implemented with ModEx rules,
whose sources are in
lima_linguisticdata/SpecificEntities/<lang>.
You can add your own entity types by writing a ModEx.
Neural NER
LIMA previously also provided neural (ner-deep) and combined
(ner-fusion) NER pipelines, based on TensorFlow. TensorFlow support has
been removed; these pipeline names are kept for compatibility but
currently run the rule-based recognizer only. A libtorch-based NER unit
(RnnNER) exists but no models are published for it yet.
Usage¶
analyzeText -l eng -p ner-rules input_file.txt
analyzeText -l fre -p ner-rules-pretok input_file.conllu
Processing units¶
ner-rules |
ner-rules-pretok |
|
|---|---|---|
| Input | ||
flattokenizer |
✓ | |
conllureader |
✓ | |
| Pre-processing | ||
simpleWord |
✓ | ✓ |
defaultProperties |
✓ | ✓ |
| Rule-based NER | ||
SpecificEntitiesModex |
✓ | ✓ |
sentenceBoundariesUpdater |
✓ | ✓ |
| Output | ||
conllDumperNer |
✓ | ✓ |
Evaluation¶
The pipelines are evaluated on pre-tokenized input (ner-rules-pretok).
processed 46435 tokens with 5616 phrases; found: 4440 phrases; correct: 2984.
accuracy: 92.45%; precision: 67.21%; recall: 53.13%; FB1: 59.35
LOC: precision: 64.31%; recall: 84.99%; FB1: 73.22 2202
MISC: precision: 90.32%; recall: 3.99%; FB1: 7.65 31
ORG: precision: 57.07%; recall: 20.10%; FB1: 29.73 580
PER: precision: 74.31%; recall: 75.47%; FB1: 74.88 1627
processed 3499655 tokens with 296413 phrases; found: 211853 phrases; correct: 128203.
accuracy: 92.12%; precision: 60.52%; recall: 43.25%; FB1: 50.45
LOC: precision: 64.89%; recall: 55.78%; FB1: 59.99 72577
MISC: precision: 57.84%; recall: 1.84%; FB1: 3.56 2144
ORG: precision: 41.70%; recall: 36.84%; FB1: 39.12 41009
PER: precision: 65.30%; recall: 64.00%; FB1: 64.64 96123
processed 3499679 tokens with 251726 phrases; found: 141976 phrases; correct: 102255.
accuracy: 92.67%; precision: 72.02%; recall: 40.62%; FB1: 51.95
LOC: precision: 68.93%; recall: 45.60%; FB1: 54.89 74057
MISC: precision: 45.09%; recall: 4.74%; FB1: 8.57 4096
ORG: precision: 53.15%; recall: 33.15%; FB1: 40.84 15262
PER: precision: 84.94%; recall: 54.04%; FB1: 66.06 48561
The rule-based recognizer processes about 6,400 tokens per second on a single thread.