![]() |
LIMA
Libre Multilingual Analyzer — C++ API
|
Multiword-token (MWT) expansion stage. More...
#include <deeplima/mwt_expander.h>
Public Types | |
| typedef std::function< void(const std::vector< segmentation::token_pos > &, uint32_t)> | callback_t |
Public Member Functions | |
| MwtExpander (const std::string &dict_fn) | |
| void | register_handler (const callback_t &fn) |
| size_t | size () const |
| void | set_require_flag (bool require_flag) |
| void | operator() (const std::vector< segmentation::token_pos > &tokens, uint32_t len) |
Multiword-token (MWT) expansion stage.
Sits between the tokenizer and the TokenSequenceAnalyzer in the inference pipeline: each surface token that matches a dictionary entry (e.g. French "du") is replaced by its syntactic sub-words ("de", "le") before tagging/parsing. The tagger and parser are trained on words (de=ADP/case, le=DET/det), so an unsplit surface token is out-of-distribution; expanding here keeps the neural models in-distribution and lets the CoNLL-U output carry the UD "N-M surface" range line.
The dictionary is the one produced by deeplima-mwt-dict / extract_mwt_dict: surface <TAB> count <TAB> word1 <TAB> word2 [<TAB> ...]
This is the dictionary-only path (closed-class, concatenative MWTs: French, German/Italian/Iberian contractions). Productive or non-concatenative MWTs (Hebrew/Arabic) would need a seq2seq fallback, not implemented here.
Definition at line 40 of file mwt_expander.h.
| typedef std::function<void(const std::vector<segmentation::token_pos>&, uint32_t)> deeplima::MwtExpander::callback_t |
Definition at line 43 of file mwt_expander.h.
|
inlineexplicit |
Definition at line 45 of file mwt_expander.h.
|
inline |
Definition at line 70 of file mwt_expander.h.
|
inline |
Definition at line 50 of file mwt_expander.h.
|
inline |
Definition at line 65 of file mwt_expander.h.
|
inline |
Definition at line 55 of file mwt_expander.h.