Principles
The LIMA linguistic processing module is designed to work either with a direct load of the dynamic libraries in the program or with a client-server architecture. The API has then been designed as sufficiently generic for this purpose. In particular, it uses a callback strategy to manage in a single function several output formats.
The same API is also used either for analysis of simple text or for analysis of structured XML documents.
Main classes
AbstractLinguisticProcessingClient is the main class that implements the generic LIMA API. Its main function to launch the analysis on a text is AbstractLinguisticProcessingClient::analyze().
The callback is implemented through a handler, AbstractAnalysisHandler, that specifies the interface to handle the different events generated by the LIMA analyzer. This handler is closely related to the configuration of the analyzer concerning its output format. The choice of LIMA output is specified by setting a dumper as the last pipeline unit of the pipeline. A dumper is a class that takes the results of the linguistic analysis produced by LIMA and outputs a given subset of the information in a given format. The list of possible dumpers is given in LIMA configuration files.
Implementations of the AbstractLinguisticProcessingClient included by default in the LIMA analyzer are :
- mm-core-client : the main client that performs the analysis process on a simple text
- lp-xmlreader-client : a client that performs the analysis of a structured XML document (XML parsing and spotting of parts of text to analyze according to an explicit configuration). This client uses internally a core client to process the parts of texts identified
to call the LIMA analyzer
To call the LIMA analyzer, one must indicate
- the name of a pipeline to be used, specifying the set of processing units that are to be activated for the text analysis: such pipelines and processing units are configured in the LIMA configuration files (lima-analysis.xml and lima-lp-xxx.ml, where xxx is the 3-letter code of the language considered for analysis)
- the handler, adapted to the chosen dumper, and specifying what to do with this output (may be written to a file or stored in memory for further processing...
The program analyzeText.cpp contains several examples of calls of the LIMA analyze function. Here are some commented examples:
Initialization of the LinguisticProcessing client
std::deque<std::string> langs(1,"fre");
std::deque<std::string> pipelines(1,"main");
AbstractLinguisticProcessingClient* client(0);
XMLConfigurationFiles::XMLConfigurationFileParser lpconfig("conf/lima-analysis.xml");
"lima-coreclient",
lpconfig,
langs,
pipelines);
map<string,string> metaData;
metaData["Lang"]="fre"
metaData["FileName"]="test.txt"
std::shared_ptr< AbstractProcessingClient > createClient(const std::string &id) const override
create an Client using the appropriate registered factory.
void configureClientFactory(const std::string &id, Common::XMLConfigurationFiles::XMLConfigurationFileParser &configuration, std::deque< std::string > langs=std::deque< std::string >(), std::deque< std::string > pipelines=std::deque< std::string >()) override
configure the corresponding clientFactory so that it can create clients.
static const LinguisticProcessingClientFactory & single()
const singleton accessor
static LinguisticProcessingClientFactory & changeable()
singleton accessor
An example of call to analyze() with simple text output
client->analyze("ceci est un texte UTF-8",
metaData,
"main");
An example of call to analyze() with BoW output
BoW is the LIMA binary format for an extended bag-of-words representation, including compounds. This binary format is readable with the readBoWFile utility program (–xml outputs an XML representation of the format).
client->analyze(text,metaData,"main");