LIMA
Libre Multilingual Analyzer — C++ API
Loading...
Searching...
No Matches
LIMA

Principles

The LIMA linguistic processing module is designed to work either with a direct load of the dynamic libraries in the program or with a client-server architecture. The API has then been designed as sufficiently generic for this purpose. In particular, it uses a callback strategy to manage in a single function several output formats.

The same API is also used either for analysis of simple text or for analysis of structured XML documents.

Main classes

AbstractLinguisticProcessingClient is the main class that implements the generic LIMA API. Its main function to launch the analysis on a text is AbstractLinguisticProcessingClient::analyze().

The callback is implemented through a handler, AbstractAnalysisHandler, that specifies the interface to handle the different events generated by the LIMA analyzer. This handler is closely related to the configuration of the analyzer concerning its output format. The choice of LIMA output is specified by setting a dumper as the last pipeline unit of the pipeline. A dumper is a class that takes the results of the linguistic analysis produced by LIMA and outputs a given subset of the information in a given format. The list of possible dumpers is given in LIMA configuration files.

Implementations of the AbstractLinguisticProcessingClient included by default in the LIMA analyzer are :

  • mm-core-client : the main client that performs the analysis process on a simple text
  • lp-xmlreader-client : a client that performs the analysis of a structured XML document (XML parsing and spotting of parts of text to analyze according to an explicit configuration). This client uses internally a core client to process the parts of texts identified

to call the LIMA analyzer

To call the LIMA analyzer, one must indicate

  • the name of a pipeline to be used, specifying the set of processing units that are to be activated for the text analysis: such pipelines and processing units are configured in the LIMA configuration files (lima-analysis.xml and lima-lp-xxx.ml, where xxx is the 3-letter code of the language considered for analysis)
  • the handler, adapted to the chosen dumper, and specifying what to do with this output (may be written to a file or stored in memory for further processing...

The program analyzeText.cpp contains several examples of calls of the LIMA analyze function. Here are some commented examples:

Initialization of the LinguisticProcessing client

// several languages or pipelines may be configured in the same
// analyzer client, the ones to actually use are specified on the call of the
// analyze() function
// here, we only configure the only elements that we want to use
std::deque<std::string> langs(1,"fre"); // the configured languages
std::deque<std::string> pipelines(1,"main"); // the configured pipelines
// the configuration file lima-analysis.xml containing all the configuration elements
// for the different languages and pipelines is given as argument to the
// client factory
// Then, an analyzer client is created by the factory: here, we choose a
// local client that will load the implementation of the analyzer in
// dynamic libraries
AbstractLinguisticProcessingClient* client(0);
XMLConfigurationFiles::XMLConfigurationFileParser lpconfig("conf/lima-analysis.xml");
"lima-coreclient",
lpconfig,
langs,
pipelines);
// some metadata are used by the analysis: the minimal metadata required
// are a language and a filename
// (some other data can be given, such as date, that allows to
// normalize relative dates)
map<string,string> metaData;
metaData["Lang"]="fre"
metaData["FileName"]="test.txt"
std::shared_ptr< AbstractProcessingClient > createClient(const std::string &id) const override
create an Client using the appropriate registered factory.
void configureClientFactory(const std::string &id, Common::XMLConfigurationFiles::XMLConfigurationFileParser &configuration, std::deque< std::string > langs=std::deque< std::string >(), std::deque< std::string > pipelines=std::deque< std::string >()) override
configure the corresponding clientFactory so that it can create clients.
static const LinguisticProcessingClientFactory & single()
const singleton accessor
Definition Singleton.h:51
static LinguisticProcessingClientFactory & changeable()
singleton accessor
Definition Singleton.h:71

An example of call to analyze() with simple text output

// for simple text output, the handler only redirects the stream of
// data produced by the analyzer on a ostream (here, a file)
//ofstream fout("output.txt");
//SimpleStreamHandler handler(&fout);
//client->setAnalysisHandler(&handler);
client->analyze("ceci est un texte UTF-8",
metaData,
"main");

An example of call to analyze() with BoW output

BoW is the LIMA binary format for an extended bag-of-words representation, including compounds. This binary format is readable with the readBoWFile utility program (–xml outputs an XML representation of the format).

// for BoW output, we use a specific handler that handles the binary
// format and write it to a file
// ofstream fout("output.bin");
// BowTextWriter bowWriterHandler(&fout);
// client->setAnalysisHandler(&bowWriter);
client->analyze(text,metaData,"main");