Discussions
Categories
Groups
Community Home
Categories
INTERNAL ENABLEMENT
POPULAR
PUBLIC CLOUD
PRIVATE CLOUD
Quick Links
MY LINKS
HELPFUL TIPS
Back to website
Home
Web CMS (TeamSite)
Extractor and PreProcessor
yuvpurush
Hi,
Can anybody suggest me where exactly the pre processor and extractor differ. Both are designated to extract data based on some predefined pattern. Where to use pre processor and where to use extractor?
Rgds
Yuv
Find more posts tagged with
Comments
Migrateduser
Why not use the powerful,documented and tested regular expression engine provided by MetaTagger .vs. write your own pattern matcher to report metadata, if that is all you want to do in a pre-processor.
But, if your intent is to prepare the document(crackedText) by modifying(adding/deleting) content (because you know the document structure, and your taxonomy can benefit from it), a pre-processor is the ideal way forward.
Please refer to the section "Problems that Processors Can Address" in the MetaTagger - User Guide.
Migrateduser
preprocessors are a generalized plug-in mechanism for modifying the cracked text and/or the metadata record prior to the field engines.
extractors models are a type of field engine that is good for semantic pattern extraction
If the question is really about what makes extractors better than Perl for some tasks, there are a number of things:
1) Simplicity: With Perl you need to read the file, parse the metadata, etc. In extractors you just define the patterns you need.
2) Compilable Lexicons: With Perl, building a pattern that recognizes a large number of known words can be very complex and slow. Extractors can utilize compiled recognizer models to support large word-lists for parts of patterns. Consider building in perl the simple extractor pattern "FIRSTNAME LASTNAME" using a list of 50,000 first names and 100,000 last names.
3) Attribute Matching: With Perl, patterns match only the symbols (characters) in the input text. With Extractors, you can match either the symbols or the attributes of each token in the input, such as membership in one or more classes. Examples:
"[ FIRSTNAME 'Sriram' /^.*ay$/]" - Match any first name in the lexicon, the token 'Sriram' or the regular expression /^.*ay$/
"[ FIRSTNAME !'Clark' /^S.*$/]" - Match a first name that is not Clark and begins with a capital S.
4) Attribute Assignment: With Extractors, you can add new attributes to tokens or token sequences and create matching patterns. Because Perl is limited to literal matching, it would require back-flips.
That said, Perl is a general-purpose language. You can write lots of things in general purpose languages. That does not make the langauge the same as the tool created with it. Writing a database in Java does not make Java a database ....
yuvpurush
Thanks a lot for the clarification