Discussions
Categories
Groups
Community Home
Categories
INTERNAL ENABLEMENT
POPULAR
PUBLIC CLOUD
PRIVATE CLOUD
Quick Links
MY LINKS
HELPFUL TIPS
Back to website
Home
Web CMS (TeamSite)
extractor language help
yanSi
I've got an input text file with sections separated by language tags like this:
eng content
fre content
I'm using an extractor to only report the french section by matching on the french tags, but I can't get it to work using
iwmtdebug -script myScript.txt -dump_xml input_file.txt
My pattern file has
open_fre /.*/* close_fre
I can match the open_fre and close_fre tokens fine, but when I combine them like that, nothing matches.
Find more posts tagged with
Comments
Migrateduser
Seems like this problem would be easier solved by creating a plug-in transconverter (like the XSLT transconverter sample). One content processor would extract the english and push it though the english analysis modules, another would do the same for french. See the archived presentation on "Cooking with MetaTagger Plugins"
yanSi
Thanks. I think I'll go with state machines instead. From what I've read, the sections of text can be extracted into the story buffer where it is passed on to the recognizers/classifers.
Migrateduser
I'd strongly advise avoiding SMIL. They are there for compatibility purposes, but the vast majority of customers use the plug-in framework for preprocessing input text.