Discussions
Categories
Groups
Community Home
Categories
INTERNAL ENABLEMENT
POPULAR
PUBLIC CLOUD
PRIVATE CLOUD
Quick Links
MY LINKS
HELPFUL TIPS
Back to website
Home
Web CMS (TeamSite)
MT vocab to evaluate certain fields
motlnt
I have a particular requirement to create a DCR and then set metadata for that DCR.
I have two items in DCR called description and features.
I have a keyword vocabulary and i want to have a vocab keyword box in MT UI that evaluates only description and features fields from DCR and automatically identify keywords???
Teamsite version: 5.5
MT version: 3.1
Thanx in advance!!
Find more posts tagged with
Comments
nipper
Hi Manju
Think you would be best writing your own preprocessor to filter out all of the other fields. (I
assume there are more than the 2 fields in the DCR). This preprocessor should just filter out
the 2 fields and the other XML & pass it to MT.
Now if one of the fields is a Text area populated by visual formatter, then the problem is more difficult
since your preprocessor would need to filter the XMl then the HTML.
HTH
Andy
motlnt
Thanx Andy..
Do you have any sample preprocessor written for this type of functionality.
And yes, the textareas are populated by visual formater so i need to filter out XMl tags as well as HTML tags.
Thanx in advance..
Regards,
nipper
Nope, but renew my contract & I will see.
Actually is is not to hard,
For XML alone, you have a perl script open the DCR, use DCRparser to
extract the appropriate fields, Dump that information into a file. You will need
to take another step by filtering out all the HTML, but you *should* be able to
invoke the MT HTML parser for that. Worth a shot.
There is a section in the MT Advanced Users Guide on creating your own Preprocessor.
HTH
Andy
motlnt
I wish i had those rights
Anyways, I want to fetch the description and features from DCR XML file and then pass it to recogniser to evaluate the keywords from description and features data.
In Avanced guide, it mentions that
Whether you want to isolate structured text or tag files directly, you must add the
state-machine preprocessor to a preprocessor group (for example, a group called
“htmlTitle”) and then add the group to a content processor. If you are isolating text, include
the preprocessing group in the same content processor that references the recognition or
classification indexes you intend to use.
I didn't understand, how to put this SMIL in processor group or the group in content processor.
It is not mentioned in detail in this document also.
I found that this advaced guide is available for MT versions 3.6 and 4.0
I wonder that this feature is available with MT 3.1 as it is no where mentioned in MT3.1 admin guide.
I have written this preprocessor to get the features and description and i am just setting it in buffer A and B, i dont know where to set it to set it as an input to recognizer....
Even the place where i am trying to assign token DESCRIPTION and FEATURES is giving me a compilation error. I dont know the escape character to be used for "description"
opcodes ce_opcodes.opc
tokens
{
START_XML "<record";
GT ">";
END_XML "</record>";
DESCRIPTION "<item name="description"><value";
ENDITEM "</value></item>";
FEATURES "<item name="features"><value";
}
action NOTHING
{
noop();
}
action COLLECT
{
collect_view();
}
action DONE
{
settaga("description");
settagb("features");
exit();
}
action DESCRIPTION
{
collect_view();
appenda();
}
action FEATURES
{
collect_view();
appendb();
}
states
{
state INIT
START_XML COLLECT IN_START_XML
END_XML NOTHING INIT
DESCRIPTION NOTHING INIT
FEATURES NOTHING INIT
ENDITEM NOTHING INIT
GT NOTHING INIT
OTHER NOTHING INIT
EOS DONE INIT
;
state COLLECT_STORY
START_XML COLLECT COLLECT_STORY
END_XML COLLECT INIT
DESCRIPTION COLLECT IN_START_DESCRIPTION
FEATURES COLLECT IN_START_FEATURES
ENDITEM COLLECT COLLECT_STORY
GT COLLECT COLLECT_STORY
OTHER COLLECT COLLECT_STORY
EOS DONE INIT
;
state COLLECT_DESCRIPTION
START_XML DESCRIPTION COLLECT_DESCRIPTION
END_XML DESCRIPTION COLLECT_DESCRIPTION
DESCRIPTION DESCRIPTION COLLECT_DESCRIPTION
FEATURES COLLECT COLLECT_DESCRIPTION
ENDITEM COLLECT COLLECT_STORY
GT DESCRIPTION COLLECT_DESCRIPTION
OTHER DESCRIPTION COLLECT_DESCRIPTION
EOS DONE INIT
;
state COLLECT_FEATURES
START_XML FEATURES COLLECT_FEATURES
END_XML FEATURES COLLECT_FEATURES
DESCRIPTION DESCRIPTION COLLECT_FEATURES
FEATURES COLLECT COLLECT_FEATURES
ENDITEM COLLECT COLLECT_STORY
GT DESCRIPTION COLLECT_FEATURES
OTHER DESCRIPTION COLLECT_FEATURES
EOS DONE INIT
;
state IN_START_XML
START_XML COLLECT IN_START_XML
END_XML COLLECT IN_START_XML
DESCRIPTION COLLECT IN_START_XML
FEATURES COLLECT IN_START_XML
ENDITEM COLLECT IN_START_XML
GT COLLECT COLLECT_STORY
OTHER COLLECT IN_START_XML
EOS DONE INIT
;
state IN_START_DESCRIPTION
START_XML COLLECT IN_START_DESCRIPTION
ENDITEM COLLECT IN_START_DESCRIPTION
DESCRIPTION COLLECT IN_START_DESCRIPTION
FEATURES COLLECT IN_START_DESCRIPTION
ENDITEM COLLECT IN_START_DESCRIPTION
GT COLLECT COLLECT_DESCRIPTION
OTHER COLLECT IN_START_DESCRIPTION
EOS DONE INIT
;
state IN_START_FEATURES
START_XML COLLECT IN_START_FEATURES
ENDITEM COLLECT IN_START_FEATURES
DESCRIPTION COLLECT IN_START_FEATURES
FEATURES COLLECT IN_START_FEATURES
ENDITEM COLLECT IN_START_FEATURES
GT COLLECT COLLECT_FEATURES
OTHER COLLECT IN_START_FEATURES
EOS DONE INIT
;
}
Thanx & Regards,
Manju.
nipper
>I found that this advaced guide is available for MT versions 3.6 and 4.0
>I wonder that this feature is available with MT 3.1 as it is no where mentioned in MT3.1 admin guide.
Aren't you on 3.5 not 3.1 ?
In 3.5 the relavent section is chapter 5 of the admin guide, building a preprocessor. That has the steps to walk you
though adding a new preprocessor
motlnt
I have gone through this document.
It gives the steps to get some field details from a document and show it in MT UI like the the title example thing is given.
But my requirement is to get two field values from DCR(features and description) and use the keyword vocabulary to evaluate these fields and find out the keywords and put it in a field in MT UI for that DCR and this process is not mentioned in this document.
A bit of it was mentioned in advanced guide of 3.6 and 4.0 but not in 3.5 admin guide.
This is the paragragh mentioned in MT 4.0 advanced user guide(page no-36)....i want to do first possibility
How you configure MetaTagger depends on how you want to use the preprocessor. There
are two possibilities: 1) using the preprocessor to isolate structured text and send it on for
further processing (such as recognition or classification), or 2) using it to tag input files
directly.
So i was just wondering whether above mentioned type of functionality is provided with 3.5 or not.
I am totally confused how to go about it
nipper
> 1) using the preprocessor to isolate structured text and send it on for
>further processing (such as recognition or classification)
Yes, this is what you want to do. The MT 3.5 Admin manual also has a
chater on building a preprocessor (with the same example). The preprocessor
can so things like pull out specific tags (like in the example) or it can
just filter out non-text tags (HTML tag, or PDF, or whatever).
The output in the second example is just pure text, everything to be run
through the recognizer.
In older releases of MT (maybe as late as 3.5) IW used a 3rd party cracking utility that
would take HTML tags out, PDF tags, Word, etc and just leave text.
See how those are invoked in the MT.cfg and see of you can replicate that,
I assume the utility just dumps to STDIO.
HTH
Andy
motlnt
I tried parsing the DCR and getting the description and features attributes and dumped these two values in a test.txt file which gives me values like
<p>manju testing o/p <strong><em><u>cmos clock</u></em></strong></p>
<p>output</p>
I tried running iwgenmetadata command on test.txt as specified below :
iwgenmetadata -suggest_by_tag Recognizer_Keywords_1 /u03/iw-home/httpd/iw-bin/SPS/cms/test.txt -save
where Recognizer_Keywords_1 is an index for keyword vocabulary and an item in MT UI.
The o/p for this command is as below:
$ more test.imd
<?xml version="1.0" encoding="ISO-8859-1" ?>
<metadata>
<facet>
<facetName>Recognizer_Keywords_1</facetName>
<descriptor>
<vocab>RKW1</vocab>
<code>9340458192</code>
<label>clock</label>
</descriptor>
<descriptor>
<vocab>RKW1</vocab>
<code>9340458343</code>
<label>cmos</label>
</descriptor>
</facet>
</metadata>
The o/p is as expected but this is stored in a file test.imd.
This shows that even if the html tags are there in a txt file, it doesn't matter and it parses it properly and takes out the keywords so no need of HTML Processor.
I read in admin guide that if i want to store these metadata values as extended attributes, we should use iwmtbatch instead of iwgenmetadata
I tried using iwmtbatch command but it gives me error as below:
$ iwmtbatch
ld.so.1: iwmtbatch: fatal: libteamsite_api.so: open failed: No such file or directory
Killed
I dont know why this error is occuring, is there any configuration which i am missing.
If i try using iwmtbatch command as below:
iwmtbatch -suggest_by_tag Recognizer_Keywords_1 /u03/iw-home/httpd/iw-bin/SPS/cms/test.txt -save <<dcrname>>
Will this command save the keywords as the Recognizer_Keywords_1extended attribute of the dcrname specified.
Basically i want the keywords to be evaluated from test.txt and get saved as extended attributes for the DCRname specfied.
Thanx in advance....
nipper
The HTML tags may not hurt anything. But I have seen issues when they are not filtered. I do
remember things like the industry type would list Software Development due to those tags. May not be an issue,
but I would test more
$ iwmtbatch
ld.so.1: iwmtbatch: fatal: libteamsite_api.so: open failed: No such file or directory
Killed
Your shell is not properly setup. There is a file that will do that for you: $MT-HOME/conf/env-metatagger.sh,
source that file:
at a KSH prompt:
. ./.env-metatagger.sh
Then the command lines should work. You can also put that into your profile.
motlnt
Thanx a lot Andy...
i was trying to remember that we used to do something to run this command during Metatagger project...I'll remmber this time
Just to confirm,
If i try using iwmtbatch command as below:
iwmtbatch -suggest_by_tag Recognizer_Keywords_1 /u03/iw-home/httpd/iw-bin/SPS/cms/test.txt -save <<dcrname>>
Will this command save the keywords as the Recognizer_Keywords_1extended attribute of the dcrname specified.
Basically i want the keywords to be evaluated from test.txt and get saved as extended attributes for the DCRname specfied.
nipper
>Will this command save the keywords as the Recognizer_Keywords_1extended attribute of the dcrname specified.
>Basically i want the keywords to be evaluated from test.txt and get saved as extended attributes for the DCRname specfied.
That part I do not know, but I do not think so. I think it will attach the EAs to the txt file. The preprocessor would usually
get invoked by iwmtbatch
I do not have access to MT so I am working from memory.
Andy
motlnt
It still gives me this error
I tried setting env-metatager.sh
$ ./env-metatagger.sh
DB_HOME=/u03/metatagger/conf
MC_LOG_DIR=/u03/metatagger/logs
PATH=/u03/metatagger/bin:/u03/metatagger/thirdparty:/u03/metatagger/thirdparty/c
harsets:/sbin:/usr/bin:/usr/sbin:/usr/ucb:/etc:/usr/openwin/bin:/usr/bin/X11:/op
t/VRTSvxva/bin:/opt/VRTSvmsa/bin:/opt/hpnpl/bin:/usr/local/proctool/bin:/usr/loc
al/bin:/usr/local/sbin:/opt/NSCPcom:/opt/SUNWxntp/bin:/usr/platform/sun4u/sbin:/
usr/pb/bin:/usr/pb/sbin:/opt/EMCpower/bin/sparcv9:/opt/VRTSvcs/EMC/bin:/etc/emc/
bin:.
LD_LIBRARY_PATH=/u03/iw-home/lib:/u03/iw-home/iwopenapi
$ /u03/metatagger/bin/iwmtbatch
ld.so.1: /u03/metatagger/bin/iwmtbatch: fatal: libteamsite_api.so: open failed:
No such file or directory
Killed
nipper
Run this command:
find /u03/ -name libteamsite_api.so -print
It appears that the MT directory /u03/metatagger/lib is not getting populated by
env-metatagger.sh, that should be in LD_LIBRARY_PATH
Andy