Discussions
Categories
Groups
Community Home
Categories
INTERNAL ENABLEMENT
POPULAR
PUBLIC CLOUD
PRIVATE CLOUD
Quick Links
MY LINKS
HELPFUL TIPS
Back to website
Home
Web CMS (TeamSite)
excluding fields in dcr for summary/keywords
kellyrmilligan
Hello, I am on solaris, teamsite 6.5, MT 4.0.1
I am using the out of the box summary index to get key phrases/sentences. I am also using this for keywords. In the keyphrases, values from some fields such as author, release date, etc, are showing up. These are not really part of content. I guess I could just put them in EA's, but is there a way to exclude fields in a dcr from being looked at for the summary?
Find more posts tagged with
Comments
mikelauruhn
Hi Kelly,
What type of content are you tagging? If they are well-structured, like HTML or DCRs, you can write a relatively straightforward preprocessor that will evaluate only the fields you want MetaTagger to and ignore the others.
In our case, we wrote a DCR preprocessor in the State Machine language. It basically tells MetaTagger that for any content with no extension (a DCR), it should only evaluate the content that appears in the <headline> and <body> fields. If you have any other subject metadata already in there -- like a <keyword> field -- you could send that on to MT as well.
Hope this helps.
MikeL
kellyrmilligan
It does help, I am having some issues at the moment just getting a pre-processor to execute at all. I am pretty new to writing pre-processors, so any code snippets you have that might help would be much appreciated. At the moment, I am working on one that based on the directory the file is in, it will grab some fields out of the dcr and apply them as metadata. You would think there would be a tighter integration on the preprocessor side, but apparently not.
Migrateduser
Here is a SMIL snippet Kelly. This script captures an ABSTRACT, if present, and writes it to crackedText. CAUTION: not totally robust because if it does not find an ABSTRACT is blanks out the crackedText - which means any downstream processors will have ONLY the ABSTACT (if found) or a blank file - so use this script in a ContentProcesor that is isolated - but it will give you an idea of SMIL.
(Of course the SMIL must be compiled and placed in the DB_HOME\processors\SMIL dir.)
#
# Copyright 2001 Interwoven, Inc. All rights reserved.
#
# This machine scans a generic HTML file and extracts the title
# if it can.
#
opcodes ce_opcodes.opc
tokens
{
ABSTRACT "Abstract";
RETURN "\n";
SPACE " ";
LF "
";
}
action NOTHING
{
noop();
}
action COLLECT_AB
{
appenda();
}
action COLLECT
{
collect_story();
}
action SET_AB
{
settaga("Abstract");
}
action DONE
{
exit();
}
states
{
state INIT
ABSTRACT NOTHING DELETE_LF
OTHER NOTHING INIT
EOS DONE INIT
RETURN NOTHING INIT
;
state DELETE_LF
OTHER NOTHING DELETE_LF
RETURN NOTHING COLLECT_ABSTRACT
EOS DONE INIT
;
state COLLECT_ABSTRACT
OTHER COLLECT_AB COLLECT_ABSTRACT
RETURN COLLECT_AB CHECK_DONE
SPACE COLLECT_AB COLLECT_ABSTRACT
EOS DONE INIT
;
state COLLECT_STORY
OTHER COLLECT COLLECT_STORY
SPACE COLLECT COLLECT_STORY
RETURN COLLECT COLLECT_STORY
EOS DONE INIT
;
state CHECK_DONE
RETURN SET_AB COLLECT_STORY
LF SET_AB COLLECT_STORY
OTHER COLLECT_AB COLLECT_ABSTRACT
EOS DONE INIT
;
}