Discussions
Categories
Groups
Community Home
Categories
INTERNAL ENABLEMENT
POPULAR
PUBLIC CLOUD
PRIVATE CLOUD
Quick Links
MY LINKS
HELPFUL TIPS
Back to website
Home
Web CMS (TeamSite)
The Minimum Threshold
j0ej0e
Hi,
We're experimenting with different minimum thresholds to tag our documents more effectively. We are talking about around 500,000 documents.
Just one question, what is the difference between minimum threshold and the so called "confidence level", are both the same?
Does it mean if i set the minimum threshold higher it is telling Metatagger "unless you're <n>% confident that the document matches then don't tag it" ?
Side question : How is the minimum threshold derived?
rgds,
joe
Find more posts tagged with
Comments
j0ej0e
Hi,
The side question i meant to ask was :
How are the "scores" we see in Metatagger Studio derived? Briefly?
rgds,
joe
mikelauruhn
I don't know that I would call those "confidence" numbers per se, but yes, you would use the minthreshold for the scenario that you are describing: Tag it with any concept that scores above a certain value.
We are actually working on the same thing right now and trying to find the right threshold for tagging based on our business objectives. What we have been trying is to gather together a small set of documents that I want to test on a particular topic. These documents will vary in relevance. Then, I will run those documents against the classifier (using the Studio Test Index as opposed to Evaluate Index), collect those scores and show them to our business or editorial team.
completely_relevant_doc_1 0.680
completely_relevant_doc_2 0.634
somewhat_relevant_doc_1 0.512
somewhat_relevant_doc_2 0.491
not_at_all_relevant_doc_1 0.211
not_at_all_relevant_doc_2 0.124
If they think that the top two are the only acceptable ones, then we'll set the minthreshold at .600. If they think documents like the third one should be included, then it is somewhere around .500.
As far as how the scores are derived:
Recognition is pretty straightforward. It is a whole number representing the total number of time the label or alternate terms appear in the document.
For a classifier, it represents the similarity of the document you are testing to the documents it matched from the training set. I think the simplest way to think about it is that the test document contains x% of the terms and phrases that are prevalent in the training documents that are most similar. Something like, my document called "Heart Attacks" contains 75% of the terms and phrases that are prevalent in the most similar training set documents about "Cardiovascular Health". To that end, I would not think of something that scores .300 as being 30% confident. It just means it contains 30% of the prevalent terms from a training set doc -- and that could mean that the document is relevant. Your scores and relevance will vary based on the size of your training set and types of content in there. (this is my best hunch. if anyone from IWOV wants to correct me, please do!) I think I remember setting K to 1 and testing a document that was already in the training set and it scored 1.0 (i.e. containing 100% -- no more, no less -- of the terms and phrases in the training set document).
Good luck.
MikeL
j0ej0e
Thanks Mike for your reply, btw are your scores derived with the "- normalize" flag?
In my case its really rare to see content (not training documents) getting 0.1+ scores, with normalize flag on, we are getting scores like 0.033.
And our training documents are having scores similar to yours, from 0.2 to highest 0.5+
But we can't set minimum threshold to 0.2, because majority of our results are in the 0.01 - 0.09 range.
mikelauruhn
No, we are not using the -normalize flag in our script.
So, I have to ask... how large is your vocabulary and training set of content?
In our case, we are using pretty large training sets: between 15 and 18 documents per concept. But they are documents that we evaluated and deliberately selected. This let's us set a K number around 7 for our testing. If I test a document at K 1, its score is lower than at the K 7 level. With your scores, is there anyway to identify a distinct difference in the scores of the documents you expect to be tagged with a concept and those that don't? Even if it means relevant documents score around .07-.09 and the non-relevant documents are scoring below .05? Make sense? If you cannot make a distinction between what your relevant and non-relevant test documents are scoring, then you may want to re-evaluate the training set.
j0ej0e
Hi,
Yes you're correct, we can say those content which does not match can fall below , say the 0.05 mark.
But the major problem we are facing is the sheer amount of nodes in our controlled vocabulary.
We have about 427 nodes, and we are facing right now the problem that content between them may overlap.
For each node we have about 10 training documents, and the reason being we have a Taxonomy and the scope for each node is defined clearly.
Based on that we have to find content that is representative of that node and within the boundaries set by the Scope.
And another problem is, there is clearly not many content for some of the nodes, and right now we are considering merging or dropping some of the nodes.
rgds,
joe
Migrateduser
Joe,
Definitely a case for using more than 1 model, probably a combination of classification and recognition, though a multi-classifier voting system might work. You might consider running a small test with a classifier that only uses the scope notes as a training document.
Rgs
Clark
j0ej0e
Hi Cbreyman,
My apologies for replying so late as i have been busy working on the client's project.
Out of curiosity, what is or how do you implement the classifier + recognizer, and how do you implement a multi-classifier voting system?
Could you kindly explain briefly?
We're using Metatagger 4.0.1 and the above 2 implementations are quite new to us.
rgds,
joej0e
Migrateduser
joej0e,
Step 1 - Implement each classifier and recognizer model as separate field.
Step 2 - Implement a plug-in final processor that does the set logic on the different fields and populates the result in new field.
Example - Category Filtering
Step 1 - Create Recognizer R1, Recognizer R2
Recognizer R1 adds categories that should apply to field R1
Recognizer R2 adds categories that MUST not apply to field R2
Step 2 - Implement a plug-in processor
The plug in processor:
a) parses the input XML metadata record
b) iterates over categories in field R2 and removes them from field R1
c) writes out new XML metadata record
Example - Category Voting
Step 1 - Create Recognizer R1, Recognizer R2
Recognizer R1 adds categories that should apply to field R1
Classifier C1 adds categories that should apply to field C1
Step 2 - Implement a plug-in processor
The plug in processor:
a) parses the input XML metadata record
b) iterates over categories in field R1 and removes them if they are not in C1
c) writes out new XML metadata record
j0ej0e
Thanks for your kind assistance