Discussions
Categories
Groups
Community Home
Categories
INTERNAL ENABLEMENT
POPULAR
PUBLIC CLOUD
PRIVATE CLOUD
Quick Links
MY LINKS
HELPFUL TIPS
Back to website
Home
Web CMS (TeamSite)
Migration
abhishek_gupta
Hi,
There are a few hundred HTML pages of which we have created a TeamSite template. Now we need to convert all these existing HTML pages into the DCR format. Is there any tool or any other way via which we can convert all the existing HTML pages into the DCR format.
I have gone thru the Content_migration_best_practices.pdf file provided by Interwoven, but couldnt find any solution to my problem.
Thanks,
Abhishek
Find more posts tagged with
Comments
jbonifaci
Are all of the html files you are converting created from a template?
If so, you should be able to write a perl script to read in every html file, strip out the info that is custom to each html page and then output the xml for the dcr with the values inserted.
What you basically need to do is look at all of the pages and see if there is a way to identify the location within the html of all the information you want to strip out that is common to all of the files. Then write a perl script to look for these identifiers and strip out the values between them. You then insert the stripped out values into your xml when you output it. I've done this several times at different clients and my conversion percentages have been around 96-99%, the remaining files needing to be converted by hand.
If all of the html pages were created following no template or guidelines, the pages probably aren't good candidates for automated conversion. This really depends on the complexity of the template and html though. If your template is as simple as having only title and body elements, you should be able to strip out the title and body elements from any html page easily. If your template has 50 different fields and replicants, you probably won't be able to strip out all of the information from an html page, unless it followed a strict template.
Jeff Bonifaci
laj1
I agree with the previous poster. I've implemented this exact process before: parse, write DCR, set EAs, generate.
You read each HTML file, and grab the parts that go into the DCR. You write the DCR, and then don't forget to
set your extended attributes for datatype, etc.
Len.
Len Jaffe
My Heart Is A Flower
Update your DevNet profile - let us know who you are!
ela
Hmm, if you can't split the content into more parts, a quick and dirty solution could be to simply import (with script) the whole html content into one textarea (which allows html) in a dcr. Try to exclude unnecessary tables etc.
Cheers,
ela
jorge1
Hi,
Well, this is true if your content has some sort of pattern!!! Else I wish you luck with copy and paste!!!
Jorge
Jorge Rodrigues
Cambridge Technology Partners Switzerland S.A.
A Division of Novell, Inc.
skip11
Dude, you so need a developer.
Skip
webby
Hey, I'm a developer.
I bounce servers all the time - but bits keeps breaking off....