Discussions
Categories
Groups
Community Home
Categories
INTERNAL ENABLEMENT
POPULAR
PUBLIC CLOUD
PRIVATE CLOUD
Quick Links
MY LINKS
HELPFUL TIPS
Back to website
Home
Web CMS (TeamSite)
VFE and encoding
rwinterpacht
We have thousands of dcr's, with a Visualformat body field that was created using the utf-8 character set. Generating and pushing these files out, we're finding '?' characters in the browser, which may be related to how our Content servlet is bringing in the html from TeamSite.
One workaround to this problem is to open the dcr with the VFE config switched from "utf-8" to "latin". This makes a conversion in the body field, the user regenerates, republishes, and the page looks fine.
I'm looking for a way to do this either in PERL, or some other command line tool, without having to open each dcr to let the VFE (and the updated vfe config) take care of the problem.
Any ideas?
Raf Winterpacht
Household Intl/HSBC
TS 5.5.2 sp3
Find more posts tagged with
Comments
Dwayne
Can you set the <meta> tag in your generated page, telling the browser that the page is UTF-8 encoded? (Assuming it is, of course)
Another alternative would be in the presentation templates. Whenever the fields are read from the DCR, convert them to latin1 within the PT, and then spit out the converted text, instead of the "raw" text in the DCR. They way you would do that would vary somewhat, depending on which version of TeamSite you're using.
--
Current project: TS 5.5.2/6.1 W2K
Migrateduser
Have you identified which characters are being translated to "?" in UTF-8? If available, try to use an editor on this content that allows you to see all control characters that are not apparent via the web browser. Those will need to be translated, as well.
This can absolutely be done in Perl -- I've had to do it a couple of times, myself, but the first thing you'll need to know is the full set of characters that is getting manipulated.
Dave
rwinterpacht
Thanks for the info, I pretty much on that path, here's what I've tried in the iw_perl tags:
TeamSite::I18N_utils::utf8_normalize_string('ISO-8859-1', $outputHTML); <-- Didn't work
Then I tried going through a loop of known characters, by running it through the following method:
sub replaceSpecialChars {
my ( $origText ) =
@_
;
my $i = 0;
my
@characters
= ('á', 'á',
'Á', 'Á',
'â', 'â',
'Â', 'Â',
'à', 'à',
'À', 'À',
'å', 'å',
'Å', 'Å',
'ã', 'ã',
'Ã', 'Ã',
'ä', 'ä',
'Ä', 'Ä',
'æ', 'æ',
'Æ', 'Æ',
'ç', 'ç',
'Ç', 'Ç',
'ð', 'ð',
'Ð', 'Ð',
'é', 'é',
'É', 'É',
'ê', 'ê',
'Ê', 'Ê',
'è', 'è',
'È', 'È',
'ë', 'ë',
'Ë', 'Ë',
'í', 'í',
'Í', 'Í',
'î', 'î',
'Î', 'Î',
'ì', 'ì',
'Ì', 'Ì',
'ï', 'ï',
'Ï', 'Ï',
'ñ', 'ñ',
'Ñ', 'Ñ',
'ó', 'ó',
'Ó', 'Ó',
'ô', 'ô',
'Ô', 'Ô',
'ò', 'ò',
'Ò', 'Ò',
'ø', 'ø',
'Ø', 'Ø',
'õ', 'õ',
'Õ', 'Õ',
'ö', 'ö',
'Ö', 'Ö',
'ß', 'ß',
'þ', 'þ',
'Þ', 'Þ',
'ú', 'ú',
'Ú', 'Ú',
'û', 'û',
'Û', 'Û',
'ù', 'ù',
'Ù', 'Ù',
'ü', 'ü',
'Ü', 'Ü',
'ý', 'ý',
'Ý', 'Ý',
'ÿ', 'ÿ');
for ($i = 1; $i < $#characters; $i += 2)
{
($origText) =~ s/$characters[$i-1]/$characters[$i]/eg;
}
return $origText;
}
That didn't seem to work either, actually, it ended up putting the  characters where I saw the "?".
Raf Winterpacht
Household Intl/HSBC
TS 5.5.2 sp3
Migrateduser
I think you might just have a Perl issue at this point, which is good because it can be easily fixed.
A couple of things stand out to me:
1. Your array should really be a hash table or a two-dimensional array, since you're basically using it as such.
2. Even in an array context, your counter, $i, should start at 0 and this may actually be the reason that things are displaying funny. The array starts at $character[0] and, by definition, supports elements up to and including $character[$#character - 1]
Try that! It sounds like you're on the right path and including that in your PT should absolutely do the trick.
Dave
ela
Hi,
in case Dave's suggestion doesn't solve your problem, here's what I do:
--------------------------------------------------------------------------------------------------------------------------------
In the presentation tpl I call a separate file ReplaceChars.pm, which defines the characters to be replaced:
--------------------------------------------------------------------------------------------------------------------------------
<?xml version="1.0" encoding="UTF-8"?>
<iw_pt name="Standard Presentation Template">
<iw_perl>
<![CDATA[
use lib ('/data/home/iwts/Perllib/lib/perl5/perl5/site_perl/5.005', '/app/teamsite/iw-home/local/lib/prod');
use ReplaceChars;
]]>
</iw_perl>
<![CDATA[
<META content="text/html; charset=ISO-8859-1" http-equiv="Content-Type">
<!-- Visual Format Text Area -->
<iw_if expr="{iw_value name='section.text'/} =~ /\S/">
<iw_perl><![CDATA[
$text = iwpt_dcr_value('section.text');
$text =~ s|^\s*<p>||si;
$text =~ s|</p>\s*$||si;
$text =~ s|<p>||gsi;
$text =~ s|</p>|<br/>|gsi;
$text = ReplaceChars::replaceChars($text);
]]>
</iw_perl>
<iw_then>
<iw_value name="$texttext"/>
</iw_then>
</iw_if>
]]>
</iw_pt>
--------------------------------------------------------------------------------------------------------------------------------
In the ReplaceChars file:
--------------------------------------------------------------------------------------------------------------------------------
#!/app/teamsite/iw-home/iw-perl/bin/iwperl
package ReplaceChars;
sub replaceChars {
my $text = shift;
$text =~ s|·|\·\;|gsi;
$text =~ s|Å“|\œ\;|gsi;
$text =~ s|«|\&\#171\;|gsi;
$text =~ s|»|\&\#187\;|gsi;
return $text;
}
1;
--------------------------------------------------------------------------------------------------------------------------------
Good luck!
Eldbjoerg
rwinterpacht
Vielen dank für das hilfe! Das probiere Ich schon bald.
Bis später!
Raf Winterpacht
Household Intl/HSBC
TS 5.5.2 sp3
Migrateduser
You're welcome, but I'm curious -- which method worked out for you?
Dave
rwinterpacht
Before proceeding with working this out on the content side, we may put in a fix on the servlet.
It's funny that this came up for our WAS 5 upgrade...some tweaks had to be done for the content delivery, and apparently the encoding wasn't taken into consideration.
I still do appreciate the suggestions, and we may end up using them if the servlet fix doesn't work...will let you know.
Raf Winterpacht
Household Intl/HSBC
TS 5.5.2 sp3
ela
I looked at one of your former posts, where you write that
it ended up putting the  characters where I saw the "?"
. From your function I saw that the replaced character would then be
Â
. This is the utf-8 character for
. Utf-8 for
Â
is
ÃÂ
.
Could it be that you only need to replace the
Â
, or are all the characters you listed being displayed incorrectly?
Eldbjoerg
rwinterpacht
All characters are being displayed incorrectly. Without doing any cleanup (by passing body field to PERL methods for cleanup), special characters are appearing as ?. Special characters such as the accented "e" are for French, and also characters. By the time the html is deployed to the application server, the characters are "A0" in hex, so should be displayed by the servlet...but are not.
Our servlet is treating the html in UNICODE, not UTF-8 or ISO-8859-1, which is a default for how InputStream is handled. So we have to override the load in the right encoding, and then our problem should be fixed...again, testing this out and will let you know.
Thanks!
Raf Winterpacht
Household Intl/HSBC
TS 5.5.2 sp3
rwinterpacht
As it turns out, making the change in the servlet, and how we output the html, made all the difference. We no longer see the characters displayed as "?".
Thanks for all your help!
Raf Winterpacht
Household Intl/HSBC
TS 5.5.2 sp3