Discussions
Categories
Groups
Community Home
Categories
INTERNAL ENABLEMENT
POPULAR
PUBLIC CLOUD
PRIVATE CLOUD
Quick Links
MY LINKS
HELPFUL TIPS
Back to website
Home
Web CMS (TeamSite)
another regex
nipper
Since I cannot ge tthis stupid PDF pm to work, I figured I would try regex myself.
Trying to extract Keywords, Author, Title, and Subject.
strings blah.pdf | grep Author
/ModDate (D:20040630211317-07') /Author (Andy) /Title (AN2383/D: A Smart Antenna System for 3G Wireless Using the MSC8102 DSP Device) /Keywords (StarCore, 3G, MSC8101, MSC8102, smart antenna) /Producer (Acrobat Distiller 4.0 for Windows) /Creator (FrameMaker 5.5P4f) /CreationDate (D:20021209132352)
So I tried slurping the PDF into $txt
$txt =~ m{(?:/Author|/Keywords|/Subject|/Title): \ * (.*)/}gx;
$hit= $1;
print "Got $hit\n"
but it is always blank.
I was hoping it would hit Author or TIltle or Keywods or Subject.
Also tried just:
$txt =~ m{/Author\ * (.*)/}g;
$author = $1;
get a bunch of garbage.
if I make it non-greedy I get nothing.
Tips/Pointers/RTFMs appreciated.
Andy
Find more posts tagged with
Comments
herald10
How about
$txt =~ m/Author\s*\(([^)]*)\)/;
print "$1";
nipper
Looks good, thanks.