Discussions
Categories
Groups
Community Home
Categories
INTERNAL ENABLEMENT
POPULAR
PUBLIC CLOUD
PRIVATE CLOUD
Quick Links
MY LINKS
HELPFUL TIPS
Back to website
Home
Web CMS (TeamSite)
Perl -R implementation
nipper
So I decided to benchmark various ways to run a script, recursively, over a directory structure. I have almost 19,000 DCRs. There are (so I have seen) 4 ways, Smitty's recursive script, Fish's improvement (reference to the array rather than an array), my implementation (ls -R plus some parsing) and Binko's suggestion of File::Find
Here are my results:
Done File::Find
real 0m39.90s
user 0m0.59s
sys 0m2.92s
Done Rec array
real 0m37.26s
user 0m1.10s
sys 0m4.49s
Done Rec reference
real 0m32.00s
user 0m0.56s
sys 0m3.63s
Done LS
real 0m28.01s
user 0m0.45s
sys 0m3.83s
OK, so I win. :-)
One point due to filesystem caching, if you rerun this script, all of them drop to 4 or 5 seconds.
Attached is my code for your ridicule.
Andy
Find more posts tagged with
Comments
Adam Stoller
What happend to the other alternative I provided - getting just the modified files using iwlistmod? :-) (That seemed to be what you wanted).
Also - you might want to check out: perldoc Benchmark, for ways of doing better benchmarking of perl code.
--fish
Senior Consultant, Quotient Inc.
http://www.quotient-inc.com
ExampleXMLIN.rar
Migrateduser
nipper,
I think you missed the central point of my recommendation. If you are going to process each file independently, then you don't need to build the list of files.
I modified your test script so that each of the methods simply counted the files, and then ran the test over 10,000 files. The ls method (#1) took 4 minutes to count 10,000 files under a non-TeamSite directory, whereas Find::File took 7 seconds.
But fish is right (he usually is). If you only need to consider modified files (and if they are a small subset of all of the files in the directory) then you should probably use iwlistmod.
Brinko Kobrin
Interwoven Staff Engineer
Migrateduser
BTW, the other recursive algorithms were even faster than Find::File. They took about 3 seconds each. But my point about Find::File was that you would not have to write as much code.
Brinko Kobrin
Interwoven Staff Engineer
nipper
Fish & Brinko - point taken.
While the example I posted before was only for modified files, I have had the need to to checking of every DCR. THat is why I have pursued this. One application was that we added a new required field to the DCR. I was required to go through ALL 18K+ DCRs and detemine the contents & add the field. I did make the change Fish suggested (lsmod), I was just being lazy and reusing code.
Brinko - I do agree that File::Find is not significantly slower than the other options& certainly does make it easier. But reading Fish's perl WP, I wanted to see how it compared to a couple of other options.
I also understand I did not need to put the find results into an array, I just wanted to make as close to an apples to apples comparison.
Andy
Johnny
If you want a few more tricks, TeamSite:
irwalk is a really good OO option to File::Find.
John Cuiuli