Thursday, March 31, 2011

use encoding 'utf8" example

#!/usr/bin/perl 

use warnings;
use strict;

# Sort strings properly and match all letters
use locale;

# Read the source file as UTF-8 and set STDOUT and STDIN to UTF-8
use encoding 'utf8';


# We can write letters in the source code in UTF-8 thanks to "use encoding"
my @pole=qw(šiška marek ucho čaj žička);

# Sort knows our alphabet thanks to "use locale"
@pole=sort(@pole);           # sort array

foreach my $a (@pole) {
    # The pattern matching works fine thanks to "use locale"
    $a =~ s/\W//g;           # remove all non word characters
    print "$a \n";           # print the array field
}

# Write in UTF-8 into the file
open my $handle, ">:utf8", "file.txt" or die "Can't write to file.txt: $!";
# We can convert to upper case thanks to "use locale"
print $handle "\Uěščřžabcd\E\n";     # write upper cased string
close $handle;

use encoding 'utf8' example

#!/usr/bin/perl 

use warnings;
use strict;

# Sort strings properly and match all letters
use locale;

# Read the source file as UTF-8 and set STDOUT and STDIN to UTF-8
use encoding 'utf8';


# We can write letters in the source code in UTF-8 thanks to "use encoding"
my @pole=qw(šiška marek ucho čaj žička);

# Sort knows our alphabet thanks to "use locale"
@pole=sort(@pole);           # sort array

foreach my $a (@pole) {
    # The pattern matching works fine thanks to "use locale"
    $a =~ s/\W//g;           # remove all non word characters
    print "$a \n";           # print the array field
}

# Write in UTF-8 into the file
open my $handle, ">:utf8", "file.txt" or die "Can't write to file.txt: $!";
# We can convert to upper case thanks to "use locale"
print $handle "\Uěščřžabcd\E\n";     # write upper cased string
close $handle;

Perl Script: Delete All Files and Folders in a Directory

use strict;
use warnings;

# Start the recursive cleanup process on the target directory
cleanup("d:/tempo");

sub cleanup {
    # Get the directory path passed to the subroutine
    my $dir = shift;
    
    # Localize the directory handle and open the directory
    local *DIR;
    opendir(DIR, $dir) or die "Cannot open directory $dir: $!";
    
    # Loop through each item (file or folder) inside the directory
    for my $file (readdir(DIR)) {
        
        # Skip the special current (.) and parent (..) directory entries
        next if $file =~ /^\.{1,2}$/;
        
        # Construct the full path to the current item
        my $path = "$dir/$file";
        
        # If the item is a regular file, delete (unlink) it
        unlink $path if -f $path;
        
        # If the item is a directory, call this subroutine recursively to empty it
        cleanup($path) if -d $path;
        
        print "Deleting $path\n";   
    }
    
    # Close the directory handle once we are done reading its contents
    closedir(DIR);
    
    # Finally, remove the now-empty directory itself
    print "Removing directory $dir\n";
    rmdir $dir or print "Error removing $dir: $!\n";
}

The Pattern Component Order of Precedence

Precedence Level, Component, and Examples

Precedence LevelComponentExamples
1Parentheses(...), (?:...), (?=...)
2Quantifiers*, +, ?, {n}, {n,m}
3Sequences and AnchorsSequences (abc), Anchors (^, $, \b, \A, \z)
4Alternation`

Friday, March 25, 2011

Perl Quick Tip: Catching Fatal Errors with 'eval'

use strict;
use warnings;
use XML::LibXML;

# Initialize the XML parser
my $parser = XML::LibXML->new();

my $tree;
my $file = 'example.xml'; # (Assuming $file is defined somewhere in your script)

# The 'eval' block catches fatal errors so the script can continue running.
eval {
    # Parses the file contents into the new libXML object.
    # If the file is missing or contains invalid XML, this throws an exception.
    $tree = $parser->parse_file($file); 
};

# $@ is a special Perl variable that holds the error message from the last eval block.
# If eval succeeded, $@ will be empty (false). If it failed, $@ contains the error string.
if ($@) {
    warn "Error encountered: $@";
}

How to Count Duplicate Elements in a Perl Array

use strict;
use warnings;

# 1. Initialize an array using the qw() operator. 
# qw() stands for "quote words" and automatically creates a list of strings separated by spaces.
my @array = qw(foo bar foo bar baz foo baz bar foo);

# 2. Create an empty hash to store the count of each element.
# Hashes store data in key-value pairs. Here, the 'key' will be the word, and the 'value' will be its count.
my %counts = ();

# 3. Loop through every element in the array.
for (@array) {
    # In a 'for' loop without a named variable, Perl temporarily stores the current item in the special variable $_
    # This line looks up the current word in the hash and increments its value by 1.
    # If the word isn't in the hash yet, Perl automatically creates it with a starting value of 0, then adds 1.
    $counts{$_}++;
}

# 4. Iterate over the sorted keys of our populated hash.
# 'keys %counts' returns a list of all the unique words we found.
foreach my $key (keys %counts) {
    # Print the word (the key) and how many times it appeared (the value)
    print "$key = $counts{$key}\n";
}

Sorting Human-Readable Dates in Perl with Date::Parse

#!/usr/bin/perl 
use warnings;
use strict;

use Date::Parse;

# Read from the DATA block, sort chronologically, and print
print for sort { str2time($a) <=> str2time($b) } <DATA>;

__DATA__
March 9, 2010
April 2, 2008
January 23, 2009
April 1, 2008

reading .ini file without using a module

use strict;
use warnings;
use Data::Dumper;

open my $fh, '<', "my.ini" or die "$!\n";
my $ini = {};

{
local $/ = "";
while (<$fh>) {
next unless ( s/\[([^]]+)\]// );
my $header = $1;
$ini->{$header} = {};
while (m/(\w+)=(.*)/g) {
$ini->{$header}->{$1} = $2;
}
}
}
close $fh;

print Dumper $ini;


__DATA__
[Build]
BuildID=20110209115208
Milestone=2.0b11
SourceStamp=f9d66f4d17bf
SourceRepository=http://hg.mozilla.org/mozilla-central
; This file is in the UTF-8 encoding

[Strings]
Title=SeaMonkey Update
Info=SeaMonkey is installing your updates and will start in a few
moments

Tie::IxHash Example in perl

use Tie::IxHash;

tie my %unique => 'Tie::IxHash';

my @duplicates = (1, 1, 2, 3, 4, 5, 5, 6, 7, 7, 8, 9, 2, 3, 3,);

@unique{@duplicates} = ();

my @unique_elements = keys %unique;

print "@unique_elements\n";

How to I convert a string to a number (ASCII to Integer)

#How to I convert a string to a number

use strict;
use warnings;

sub atoi {
    # Initialize to 0 to avoid "uninitialized value" warnings
    my $t = 0; 
    
    # Process each character of the first argument
    foreach my $d (split(//, shift())) {
        $t = $t * 10 + $d;
        # Optional: Comment out the next line if you don't want to print intermediate steps
        # print "$t\n"; 
    }
    return $t;
}

# Declare $number with 'my' to comply with 'use strict'
my $number = atoi("123");

print "string '123' number value is $number\n";

How to compare 2 arrays and differences store in 3 array

#!/usr/bin/perl
use strict;
use warnings;
use List::Compare;

my @new = qw/a b c d e/;
my @old = qw/a b   d e f/;

my $lc = List::Compare->new(\@new, \@old);

# an array with the elements that are in @new and not in @old : c
my @Lonly = $lc->get_Lonly;
print "\@Lonly: @Lonly\n";

# an array with the elements that are not in @new and in @old : f
my @Ronly = $lc->get_Ronly;
print "\@Ronly: @Ronly\n";

# an array with the elements that are in both @new and @old : a b d e
my @intersection = $lc->get_intersection;
print "\@intersection: @intersection\n";

__END__
** prints

@Lonly: c
@Ronly: f
@intersection: a b d e


OR


#!/usr/bin/perl
use strict;
use warnings;
use Data::Dumper;

my @new = qw/a b c d e/;
my @old = qw/a b   d e f/;
my %new = map{$_ => 1} @new;
my %old = map{$_ => 1} @old;

my (@new_not_old, @old_not_new, @new_and_old);
foreach my $key(@new) {
    if (exists $old{$key}) {
        push @new_and_old, $key;
    } else {
        push @new_not_old, $key;
    }
}
foreach my $key(@old) {
    if (!exists $new{$key}) {
        push @old_not_new, $key;
    }
}

print Dumper\@new_and_old;
print Dumper\@new_not_old;
print Dumper\@old_not_new;

Runtime vs Compile time

The difference between compile time and run time is an example of what pointy-headed theorists call the phase distinction. It is one of the hardest concepts to learn, especially for people without much background in programming languages. To approach this problem, I find it helpful to ask

    What invariants does the program satisfy?
    What can go wrong in this phase?
    If the phase succeeds, what are the postconditions (what do we know)?
    What are the inputs and outputs, if any?

Compile time

    The program need not satisfy any invariants. In fact, it needn't be a well-formed program at all. You could feed this HTML to the compiler and watch it barf...
    What can go wrong at compile time:
        Syntax errors
        Typechecking errors
        (Rarely) compiler crashes
    If the compiler succeeds, what do we know?
        The program was well formed---a meaningful program in whatever language.
        It's possible to start running the program. (The program might fail immediately, but at least we can try.)
    What are the inputs and outputs?
        Input was the program being compiled, plus any header files, interfaces, libraries, or other voodoo that it needed to import in order to get compiled.
        Output is hopefully assembly code or relocatable object code or even an executable program. Of if something goes wrong, output is a bunch of error messages.

Run time

    We know nothing about the program's invariants---they are whatever the programmer put in. Run-time invariants are rarely enforced by the compiler alone; it needs help from the programmer.

    What can go wrong are run-time errors:
        Division by zero
        Deferencing a null pointer
        Running out of memory

    Also there can be errors that are detected by the program itself:
        Trying to open a file that isn't there
        Trying find a web page and discovering that an alleged URL is not well formed
    If run-time succeeds, the program finishes (or keeps going) without crashing.
    Inputs and outputs are entirely up to the programmer. Files, windows on the screen, network packets, jobs sent to the printer, you name it. If the program launches missiles, that's an output, and it happens only at run time

Difference between compile time and run time

The difference between compile time and run time is known as the "phase distinction." To understand how a program behaves in each phase, it is helpful to evaluate its invariants, potential errors, postconditions, and inputs/outputs.

Compile Time

  • Invariants: The program does not need to satisfy any invariants and does not even need to be well-formed initially.

  • What Can Go Wrong:

    • Syntax errors

    • Typechecking errors

    • Compiler crashes (rare)

  • Postconditions (If Successful):

    • The program is deemed well-formed.

    • It is ready to be executed (though it may still fail upon running).

  • Inputs and Outputs:

    • Inputs: The source code, along with necessary header files, interfaces, and imported libraries.

    • Outputs: Assembly code, relocatable object code, an executable program, or error messages if compilation fails.

Run Time

  • Invariants: The environment assumes nothing; invariants are solely enforced by the programmer's logic.

  • What Can Go Wrong:

    • Fatal run-time errors: Division by zero, dereferencing a null pointer, or running out of memory.

    • Program-handled errors: Attempting to open a missing file or processing a malformed URL.

  • Postconditions (If Successful): The program successfully finishes its tasks or continues operating without crashing.

  • Inputs and Outputs: Entirely dictated by the programmer. This includes interacting with files, UI windows, network packets, or external hardware (which only occurs during run time).