Perl for Text Processing
Combine regular expressions and file handling to build a log-analysis script that counts error occurrences by severity level.
Introduction
Text processing is where Perl's regular expression engine and file-handling tools come together most powerfully. This lesson builds a small but genuinely useful log analyzer: it reads a log file, extracts the severity level from each line, and reports how often each level occurred.
- How to design a regex that extracts structured fields from a log line.
- How to tally occurrences using a hash as a counter.
- How to combine file reading and regex matching into one script.
- How named captures make regex-heavy code easier to read.
Why Perl Excels at Text Processing
Regular expressions are a first-class part of Perl's syntax, not a bolted-on library, which makes pattern matching, extraction, and substitution feel completely natural. Combined with simple, fast file I/O, this is why Perl became the default tool for chewing through log files and text reports long before dedicated log-analysis platforms existed.
Anatomy of a Log Line
The examples below assume a log file with lines in this common format: a timestamp, a severity level in brackets, and a message.
2026-08-01 09:12:03 [INFO] Server started2026-08-01 09:15:41 [ERROR] Database connection failed2026-08-01 09:16:02 [WARN] Retrying connection2026-08-01 09:16:05 [INFO] Connection restored2026-08-01 09:20:11 [ERROR] Timeout while contacting payment gateway2026-08-01 09:25:47 [ERROR] Disk space below 10%2026-08-01 09:30:00 [INFO] Backup completedMatching Log Entries with Regex
A regex with two capture groups can pull the level and the message out of each line in one pass: \[(\w+)\] matches the bracketed level, and \s+(.*) captures everything after it as the message.
use strict;use warnings;
my $line = '2026-08-01 09:15:41 [ERROR] Database connection failed';
if ($line =~ /\[(\w+)\]\s+(.*)/) { print "Level: $1\n"; print "Message: $2\n";}Click Run to see what this code prints.
Counting Occurrences with a Hash
A hash makes an excellent tally counter: $counts{$level}++ automatically creates the key on first use (starting from undef, which increments to 1) and increments it on every subsequent match.
Building the Full Log Analyzer
Putting file reading, regex matching, and hash counting together produces a small but genuinely useful script: a summary of log levels, plus a list of every error message encountered.
use strict;use warnings;
my $logfile = shift @ARGV // 'app.log';
open(my $fh, '<', $logfile) or die "Cannot open '$logfile': $!\n";
my %counts;my @errors;
while (my $line = <$fh>) { chomp $line; if ($line =~ /\[(\w+)\]\s+(.*)/) { my ($level, $message) = ($1, $2); $counts{$level}++; push @errors, $message if $level eq 'ERROR'; }}close $fh;
print "=== Log Level Summary ===\n";for my $level (sort { $counts{$b} <=> $counts{$a} } keys %counts) { printf "%-10s %d\n", $level, $counts{$level};}
if (@errors) { print "\n=== Error Messages ===\n"; print "- $_\n" for @errors;}Click Run to see what this code prints.
Using Named Captures
For regexes with several groups, numbered captures like $1 and $2 quickly become hard to read. Named captures - (?<name>...) - let you refer to a group by a meaningful name through the %+ hash instead.
use strict;use warnings;
my $line = '2026-08-01 09:15:41 [ERROR] Database connection failed';
if ($line =~ /\[(?<level>\w+)\]/) { print "Level: $+{level}\n";}Click Run to see what this code prints.
Common Mistakes
- Writing a regex that is too greedy or too loose, matching more (or less) of the line than intended.
- Forgetting to chomp lines, leaving a trailing newline stuck to the captured message.
- Matching levels case-sensitively (ERROR vs error) when logs are not guaranteed to be consistent - use /i or normalize with uc/lc.
- Slurping an entire multi-gigabyte log file into memory instead of streaming it line by line.
- Hardcoding the log format so tightly that a single formatting change breaks every match.
Best Practices
- Stream log files line by line rather than reading them entirely into memory.
- Anchor and scope regexes as precisely as the data allows, to avoid accidental over-matching.
- Use named captures for regexes with more than one or two groups.
- Separate the parsing logic (extracting fields) from the reporting logic (formatting output).
- Consider case-insensitive matching (/i) unless the log format is strictly guaranteed to be consistent.
Frequently Asked Questions
Yes - modules like IO::Uncompress::Gunzip, or simply opening a pipe from a command like zcat, let you read compressed logs without decompressing them to disk first.
Always stream line by line with a while loop rather than slurping the whole file, and avoid building large in-memory arrays unless you truly need every line at once.
No - for structured formats like JSON or CSV, a dedicated parser module (JSON, Text::CSV) is more robust than a hand-written regex.
Key Takeaways
- Capture groups let you pull structured fields out of unstructured log lines.
- A hash with auto-incrementing values is a simple, effective tally counter.
- Streaming line by line keeps memory usage low even for large log files.
- Named captures (?<name>...) make multi-group regexes far more readable.
Summary
Regex plus file handling is a combination you will reach for constantly in Perl - it is fast to write and powerful enough for real production log analysis. Next, you will learn how to verify that code like this actually works, using Perl's built-in testing tools.