LearnAI ToolsCareerPractice BuildsPlayContact
PerlBeginner~1.5 hours

Log File Analyzer

Parse a web server log with regular expressions and summarize errors by hour.

Regular ExpressionsFile HandlingHashes

Overview

Perl was built for exactly this kind of job: chewing through a text file line by line and pulling structured facts out of it with a regular expression. A web server access log is a stream of lines that all follow the same shape — client address, timestamp, request, status code, size — and a regex with capture groups is the natural tool for pulling the pieces you care about out of that shape without writing a hand-rolled parser.

By the end of this tutorial you will have a script that reads a log file line by line with an open file handle, matches each line against a regular expression built to mirror the Common Log Format, and tallies every 4xx and 5xx response into a hash keyed by the hour it happened in. Every step leans on the three habits every Perl script should have from the very first line: `use strict; use warnings;` to catch mistakes early, and checked `open` calls so a missing file fails loudly instead of silently.

What You'll Build
  • A regex with named-feeling capture groups that matches one Common Log Format line.
  • An `open` call wrapped in error checking with `or die`, using the three-argument form.
  • A `while (<$fh>)` loop that processes the log one line at a time without loading the whole file into memory.
  • A hash keyed by hour (`"00"` through `"23"`) counting how many 4xx and 5xx responses occurred in each.
  • A sorted summary report printed with `sort` and `printf` for aligned columns.

Prerequisites

  • Scalars and arrays — Perl's basic `$scalar` and `@array` variable types.
  • Hashes — `%hash`, keying values with `$hash{$key}`, and iterating with `keys`/`values`.
  • Regular expressions — `=~`, capture groups `(...)`, and Perl's `$1`, `$2` match variables.
  • File handling — `open`, `close`, and reading a file a line at a time with `<$fh>`.
  • Basic control flow — `while`, `foreach`, and `if`/`unless`.

Project Structure

The whole project is one script, `log_analyzer.pl`, plus a sample log file, `access.log`, sitting next to it. The script is organized top to bottom exactly the way it runs: open the log file, loop over its lines matching each one against a regex, accumulate counts into a hash, then loop over that hash a second time to print the report. There is no need for subroutines or modules here — the entire job is a single pass over a file followed by a single pass over a hash, which is about as close to Perl's original one-liner-turned-script roots as a project gets.

`use strict;` and `use warnings;` appear on the first two lines and stay there for the rest of the tutorial. `strict` forces every variable to be declared with `my` before it is used, which catches typos in variable names at compile time instead of letting them silently create a new global. `warnings` flags things like using an undefined value in a numeric context, which is exactly the kind of mistake a log line that does not match the expected format can trigger.

Step 1: Open the Log File Safely

Perl's three-argument `open` — `open(my $fh, '<', $path)` — is the modern idiom: the mode (`'<'` for reading) is kept separate from the filename, so a filename that happens to contain characters like `>` or `|` can never be misread as part of the mode. `open` returns a true value on success and a false value on failure, so `or die "..."` is the standard Perl pattern for turning a failed open into an immediate, informative error instead of a confusing failure several lines later.

use strict; # Require every variable to be declared with 'my' before use
use warnings; # Warn about risky things like using undef in a numeric context
my $log_file = 'access.log'; # Path to the log file this script analyzes
# Three-argument open: mode ('<' for read) is separate from the filename,
# so a filename containing '>' or '|' can never be mistaken for a mode.
# "or die" aborts immediately with a clear message if the file can't be opened,
# instead of silently continuing with an unusable file handle.
open(my $fh, '<', $log_file)
or die "Could not open '$log_file': $!\n"; # $! holds the OS error string
# ... regex matching and hour tallying happen here in the next steps ...
close($fh); # Release the file handle once every line has been processed

Step 2: Match Each Line With a Regex

Reading `<$fh>` inside a `while` condition pulls one line at a time out of the file handle, which keeps memory usage flat even for a multi-gigabyte log — the whole file is never loaded into a single string or array at once. The regex below is built to match a Common Log Format line such as `192.168.1.10 - - [10/Aug/2026:14:23:01 +0000] "GET /index.html HTTP/1.1" 404 512`, with parentheses marking every piece the rest of the script needs to pull out.

# Matches a Common Log Format line and captures the two-digit hour and the
# three-digit HTTP status code. Everything outside a capture group here is
# just structural text the regex must see but doesn't need to remember.
my $log_pattern = qr/
^\S+ \s+ \S+ \s+ \S+ \s+ # client IP, ident, userid (ignored)
\[ \d{2}\/\w{3}\/\d{4}:(\d{2}):\d{2}:\d{2} \s [+-]\d{4} \] \s+ # (1) hour
"[^"]*" \s+ # the quoted request line (ignored)
(\d{3}) \s+ # (2) HTTP status code
\S+ # response size (ignored)
/x; # /x lets the pattern span multiple lines with whitespace and comments
while (my $line = <$fh>) { # Reads and processes one line per iteration
chomp $line; # Strip the trailing newline before matching
if ($line =~ $log_pattern) {
my ($hour, $status) = ($1, $2); # $1/$2 hold the two capture groups on a successful match
# ... tally this hour/status pair in the next step ...
}
# Lines that don't match the expected format are silently skipped —
# a real log can contain the occasional malformed or truncated line.
}
Example Usage

Click Run to see what this code prints.

Step 3: Extract the Hour and Status Code

Step 2 already captures the hour and status code into `$1` and `$2`; this step is about deciding what counts as an "error" worth tallying. HTTP status codes in the 400s mean a client error and the 500s mean a server error, so checking the first digit of the three-digit code with a small regex — `$status =~ /^[45]/` — is enough to classify a response as an error without needing a lookup table of every individual code.

# A status code is an "error" if it starts with 4 (client error) or 5 (server
# error) — 2xx success and 3xx redirect codes are not counted here.
sub is_error_status {
my ($status) = @_; # @_ holds this subroutine's arguments
return $status =~ /^[45]/; # Matches if the first digit is 4 or 5
}
# Used inside the while loop from Step 2:
if ($line =~ $log_pattern) {
my ($hour, $status) = ($1, $2);
if (is_error_status($status)) {
print "Error at hour $hour: status $status\n"; # Replaced by hash tallying in Step 4
}
}

Step 4: Tally Errors in a Hash Keyed by Hour

A Perl hash is the natural structure for "count how many times each hour appears" — `$errors_by_hour{$hour}++` both creates the key with a starting value of `0` the first time a given hour is seen and increments it on every subsequent match, because an undefined hash value autovivifies to `0` the moment `warnings` lets it be used in a numeric context like `++`. No pre-initialization loop over all 24 hours is needed; the hash only ever grows entries for hours that actually appear in the log.

my %errors_by_hour; # Hash: hour string ("00".."23") => count of 4xx/5xx responses that hour
my $total_lines = 0; # Total lines read, matching or not — used in the summary
my $total_errors = 0; # Total error responses across every hour
while (my $line = <$fh>) {
chomp $line;
$total_lines++;
if ($line =~ $log_pattern) {
my ($hour, $status) = ($1, $2);
if (is_error_status($status)) {
$errors_by_hour{$hour}++; # Autovivifies to 1 on the first error seen for this hour
$total_errors++;
}
}
}

Step 5: Print a Sorted Summary Report

`keys %errors_by_hour` returns the hash's keys in no particular order, so `sort` is needed before printing anything readable. Because the hour keys are two-digit strings like `"09"` and `"14"`, Perl's default `sort` (which compares as strings) already puts them in the right chronological order — the leading zero keeps `"09"` sorting before `"10"` the same way a numeric sort would. `printf` then lines up the hour and count into fixed-width columns.

print "\n===== ERROR SUMMARY BY HOUR =====\n";
printf "%-8s %s\n", "Hour", "Errors";
foreach my $hour (sort keys %errors_by_hour) { # String sort works here because hours are zero-padded "00".."23"
printf "%-8s %d\n", "$hour:00", $errors_by_hour{$hour};
}
print "\nTotal lines processed: $total_lines\n";
print "Total error responses: $total_errors\n";

Complete Code

Here is the full script assembled in the order it runs, ready to save as `log_analyzer.pl` and run with `perl log_analyzer.pl` alongside an `access.log` file in the same directory.

use strict;
use warnings;
my $log_file = 'access.log';
open(my $fh, '<', $log_file)
or die "Could not open '$log_file': $!\n";
# Matches a Common Log Format line and captures the two-digit hour and the
# three-digit HTTP status code.
my $log_pattern = qr/
^\S+ \s+ \S+ \s+ \S+ \s+
\[ \d{2}\/\w{3}\/\d{4}:(\d{2}):\d{2}:\d{2} \s [+-]\d{4} \] \s+
"[^"]*" \s+
(\d{3}) \s+
\S+
/x;
# A status code is an "error" if it starts with 4 (client error) or 5 (server error).
sub is_error_status {
my ($status) = @_;
return $status =~ /^[45]/;
}
my %errors_by_hour; # hour string ("00".."23") => count of 4xx/5xx responses
my $total_lines = 0;
my $total_errors = 0;
while (my $line = <$fh>) {
chomp $line;
$total_lines++;
if ($line =~ $log_pattern) {
my ($hour, $status) = ($1, $2);
if (is_error_status($status)) {
$errors_by_hour{$hour}++;
$total_errors++;
}
}
}
close($fh);
print "\n===== ERROR SUMMARY BY HOUR =====\n";
printf "%-8s %s\n", "Hour", "Errors";
foreach my $hour (sort keys %errors_by_hour) {
printf "%-8s %d\n", "$hour:00", $errors_by_hour{$hour};
}
print "\nTotal lines processed: $total_lines\n";
print "Total error responses: $total_errors\n";

Sample Run

Sample Run

Click Run to see what this code prints.

Extend This Project

  • Break the error tally into two hashes, `%client_errors_by_hour` and `%server_errors_by_hour`, so 4xx and 5xx are reported separately.
  • Add a hash-of-hashes keyed by hour and then by exact status code (e.g. `$errors{"14"}{"404"}++`) for a finer-grained breakdown.
  • Accept the log filename from `@ARGV` instead of hardcoding `access.log`, so the script can analyze any file passed on the command line.
  • Track the top 5 most-requested URLs by adding a second regex capture group for the request path and tallying it in its own hash.
  • Use `Time::Piece` to parse the full timestamp and report errors bucketed by day as well as by hour.

Summary

You built a log analyzer that reads a file one line at a time, pulls structured data out of unstructured text with a regex and capture groups, and summarizes it with a hash keyed by hour. This read-match-tally-report shape — an `open`ed file handle feeding a `while` loop, a regex doing the extraction, and a hash doing the counting — is the backbone of a huge share of real-world Perl scripts, from log analysis to report generation.