LearnAI ToolsCareerPractice BuildsPlayContact
Lesson 1619 min read

Regular Expressions

Learn the =~ binding operator, m// matching, common regex metacharacters, and how to capture parts of a match in Perl.

Introduction

Regular expressions (regex) describe patterns of text, letting you test whether a string matches a shape, and pull specific pieces out of it. Perl's regex support is not a bolted-on library — it is a core part of the language's grammar, which is why Perl remains one of the most natural languages for pattern-based text processing.

This lesson focuses on matching: testing whether a string fits a pattern and capturing pieces of it. The next lesson builds on this to cover substitution and translation.

What You Will Learn
  • How the =~ operator binds a string to a pattern.
  • How m// matches a pattern against a string.
  • The most common regex metacharacters and character classes.
  • How to capture parts of a match with parentheses.

The =~ Binding Operator

The =~ operator tells Perl "apply this pattern-matching operation to this variable." Without it, regex operators like m// implicitly apply to the default variable $_. With =~, you can target any variable you like.

use strict;
use warnings;
my $email = "alice\@example.com";
if ($email =~ /\@/) {
print "Looks like it contains an \@ sign.\n";
} else {
print "Missing the \@ sign.\n";
}
Program Output

Click Run to see what this code prints.

The Negated Form

There is also !~, which returns true when the pattern does NOT match: if ($email !~ /\@/) { print "Invalid email\n"; }

The m// Match Operator

m// (the m is often omitted when using slashes, so /pattern/ alone is shorthand for m/pattern/) tests whether a string matches a pattern. You can use a different delimiter besides slashes — handy when your pattern itself contains slashes, such as file paths — by writing m{pattern} or m!pattern! instead.

use strict;
use warnings;
my $path = "/usr/local/bin/perl";
if ($path =~ m{/local/}) {
print "Path includes /local/\n";
}
my $sentence = "Perl was released in 1987.";
if ($sentence =~ /\d{4}/) {
print "Found a 4-digit year: $&\n";
}
Program Output

Click Run to see what this code prints.

The special variable $& holds the exact text that the whole pattern matched, which is convenient for quick debugging, though relying on named or numbered captures is usually clearer for real code.

Common Metacharacters

Metacharacters are symbols in a regex that have special meaning rather than matching themselves literally. Here are the ones you will use constantly.

MetacharacterMeaning
.Matches any single character except a newline.
^Anchors the match to the start of the string (or line, in multiline mode).
$Anchors the match to the end of the string (or line, in multiline mode).
\dMatches any digit (0-9).
\wMatches any "word" character: letters, digits, or underscore.
\sMatches any whitespace character (space, tab, newline).
\bMatches a word boundary (the edge between \w and non-\w characters).
use strict;
use warnings;
my @lines = ("apple: 5", "banana", "cherry: 12");
foreach my $line (@lines) {
if ($line =~ /^\w+: \d+$/) {
print "$line -- matches the pattern\n";
} else {
print "$line -- does not match\n";
}
}
Program Output

Click Run to see what this code prints.

Character Classes

Square brackets define a character class: a set of characters where any one of them satisfies that position in the pattern. A caret right after the opening bracket negates the class, matching anything NOT in the set. A hyphen inside the brackets defines a range, like [a-z] or [0-9].

use strict;
use warnings;
my @codes = ("A1", "Z9", "!!", "b2");
foreach my $code (@codes) {
if ($code =~ /^[A-Z][0-9]$/) {
print "$code is a valid code.\n";
} else {
print "$code is NOT valid.\n";
}
}
Program Output

Click Run to see what this code prints.

Quantifiers

Quantifiers control how many times the preceding token may repeat: * means zero or more, + means one or more, ? means zero or one (optional), and {n,m} means between n and m repetitions.

use strict;
use warnings;
my $phone = "call 555-2671 or 555-9012 today";
if ($phone =~ /\d{3}-\d{4}/) {
print "Found a phone number pattern.\n";
}
my $optional = "color";
my $optional2 = "colour";
foreach my $word ($optional, $optional2) {
print "$word matches\n" if $word =~ /colou?r/;
}
Program Output

Click Run to see what this code prints.

Capture Groups

Parentheses in a pattern do double duty: they group parts of the pattern together, and they capture whatever matched inside them into numbered variables $1, $2, $3, and so on, in the order the opening parentheses appear.

use strict;
use warnings;
my $date = "2026-08-05";
if ($date =~ /^(\d{4})-(\d{2})-(\d{2})$/) {
print "Year: $1\n";
print "Month: $2\n";
print "Day: $3\n";
}
Program Output

Click Run to see what this code prints.

Named Captures

For clarity in longer patterns, use named captures: /(?<year>\d{4})-(?<month>\d{2})-(?<day>\d{2})/ and then access them via %+, e.g. $+{year}.

Common Mistakes

Avoid These Mistakes
  • Forgetting to escape special regex characters like . or $ when you mean them literally (use \. and \$).
  • Using =~ but forgetting the pattern needs slashes: $str =~ "abc" does not do a regex match the way $str =~ /abc/ does.
  • Expecting $1 to still hold a value after a later, unrelated match has run and overwritten it.
  • Writing an overly greedy pattern like ".*" when a more specific character class would avoid matching too much text.
  • Forgetting anchors (^ and $) and accidentally matching a pattern anywhere in the string instead of the whole string.

Best Practices

  • Use the /x modifier for complex patterns to spread them across multiple lines with comments: /\d{4} # year\n -\d{2} # month/x.
  • Prefer named captures (%+) over numbered ones ($1, $2) once a pattern has more than two or three groups.
  • Anchor patterns with ^ and $ whenever you intend to validate an entire string, not just find a substring anywhere inside it.
  • Test regex patterns against edge cases (empty strings, unexpected formats) before trusting them in production code.
  • Choose delimiters other than / (like m{} or m||) when your pattern itself needs to contain slashes, to avoid excessive escaping.

Frequently Asked Questions

=~ is the binding operator that connects a variable to a pattern-matching operation; m// (or its slash shorthand //) is the actual match operator being applied. They are almost always used together.

Yes. Add the /i modifier, like /perl/i, to make a match case-insensitive.

The match expression simply evaluates to false, and any capture variables ($1, $2, ...) from a previous successful match remain unchanged, which can be a source of subtle bugs.

Use the /s modifier to let . match newlines too, and/or the /m modifier to make ^ and $ match at the start/end of each line rather than the whole string.

Key Takeaways

  • =~ binds a variable to a regex operation; !~ is its negated form.
  • m// (or plain //) tests whether a string matches a pattern.
  • Metacharacters like ., \d, \w, and \s match categories of characters; ^ and $ anchor the match.
  • Quantifiers (*, +, ?, {n,m}) control how many times something can repeat.
  • Parentheses create capture groups, accessible afterward as $1, $2, $3, or via named captures in %+.

Summary

Matching patterns is only half the story. Next, you will learn how to actually transform text using substitution (s///) and character translation (tr///), which is where regular expressions become truly powerful editing tools.

Next Lesson →

Pattern Matching & Substitution