🌙
☀️ Dark
The Engineer's Bible | Volume 1: Foundations | Reading Time: ~22 min

Chapter 23: The Language of Patterns

Learning Objectives

Prerequisites

Strings and String manipulation.

Why Does This Exist?

You can't solve complex text pattern matching (like validating an email) with simple `==` or `.contains()` checks without writing hundreds of lines of brittle code.

History

Created by Ken Thompson in the late 1960s based on formal language theory, originally for the QED text editor, eventually leading to `grep`.

Mental Model

Imagine a tiny robot reading one character of your text at a time, strictly following a set of matching rules you provided.

Internal Working

Regex engines often compile the pattern into a Finite State Machine. They process strings character by character, branching, backtracking, and matching based on greediness vs laziness.

Syntax

Python's `re` module rules:

python
1import re
2pattern = r'\d{3}-\d{4}'
3match = re.search(pattern, "Call 555-1234 now!")
4print(match.group())  # 555-1234

Visual Explanation

Regex Pattern: ^\d{3}-\d{4}$ ^ : Start of string \d{3} : Exactly 3 digits - : A literal dash \d{4} : Exactly 4 digits $ : End of string Matches: "123-4567"

Tiny Example

python
1import re
2email = "test@example.com"
3if re.match(r'^[\w\.-]+@[\w\.-]+\.\w+$', email):
4    print("Valid!")

Walkthrough

Common Mistakes

Greedy Matching

`.*` matches as much as possible. If you want it to stop at the first match, use `.*?` (lazy matching). Forgetting to escape special chars like `.` is also common.

Debugging

Always use tools like https://regex101.com to visualize what your regex is matching and where it backtracks.

Mini Project

Time: 20 min.

Goal: Write regex patterns to (1) validate a UK postcode, (2) extract all URLs from a string, (3) replace phone numbers with [REDACTED].

💡 See One Approach (Mini Project)

One valid solution — yours may differ.

python
import re

# 1. Validate UK postcodes
uk = re.compile(r"^[A-Z]{1,2}[0-9][0-9A-Z]?\s?[0-9][A-Z]{2}$", re.IGNORECASE)
tests = ["SW1A 1AA", "EC1A 1BB", "bad-code", "12345", "M1 1AE"]
print("=== UK Postcode Validation ===")
for p in tests:
    ok = bool(uk.match(p.strip()))
    print(f"  {p:15} {'OK' if ok else 'INVALID'}")

# 2. Extract URLs from text
text = "Visit https://example.com or read http://docs.python.org/3 for docs."
urls = re.findall(r"https?://[^\s]+", text)
print(f"\nURLs found ({len(urls)}): {urls}")

# 3. Redact phone numbers
sensitive = "Call 07700 900123 or +44 7911 123456 anytime."
redacted  = re.sub(r"(\+44\s?|0)7\d{3}\s?\d{6}", "[REDACTED]", sensitive)
print(f"\nBefore: {sensitive}")
print(f"After:  {redacted}")

Bigger Project

Time: 1.5 hr.

Goal: Build a log parser. Given an nginx access log, use regex to extract IP, timestamp, method, path, and status into a dictionary. Print stats.

💡 See One Approach (Bigger Project)

One valid solution — yours may differ.

python
import re
from collections import defaultdict

# Simulated nginx access log lines
raw_lines = [
    '10.0.0.1 - - [01/Jan/2026:12:00:01 +0000] "GET /index.html HTTP/1.1" 200 1234',
    '10.0.0.2 - - [01/Jan/2026:12:00:02 +0000] "POST /api/login HTTP/1.1" 401 89',
    '10.0.0.1 - - [01/Jan/2026:12:00:03 +0000] "GET /dashboard HTTP/1.1" 200 5678',
    '10.0.0.3 - - [01/Jan/2026:12:00:04 +0000] "GET /admin HTTP/1.1" 403 56',
    '10.0.0.2 - - [01/Jan/2026:12:00:05 +0000] "POST /api/login HTTP/1.1" 200 312',
    '10.0.0.1 - - [01/Jan/2026:12:00:06 +0000] "DELETE /api/user/7 HTTP/1.1" 204 0',
]

# Named-group regex — each (?P<name>...) creates a named capture group
PATTERN = re.compile(
    r"(?P<ip>[\d.]+) - - \[(?P<ts>[^\]]+)\] "
    r'"(?P<method>[A-Z]+) (?P<path>[^\s]+) HTTP/[\d.]+" '
    r"(?P<status>\d+) (?P<bytes>\d+)"
)

entries = [m.groupdict() for line in raw_lines if (m := PATTERN.match(line))]
print(f"Parsed {len(entries)} log entries\n")

ip_counts     = defaultdict(int)
status_counts = defaultdict(int)
total_bytes   = 0

for e in entries:
    ip_counts[e["ip"]] += 1
    status_counts[e["status"]] += 1
    total_bytes += int(e["bytes"])

print("Requests per IP:")
for ip, n in sorted(ip_counts.items(), key=lambda x: -x[1]):
    print(f"  {ip:15} {n} request(s)")

print("\nStatus codes:")
labels = {"200": "OK", "204": "No Content", "401": "Unauthorized", "403": "Forbidden"}
for code, n in sorted(status_counts.items()):
    print(f"  HTTP {code} {labels.get(code, '?'):15} x{n}")

print(f"\nTotal bytes transferred: {total_bytes:,}")

Production Usage

Log parsing (ELK stack), form validation, search engines, web scraping, and code linters.

Best Practices

Interview Questions

🟢 Easy:What does `\d` mean?

🔍 Reveal Answer
Matches any digit character (0-9).

🟡 Medium:What's the difference between `.*` and `.*?`?

🔍 Reveal Answer
`.*` is greedy (consumes as much as possible), `.*?` is lazy (consumes as little as possible to satisfy the match).

🔴 Hard:What is a non-capturing group?

🔍 Reveal Answer
`(?:...)` groups tokens for quantifiers or OR conditions without saving the match for backreferencing, saving memory.

Revision Sheet

✅ I can write regex patterns to validate, extract, and transform text without custom parsing code.

Connections

← Previous (Chapter 14) Chapter 14 Next (Chapter 16) → Chapter 16