Chapter 23: The Language of Patterns
Learning Objectives
- Understand Regex syntax (character classes, quantifiers, groups).
- Apply Regex to real-world use cases like validation and parsing.
Prerequisites
Strings and String manipulation.
Why Does This Exist?
You can't solve complex text pattern matching (like validating an email) with simple `==` or `.contains()` checks without writing hundreds of lines of brittle code.
History
Created by Ken Thompson in the late 1960s based on formal language theory, originally for the QED text editor, eventually leading to `grep`.
Mental Model
Imagine a tiny robot reading one character of your text at a time, strictly following a set of matching rules you provided.
Internal Working
Regex engines often compile the pattern into a Finite State Machine. They process strings character by character, branching, backtracking, and matching based on greediness vs laziness.
Syntax
Python's `re` module rules:
- `[]`: Character classes
- `*, +, ?, {n,m}`: Quantifiers
- `()`: Capture groups
- `^, $`: Anchors
- `\d, \w, \s`: Special chars
1import re
2pattern = r'\d{3}-\d{4}'
3match = re.search(pattern, "Call 555-1234 now!")
4print(match.group()) # 555-1234Visual Explanation
Tiny Example
1import re
2email = "test@example.com"
3if re.match(r'^[\w\.-]+@[\w\.-]+\.\w+$', email):
4 print("Valid!")Walkthrough
- `^` matches start.
- `[\w\.-]+` matches one or more word chars, dots, or dashes.
- `@` matches literal @.
- `[\w\.-]+` matches the domain name.
- `\.` matches a literal dot.
- `\w+` matches the TLD (com).
- `$` matches the end.
Common Mistakes
Greedy Matching
`.*` matches as much as possible. If you want it to stop at the first match, use `.*?` (lazy matching). Forgetting to escape special chars like `.` is also common.
Debugging
Always use tools like https://regex101.com to visualize what your regex is matching and where it backtracks.
Mini Project
Time: 20 min.
Goal: Write regex patterns to (1) validate a UK postcode, (2) extract all URLs from a string, (3) replace phone numbers with [REDACTED].
💡 See One Approach (Mini Project)
One valid solution — yours may differ.
import re
# 1. Validate UK postcodes
uk = re.compile(r"^[A-Z]{1,2}[0-9][0-9A-Z]?\s?[0-9][A-Z]{2}$", re.IGNORECASE)
tests = ["SW1A 1AA", "EC1A 1BB", "bad-code", "12345", "M1 1AE"]
print("=== UK Postcode Validation ===")
for p in tests:
ok = bool(uk.match(p.strip()))
print(f" {p:15} {'OK' if ok else 'INVALID'}")
# 2. Extract URLs from text
text = "Visit https://example.com or read http://docs.python.org/3 for docs."
urls = re.findall(r"https?://[^\s]+", text)
print(f"\nURLs found ({len(urls)}): {urls}")
# 3. Redact phone numbers
sensitive = "Call 07700 900123 or +44 7911 123456 anytime."
redacted = re.sub(r"(\+44\s?|0)7\d{3}\s?\d{6}", "[REDACTED]", sensitive)
print(f"\nBefore: {sensitive}")
print(f"After: {redacted}")
Bigger Project
Time: 1.5 hr.
Goal: Build a log parser. Given an nginx access log, use regex to extract IP, timestamp, method, path, and status into a dictionary. Print stats.
💡 See One Approach (Bigger Project)
One valid solution — yours may differ.
import re
from collections import defaultdict
# Simulated nginx access log lines
raw_lines = [
'10.0.0.1 - - [01/Jan/2026:12:00:01 +0000] "GET /index.html HTTP/1.1" 200 1234',
'10.0.0.2 - - [01/Jan/2026:12:00:02 +0000] "POST /api/login HTTP/1.1" 401 89',
'10.0.0.1 - - [01/Jan/2026:12:00:03 +0000] "GET /dashboard HTTP/1.1" 200 5678',
'10.0.0.3 - - [01/Jan/2026:12:00:04 +0000] "GET /admin HTTP/1.1" 403 56',
'10.0.0.2 - - [01/Jan/2026:12:00:05 +0000] "POST /api/login HTTP/1.1" 200 312',
'10.0.0.1 - - [01/Jan/2026:12:00:06 +0000] "DELETE /api/user/7 HTTP/1.1" 204 0',
]
# Named-group regex — each (?P<name>...) creates a named capture group
PATTERN = re.compile(
r"(?P<ip>[\d.]+) - - \[(?P<ts>[^\]]+)\] "
r'"(?P<method>[A-Z]+) (?P<path>[^\s]+) HTTP/[\d.]+" '
r"(?P<status>\d+) (?P<bytes>\d+)"
)
entries = [m.groupdict() for line in raw_lines if (m := PATTERN.match(line))]
print(f"Parsed {len(entries)} log entries\n")
ip_counts = defaultdict(int)
status_counts = defaultdict(int)
total_bytes = 0
for e in entries:
ip_counts[e["ip"]] += 1
status_counts[e["status"]] += 1
total_bytes += int(e["bytes"])
print("Requests per IP:")
for ip, n in sorted(ip_counts.items(), key=lambda x: -x[1]):
print(f" {ip:15} {n} request(s)")
print("\nStatus codes:")
labels = {"200": "OK", "204": "No Content", "401": "Unauthorized", "403": "Forbidden"}
for code, n in sorted(status_counts.items()):
print(f" HTTP {code} {labels.get(code, '?'):15} x{n}")
print(f"\nTotal bytes transferred: {total_bytes:,}")
Production Usage
Log parsing (ELK stack), form validation, search engines, web scraping, and code linters.
Best Practices
- Use raw strings `r'...'` in Python.
- Compile with `re.compile()` for performance.
- Test edge cases exhaustively.
Interview Questions
🟢 Easy:What does `\d` mean?
🔍 Reveal Answer
🟡 Medium:What's the difference between `.*` and `.*?`?
🔍 Reveal Answer
🔴 Hard:What is a non-capturing group?
🔍 Reveal Answer
Revision Sheet
- Regex is a powerful DSL for text matching.
- Master character classes, quantifiers, and groups.
Connections
- Past Chapters: Strings.
- Future Chapters: Web Scraping and Data pipelines.