🧰 Standard Library Tour, lesson 6 of 6
Regular expressions & text tools
Find, extract and clean text with re, string and textwrap.
12 min
2 exercises
1 quiz
0/3 solved
Getting Python ready… examples can run in a moment.
A regular expression (regex) is a mini-language
for text patterns, like "three digits, a dash, four
digits". Write patterns as raw strings (r"...")
so backslashes stay intact.
\ddigit,\wletter/digit/_,\sspace,.any character+one or more,*zero or more,?optional,{3}exactly three[abc]any one of these characters( )captures a part of the match
Find all matches
Change the first pattern to r"\d+". What does it
find now?
What does this print?
search and groups. re.search(pattern, text)
returns the first match (or None). Parentheses
capture parts you can pull out with .group(1) or
.groups().
re.fullmatch only succeeds if the whole string
matches, which is ideal for validating input.
Extract parts, validate codes
Add flags=re.IGNORECASE to the fullmatch call
and "ab-12" is still invalid. Why? (Count the digits.)
Replace and split
re.sub replaces every match; re.split cuts the
text wherever the pattern matches. Try
re.sub(r"\d+", "#", note): what changes?
Find the hashtags
Write find_hashtags(text) that returns every
hashtag's word without the #, in lowercase, in
the order they appear.
string and textwrap. The string module holds
handy constants such as string.ascii_lowercase,
string.digits and string.punctuation.
textwrap wraps long text to a width, perfect for
small screens.
string constants and textwrap
Change width=28 to 20 and watch the text reflow.
Normalize phone numbers
Write normalize_phone(text): remove every non-digit
with re.sub. If exactly 10 digits remain, return
them formatted as "555-123-4567"; otherwise return
None.