Regular expressions & text tools

Standard Library Tour, lesson 6 of 6

🧰 Standard Library Tour, lesson 6 of 6

Regular expressions & text tools

Find, extract and clean text with re, string and textwrap.

12 min

2 exercises

1 quiz

0/3 solved

Getting Python ready… examples can run in a moment.

A regular expression (regex) is a mini-language for text patterns, like "three digits, a dash, four digits". Write patterns as raw strings (r"...") so backslashes stay intact.

  • \d digit, \w letter/digit/_, \s space, . any character
  • + one or more, * zero or more, ? optional, {3} exactly three
  • [abc] any one of these characters
  • ( ) captures a part of the match
Example

Find all matches

Change the first pattern to r"\d+". What does it find now?

Quiz

What does this print?

search and groups. re.search(pattern, text) returns the first match (or None). Parentheses capture parts you can pull out with .group(1) or .groups().

re.fullmatch only succeeds if the whole string matches, which is ideal for validating input.

Example

Extract parts, validate codes

Add flags=re.IGNORECASE to the fullmatch call and "ab-12" is still invalid. Why? (Count the digits.)

Example

Replace and split

re.sub replaces every match; re.split cuts the text wherever the pattern matches. Try re.sub(r"\d+", "#", note): what changes?

Exercise

Find the hashtags

Write find_hashtags(text) that returns every hashtag's word without the #, in lowercase, in the order they appear.

string and textwrap. The string module holds handy constants such as string.ascii_lowercase, string.digits and string.punctuation. textwrap wraps long text to a width, perfect for small screens.

Example

string constants and textwrap

Change width=28 to 20 and watch the text reflow.

Exercise

Normalize phone numbers

Write normalize_phone(text): remove every non-digit with re.sub. If exactly 10 digits remain, return them formatted as "555-123-4567"; otherwise return None.