🚀 Mini Projects, lesson 2 of 6
Project: text statistics
Build an analyzer that cleans words, counts them and prints a report.
12 min
3 exercises
1 quiz
0/4 solved
Getting Python ready… examples can run in a moment.
Writers, students and search engines all need text statistics. You'll build an analyzer in three steps:
- Clean words: lowercase, punctuation stripped
- Stats: word count, unique words, average length, longest word, sentences
- Report: the most common words and a tidy summary
Why clean the words?
Without cleaning, "Hello," and "Hello..." count as
different words. strip removes the given characters
from both ends only.
What does this print?
Step 1: clean words
Write get_words(text): split the lowercased text,
strip punctuation from both ends of each piece with
string.punctuation, and skip pieces that end up
empty (like "--").
Step 2: statistics. Return them together in a
dictionary. Two details to handle: an empty text (no
dividing by zero, no max() of an empty list), and
sentences, which you can count as groups of .,
! or ? with re.findall(r"[.!?]+", text).
Counting sentence endings
The + makes "..." and "?!" count once each. Try
r"[.!?]" without the plus.
Step 2: the stats dictionary
Write stats(text) returning a dict with:
"words": number of words"unique": number of different words"avg_len": average word length, rounded to 1 decimal (0 if there are no words)"longest": the longest word (the first one on a tie;""if there are no words)"sentences": number of[.!?]+groups
Step 3: the report. Words like "is" and "and" are
everywhere, so skip them when finding the top words.
Counter.most_common keeps first-seen order on ties.
Step 3: top words and report
top_words(text, n)returns thenmost common(word, count)pairs, ignoring words inSTOP_WORDS.report(text)returns this multi-line string (top 3 words), shown here for "Cats chase mice. Mice fear cats!":
Words: 6
Unique: 4
Average length: 4.2
Longest: chase
Sentences: 2
Top words: cats (2), mice (2), chase (1)