What you will build
Word rule: a word is a continuous run of Unicode letters or digits.
gpu-0 becomes gpu and 0; café remains one word.Five guided demos
Complete these in order. Each demo gives exact input and expected output.
Demo 1 · Count one sentence
Practise split(), loops, dictionaries, and dict.get().
Input
the cat and the hat and the bat
Expected output
{'the': 3, 'cat': 1, 'and': 2, 'hat': 1, 'bat': 1}Demo 2 · Parse real log text
Handle punctuation, case, digits, Unicode, and multiple lines.
Input
ERROR: disk full on gpu-0 error: Disk full on gpu-1 Warning: café wifi down, retrying...
Expected output with fold=True
{'error': 2, 'disk': 2, 'full': 2, 'on': 2,
'gpu': 2, '0': 1, '1': 1, 'warning': 1,
'café': 1, 'wifi': 1, 'down': 1, 'retrying': 1}Demo 3 · Select the top N
Sort by count descending, then word ascending.
Input
counts = {'b': 2, 'a': 2, 'c': 5, 'd': 1}
n = 3Expected output
5 c 2 a 2 b
Demo 4 · Make reusable functions
Create count_words(), top_n(), and a frozen dataclass.
Input
io.StringIO("the cat and the hat")
top_n(..., 3)Expected output
2 the 1 and 1 cat
Demo 5 · Build the CLI
Add argparse, stdin, merged files, logging, and exit codes.
Command
echo "x y x z z z" | python demo-v5.py -n 2
Expected output
3 z 2 x
Final assignment downloads
Keep these files together. Implement the three functions in wordfreq.py, then run the tests.
Run and verify
python -m unittest -v python wordfreq.py -n 5 app1.log echo "a b a" | python wordfreq.py
Done when: all tests pass, count and word are separated by a tab, ties are deterministic, and files are streamed line by line.