articles has columns article_id and text.
Return the top_n most frequent words across all the text, as a DataFrame with
columns word and count. Lowercase everything first and strip punctuation, so
"Data," and "data" are the same word. Sort by count descending then word
alphabetically, and renumber the index from 0.
Input
articles =
article_id text
0 1 Data, data everywhere!
1 2 The data is fine.
top_n = 3
Output
word count
0 data 3
1 everywhere 1
2 fine 1