Text mining James Joyce's Finnegans Wake. All measures are
dictionary heuristics, not philological classification.
1. Basic numbers
Words (tokens): 219,191. Distinct forms: 57,718.
Type-Token Ratio: 0.2633. Hapax legomena (words used
exactly once): 45,870, i.e.
79.47% of all forms. Mean word length:
4.71 letters.
Recognised as English: 85.06% of occurrences.
Outside all 14 dictionaries (pure coinages):
13.6% of occurrences.
2. How many words from which languages
Words that fail the English dictionary test and appear on the
frequency list of exactly one foreign language (exclusive match;
OpenSubtitles 50k lists plus hunspell for Irish and Latin):
Words outside all dictionaries that split into two dictionary parts:
13,773 (14,332 occurrences).
Part of speech guessed from the suffix: other 6904 (50%), noun 4335 (31%), verb 1801 (13%), adjective 527 (4%), adverb 206 (1%).
Most frequent language pairs inside words:
English + English: 9680 words (10093 occurrences) — willingdone = willing+done; earwicker = ear+wicker; lipoleums = lip+oleums; prankquean = prank+quean
English + French: 370 words (380 occurrences) — macdougal = mac+dougal; pritchards = prit+chards; bannistars = banni+stars; catholick = catho+lick
English + Norwegian: 357 words (376 occurrences) — kiddling = kidd+ling; fortissa = fort+issa; rossies = ros+sies; tolbris = tol+bris
English + Latin: 352 words (358 occurrences) — sandhyas = sand+hyas; pelagiarist = pelagia+rist; allinoilia = allino+ilia; gregorius = gregor+ius
German + English: 338 words (339 occurrences) — camiflag = cami+flag; bourgeoismeister = bourgeois+meister; negativisticists = negativistic+ists; fornicationists = fornication+ists
English + Dutch: 326 words (335 occurrences) — glendalough = glenda+lough; chickchilds = chick+childs; clottering = clot+tering; spanglers = spang+lers
English + Italian: 315 words (328 occurrences) — balenoarch = baleno+arch; allaboy = alla+boy; hairing = hai+ring; shaunti = sha+unti
English + Swedish: 283 words (296 occurrences) — fenians = fen+ians; finglas = fin+glas; sotisfiction = sotis+fiction; glasstone = glas+stone
English + Irish: 256 words (264 occurrences) — smolking = smol+king; taytotally = tayto+tally; mullingar = mullin+gar; yearlyng = year+lyng
English + Spanish: 195 words (201 occurrences) — breathings = brea+things; mamalujo = mama+lujo; beseated = bese+ated; paloola = palo+ola
English + Portuguese: 191 words (206 occurrences) — praties = pra+ties; margareena = marga+reena; remembore = remem+bore; remembored = remem+bored
English + Polish: 140 words (146 occurrences) — walhalloo = wal+halloo; semitary = semi+tary; synamite = syna+mite; plikplak = plik+plak
English + Finnish: 126 words (134 occurrences) — pepette = pep+ette; minnelisp = minne+lisp; peepette = peep+ette; garonne = garo+nne
French + Italian: 18 words (19 occurrences) — culious = culi+ous; salvatorious = salvatori+ous; arrivaliste = arriva+liste; capellisato = capelli+sato
French + Latin: 17 words (17 occurrences) — chronometrum = chrono+metrum; infirmierity = infirmier+ity; indivisibles = indi+visibles; deiectiones = deiectio+nes
Two foreign languages co-occurring in one sentence:
German + Latin: 15 sentences — e.g. „blau” and „dionysius” in one sentence
Irish + Latin: 14 sentences — e.g. „canavan” and „puritas” in one sentence
French + Latin: 10 sentences — e.g. „plaine” and „isthmon” in one sentence
German + French: 8 sentences — e.g. „isst” and „fuit” in one sentence
Irish + Italian: 7 sentences — e.g. „meagher” and „questa” in one sentence
Danish + Latin: 7 sentences — e.g. „nolans” and „cius” in one sentence
Latin + Dutch: 7 sentences — e.g. „epistola” and „nieuw” in one sentence
French + Dutch: 6 sentences — e.g. „trou” and „boobytrap” in one sentence
French + Irish: 6 sentences — e.g. „fion” and „seinn” in one sentence
Italian + Latin: 6 sentences — e.g. „questa” and „puella” in one sentence
4. Sound patterns
Vowel share: 39.68% in English words,
38.92% in coinages.
Bigrams overrepresented in the coinages relative to the book's own
English (how many times more frequent): yn (×9.9), yb (×9.31), ko (×7.91), lb (×6.95), ym (×6.8), ka (×6.26), rh (×5.67), eb (×5.36), eu (×4.39), oh (×4.12), iu (×3.99), yl (×3.92), ky (×3.67), lm (×3.64).
Most frequent consonant clusters ≥3 letters in coinages: str (181, e.g. oystrygods), ght (179, e.g. wielderfight), rth (158, e.g. eggeberth), ngs (140, e.g. ringsome), tch (138, e.g. detch), lls (129, e.g. wallhall's), nds (118, e.g. fondseed), ttl (113, e.g. lovelittle), ngl (109, e.g. sanglorians), nch (104, e.g. ffrinch), cks (103, e.g. carhacks), thr (101, e.g. threefoiled).
Alliteration: 1,791 runs of ≥3 words on the same
letter. Longest:
9 words on „h”: hicky hecky hock huges huges huges hughy hughy hughy
8 words on „t”: tilling teel tum telling toll teary turty taubling
8 words on „t”: the ten ton tonuant thunderous tenor toller the
8 words on „t”: the twenty twotoosent time thwealthy took thousands the
8 words on „t”: the tumble the toss tot the trouble the
7 words on „b”: bald black bronze brown brindled betteraved blanchemanged
7 words on „b”: boosted blasted bleating blatant bloaten blasphorus blesphorous
Adjacent repetitions: you you (×10), that that (×6), nin nin (×6), hee hee (×5), well well (×5), order order (×5), whoishe whoishe (×5), shee shee (×4), rain rain (×4), had had (×4).
The book split into 100 equal windows (~2,200 words each). Move the
slider to read the parameters at that point of the text. Vertical
dashes mark the boundaries of Books I–IV; the thick grey line is the
slider position.