ToK-explorer
  • Home
  • Relative word frequencies
  • Baselines
  • Relative word frequencies (words)
  • Word baselines
  • About
  1. Overview
  2. About
  • Overview
    • About
    • Speakers
    • Relative word frequencies
    • Baselines
    • Relative word frequencies (words)
    • Word baselines
    • damer & fru
    • Paper figures (all plot_ functions)
  • Chamber 1
    • Speakers
    • Baselines
    • Relative word frequencies
    • Word baselines
    • Relative word frequencies (words)
    • damer & fru
  • Chamber 2
    • Speakers
    • Baselines
    • Relative word frequencies
    • Word baselines
    • Relative word frequencies (words)
    • damer & fru

On this page

  • Terminology
  • Corpus figures
    • Before 1920
    • Full period (1900–1940)
    • 1920 onwards
  • Data provenance
  1. Overview
  2. About

ToK-explorer

Authors

Mathias Johansson

Ulrika Holgersson

This site accompanies the Tal om Kvinnor project at Lund University. It reports how the Swedish Riksdag spoke about women between 1900 and 1940, using the openly published Swedish Parliament Corpus as its raw material.

Each page presents one slice of that story. Baselines report raw utterance and word counts; Relative word frequencies report each category’s share of the total. Every headline view is available for the whole period and for either chamber separately.

Terminology

Three keyword categories are used throughout the site:

  • kvinna 1 — a single broad wildcard kvinn* that matches any word starting with kvinn.
  • Kvinna 2 — a hand-picked list of generic feminine terms, pronouns and kin relations (dam, dotter*, flick*, mor, …).
  • Kvinna 3 — a large list of specific and often work-related female titles (barnmorska, hushållerska, piga*, …).

The categories are guaranteed pairwise disjoint by an automated check; where a name appears in more than one list a Kvinna 2 entry always wins over Kvinna 3. Older pages sometimes show a kvinna all line — that is the union of all three, i.e. any utterance/word matching at least one category. It sits above the individual category lines by construction.

Corpus figures

The three heatmaps below split the corpus by chamber and by the speaker’s recorded gender, and count both utterances and words. The first heatmap covers everything before 1920; the second covers the whole 1900–1940 window; the third covers 1920 onwards. The two blocks of years bracket the introduction of women’s suffrage in 1919, which lets us see how the gendered composition of the recorded discourse shifts once women can vote.

Before 1920

57,417 utterances · 38,303,709 words

Chamber 1 · man Chamber 1 · woman Chamber 1 · unknown Chamber 2 · man Chamber 2 · woman Chamber 2 · unknown
metric
utterances 20567 0 1264 33602 0 1984
words 13086863 0 632490 23429868 0 1154488

Full period (1900–1940)

141,379 utterances · 92,595,326 words

Chamber 1 · man Chamber 1 · woman Chamber 1 · unknown Chamber 2 · man Chamber 2 · woman Chamber 2 · unknown
metric
utterances 57448 123 2141 78034 558 3075
words 36588819 94426 1123647 52670647 327998 1789789

1920 onwards

83,962 utterances · 54,291,617 words

Chamber 1 · man Chamber 1 · woman Chamber 1 · unknown Chamber 2 · man Chamber 2 · woman Chamber 2 · unknown
metric
utterances 36881 123 877 44432 558 1091
words 23501956 94426 491157 29240779 327998 635301

Data provenance

The two upstream sources are pinned by SHA-256 in tok_preparer/src/download.py. The exact commits used to render this build are reported below.

Riksdagen speeches:  records_speeches v1.6.0
                     sha256: a1c1310971214d4f836b72f48503b8d86bd3ec80763738e884a02e45a8f49ebf
Riksdagen persons:   persons.sqlite   v1.2.2
                     sha256: dfc0db29039715440b53d9820873c8c163fc72b5acdd207917e58bd4869ab5fd

ToK-explorer  parent HEAD: 346383dbe094 — Expand figure gallery based on new corpus
tok_preparer  submodule HEAD: 3e697c4b88f4 — remove pandas FutureWarning