Skip to contents

A function for retrieving all speeches in a debate transcript (publication type "referat"), one row per speech. Handles both the transcript format used up to the 2015-2016 session and the format used from the 2016-2017 session onward, returning the same variables for both. Plenary sittings, open hearings, and meetings in the European Committee are all published as transcripts.

Usage

get_speeches(publicationid = NA, good_manners = 0, link = TRUE)

Arguments

publicationid

Character string, or a vector of strings, indicating the id of the transcript to retrieve. Ids can be found with get_session_publications (type = "referat")

good_manners

Numeric. Seconds delay between calls when making multiple calls to the same function. Note that the Stortinget API is limited to 100 calls per minute (see https://data.stortinget.no/nyhetsoversikt/begrensning-pa-api-kall/).

Logical. Whether to add person ids linked from the names of speakers and chairs (see speaker_links). Defaults to TRUE.

Value

A data.frame with the following variables:

response_dateDate and time of retrieval (the transcripts have no response date of their own)
publication_idId of the transcript
session_idId of the parliamentary session (see get_parlsessions), from meeting_date
meeting_orderOrder of the meeting within the transcript (some transcripts hold several meetings)
meeting_idMeeting id (see get_session_meetings), when given in the transcript
meeting_titleRaw meeting heading (e.g. "Møte torsdag den 13. februar 2014 kl. 10")
meeting_dateDate of the meeting, from meeting_title and the publication id (see details)
sectionName of the XML element directly containing the speech (e.g. sak, spm, formalia)
case_idId of the case the speech belongs to (see get_case), when given in the transcript
agenda_noAgenda item number (from 2016-2017 onward)
agenda_mergedAgenda item numbers debated together, comma separated (from 2016-2017 onward)
speech_orderOrder of the speech element within the transcript
speech_partOrder of the speaker within the speech element (usually 1)
speech_idSpeech element id (from 2016-2017 onward)
speech_typeType of speech ("hovedinnlegg", "replikk", or "presinnlegg"; see details)
speaker_rawSpeaker as written in the transcript
speaker_titleTitle parsed from speaker_raw (e.g. "Statsråd", "Presidenten")
speaker_nameName parsed from speaker_raw
speaker_partyParty parsed from speaker_raw, as a party id of get_all_parties (else NA)
speech_timeTime stamp parsed from speaker_raw (hh:mm:ss)
person_idId of the speaker (see get_mp), when given in the transcript
linked_person_idId of the speaker, linked from speaker_name (with link = TRUE)
link_methodHow linked_person_id was linked (with link = TRUE; see speaker_links)
chair_nameName of the sitting chair (president or meeting leader)
chair_idId of the sitting chair, when given in the transcript
chair_linked_idId of the sitting chair, linked from chair_name (with link = TRUE)
textSpeech text, one line per paragraph

Details

Some speech elements in the transcripts hold more than one speaker (e.g. a question and an answer in a hearing). These are split into one row per speaker; speech_order identifies the speech element and speech_part the speaker within it.

The transcripts before about 2005 mostly do not say whether a speech is a main speech ("hovedinnlegg") or a reply ("replikk"); speech_type is then NA, except for the president's remarks ("presinnlegg").

The meeting date is given both in the meeting heading (weekday, day, month, and, from 2007 onward, year) and in the publication id, and both contain occasional errors. When they agree, that date is used; when they disagree, the one whose weekday matches the weekday in the heading is used, and NA if that does not settle it.

Speakers are only identified by the API (person_id) in transcripts from the 2016-2017 session onward, and not for all speeches (e.g. not for the president or committee chair). Note that these ids are not always correct: speeches are sometimes tagged with the id of another person than the one named in speaker_raw. A numeric suffix that some of these ids carry (e.g. "ARK_775612110") is removed. The raw speaker string is always kept in speaker_raw; speaker_title, speaker_name, speaker_party, and speech_time are parsed from it.

With link = TRUE (the default), person ids linked from the names of the speakers and chairs are added from the speaker_links dataset (linked_person_id, link_method, and chair_linked_id), also for transcripts before 2016-2017. These links are made by the package, not given by the API (person_id is kept as the API gives it), and they only cover the sessions in speaker_links; for later sessions, the linked ids are NA.

The sitting chair (president or meeting leader) is tracked through the transcript: the chair named at the start of each meeting, updated at the transcript's notes on changes of chair (e.g. "X hadde her overtatt presidentplassen"). When a note about the chair cannot be read, the chair is set to NA until the next recognized change. chair_id is only given for the chair named at the start of a meeting, and only when the transcript includes it (a numeric suffix that some of these ids carry, e.g. "OLET_62710109", is removed).

Examples


if (FALSE) { # \dontrun{
speeches <- get_speeches("s140213")
head(speeches[, c("speech_type", "speaker_name", "person_id", "linked_person_id")])

# Without the linked ids
speeches <- get_speeches("s140213", link = FALSE)
} # }