Does the party out of power reach for evidence more often than the party in power? We built the instrument that can answer it, and then asked whether the instrument survives twenty-eight years of changing language.
Stop 1 of 6
Transcripts are prose. Before anything can be modelled, every speaking turn is attributed to a member, split into sentences, and joined to what we know about that member on that day.
Two corpora built by the same code: 1997–2016 and 2017–2024. The earliest Congresses are thin because the archive holds text for fewer of their hearings.
Stop 2 of 6
A sentence is empirical if it describes the state of the world, makes a causal claim, cites data or asks for evidence. Opinion, procedure and intent are not. These four split our two coders. Call them yourself.
One from each end of the period, coded under the same written protocol. Each is split in half: one half trains the models, the other is held back to grade them.
Stop 3 of 6
RoBERTa, DeBERTaV3 and ModernBERT, fine-tuned and compared against an 8B-parameter LLM prompted with the protocol's definition. The instrument is all of them: 3 encoders × 2 class weightings × 5 seeds, and a sentence's score is the mean of the 30 probabilities.
Five-fold cross-validation grouped by speaker, so no member's sentences are in both a training and a test fold, repeated over five seeds. AUC is on the pooled out-of-fold predictions of the mixed-era training set; each encoder is shown at its best class weighting. The LLM is Llama-3-8B-Instruct, 4-bit, scored from the logits of its two answer words. Averaging reduces variance and keeps the model's uncertainty, which matters because the analysis averages scores rather than counting labels.
Stop 4 of 6
A classifier trained on one period can get worse on another, and if its error changes where its training data end, it manufactures a trend. So we trained two 30-model ensembles that differ only in their labels: one on legacy labels alone, one on legacy and modern together.
Both ensembles score the same held-out sentences from each era. Adding modern labels improves ranking and calibration on modern text, and costs no ranking on older text.
Cohen's κ on the modern held-out set. The ensemble agrees with the gold standard about as well as one trained coder agrees with another.
The mixed-era ensemble's labels come from different coders on either side of 2016, so it could learn to score the era instead of the sentence. Then its mean score would jump between the 114th and 115th Congresses. It does not.
The jump at 2016 is ranked against the twelve steps between other adjacent Congresses (a placebo test). A real change in Congress would also move it, so the mixed-era ensemble is compared with the legacy-only one, which never saw modern labels, on the same speeches with a legislator-paired bootstrap. Their jumps differ, but by no more than at ordinary boundaries: the gap between them drifts gradually over the whole period rather than stepping where the training labels change.
Stop 5 of 6
Live
Held-out sentences the ensemble never trained on, with the published 30-model score beside our coders' call.
Live scoring is off on this copy of the page. The scores above are the published ensemble's.
Stop 6 of 6
Each speech scores the mean of its sentences. A regression with committee and Congress fixed effects then compares minority and majority members speaking in the same committee in the same Congress.