# Supreme Court votes from each justice's own record: rules Written 2026-10-04, before any audio was downloaded or measured. Committed to git before the data exists on this machine; the commit time is the freeze. Nothing below changes after measurements are seen. If a rule turns out to be unworkable, the change is a new dated section in this file, made before the held-out terms are run, saying what changed and why. ## Question After an oral argument and before the decision, does how a justice spoke to each side, measured against that justice's own record of past arguments, predict whether they vote for the petitioner? Does it predict better than the published pitch method, than question counts, and than knowing the justice? Prior knowledge, declared: Dietrich, Enos and Sen (Political Analysis, 2019) found that justices speak at a higher pitch to the side they vote against. The literature on question counts finds that the side asked more questions tends to lose. Neither finding is used to choose anything below except baselines B2 and B3. ## Data - Arguments: Oyez API (`api.oyez.org`), terms OT2015 to OT2023. Audio (MP3) and the speaker-labelled transcript with turn start and stop times, as published. - Votes: Supreme Court Database, release 2026_01, justice-centered, docket (`SCDB_2026_01_justiceCentered_Docket.csv`). - Vote for the petitioner = (`majority` = 2 and `partyWinning` = 1) or (`majority` = 1 and `partyWinning` = 0). Rows with `partyWinning` = 2 or missing, `majority` missing, or a justice not participating are excluded. Only `decisionType` 1, 6 and 7 are used. - Joined on docket number and term. ## Exclusions - Arguments held by telephone (2020-05-04 to 2021-06-30, all of OT2020 and the May 2020 sittings): different audio and an imposed order of questioning. - Original-jurisdiction cases. - A docket argued twice: the later argument is used. - An argument without audio or without a timed transcript. ## Sides Each transcript section is assigned to a side from its advocate's Oyez description. "petitioner" or "appellant", or an amicus "supporting" either, is the petitioner's side (P). "respondent" or "appellee", or an amicus supporting either, is the respondent's side (R). A rebuttal belongs to its advocate's side. A section whose side cannot be read this way is excluded. A justice's turn belongs to the side of the section it is in. ## Measurements For each justice, argument and side, from that justice's own turns only: | Channel | Definition | |---|---| | C1 pitch | median F0 in semitones re 100 Hz over voiced frames (Praat via parselmouth, floor 75 Hz, ceiling 500 Hz, 10 ms step), keeping frames within one octave of the justice's median F0 in that argument | | C2 loudness | median RMS level in dB over the same voiced frames | | C3 speaking rate | words in the justice's turns / summed duration of those turns, per minute | | C4 words | words the justice spoke to the side / duration of the side's sections, per minute | | C5 turns | number of the justice's turns / duration of the side's sections, per minute | | C6 cut-ins | turns that start right after an advocate turn whose text ends in "--", per minute of the side's sections | | C7 response gap | median seconds from the end of the preceding advocate turn to the start of the justice's turn; gaps below 0 count as 0, gaps above 5 s are left out | A side is read for a justice only with at least 2 turns and 20 words to that side; otherwise that justice has no reading for the argument (counted, not imputed in the baselines below; in models P and B4 every input is set to 0, the justice's usual, and a 0/1 "no reading" input is added). The side difference for each channel is D = value(P) - value(R). ## The record A justice's record is the set of their D values over their reference arguments: all of their arguments in the training terms, plus, for a held-out argument, their held-out arguments argued before it. Outcomes are never used to build a record. Each D is placed in its own record as a normal score: Blom, z = Phi^-1((r - 3/8)/(n + 1/4)), r the mid-rank among the n reference values. A justice with fewer than 20 reference arguments is not evaluated for that argument (counted; this covers Justice Jackson's first 20 arguments). ## Models All fitted on the training terms only. Coefficients are then frozen. - B0: always the petitioner. - B1: the justice's share of votes for the petitioner in the training terms (pooled share for a justice with none), predict petitioner if at least 0.5. - B2, pitch (published): vote against the side spoken to at the higher pitch, raw C1 difference; no reading or a tie falls back to B1. - B3, questions: vote against the side that got more of the justice's turns, raw C5 difference; no reading or a tie falls back to B1. - B4, raw channels: logistic regression on the raw D of C1-C7, each standardized over all training rows, plus one intercept per justice, plus the no-reading input; L2, C = 1.0. - P, Phosphoros: the same logistic regression on the record normal scores z of C1-C7 instead of the raw D. A model predicts the petitioner when its probability is at least 0.5. ## Splits - Training: OT2015-OT2021 (OT2020 excluded by the telephone rule). - Held out: OT2022 and OT2023. Run once, after the code and the fitted coefficients are committed. - Live: OT2026, after the held-out result, under a separate dated section. ## Success P succeeds only if all three hold on the held-out votes: 1. Vote accuracy of P exceeds the best of B0-B3 by at least 2.0 percentage points, and the 95% interval of that difference from 2,000 bootstrap resamples of cases (seed 20261005) lies above 0. 2. P beats that same baseline in OT2022 and in OT2023 separately. 3. P's log loss is lower than B1's. Reported whatever happens, not part of the criterion: P against B4 (does the justice's own record add anything to the raw differences?), accuracy per justice, case outcomes from the majority of predicted votes, calibration, and every count of exclusions. If P fails, the result is published as a failure and the method is not tuned on the held-out terms. ## Amendment 1, 2026-10-04: response gap (C7) from the audio Found while testing the extraction code on one training argument, before any channel was computed over the data: Oyez cuts turns back to back (4,713 of 4,713 turn boundaries in 20 arguments have a gap of exactly 0 s), so C7 as written is always 0. C7 is now read from the audio: at the boundary b where a justice's turn follows an advocate's turn, the gap is the time from the last voiced frame at or before b to the first voiced frame after b, each searched within 5 s of b. If b falls inside voiced speech the gap is 0 (one voice runs into the other). Voiced means F0 > 0 in the same pitch track as C1, any speaker. Gaps above 5 s are left out, as before. Nothing else changes. ## Amendment 2, 2026-10-04: how the fit is frozen Nothing more for this test goes into git. The training fit is frozen by research/scotus/FROZEN, which fit writes with a UTC time and the SHA-256 of fit.json and train_record.json; the held-out run refuses to start if either file no longer matches. The rules and code above were committed before any data was measured (c25d8eb, 3162ebd, 15494a4). ## Result, 2026-10-04: held-out run (once), FAILED the criterion OT2022-OT2023: 107 cases, 831 justice votes (157 rows without a usable SCDB vote, 2 original-jurisdiction cases excluded). Vote accuracy: P 68.95%, B4 (raw channels) 69.80%, B0 always petitioner 66.31%, B1 66.31%, B2 pitch 60.77%, B3 questions 60.53%. Best baseline B0. Gain +2.65 points, 95% case-bootstrap interval -0.72 to +5.93: criterion 1 fails. Criterion 2 passes (OT2022 64.7 vs 62.0; OT2023 73.0 vs 70.4). Criterion 3 passes (log loss 0.607 vs 0.643). Full output: research/scotus/heldout_results.json. Not tuned afterwards. ## Phase 2, 2026-10-05: improving the method, then a new held-out test OT2022-OT2023 have been seen and are not used to choose anything below. The method is improved on the training terms only and then tested once on OT2024 and OT2025, which nobody has measured. Candidates, fixed now: - V0: model P as frozen in phase 1 (the reference). - V1: V0 without procedural turns: a justice's turn of fewer than 5 words, and the presiding justice's turns that open or close an argument or a section ("we'll hear argument", "thank you, counsel", "the case is submitted", "rebuttal", "minutes remaining"). - V2: V1 with partial readings: C4-C6 are read whenever the justice spoke in the argument (zero turns to one side is a value); C1-C3 and C7 need 2 turns and 20 words to that side, otherwise that channel is 0 with a 0/1 missing input. - V3: V2 plus the bench: for each of C4-C6, the mean of the other justices' D in the same argument. - V4: V3 plus format: each channel interacted with a 0/1 input for OT2021 and later (questioning in turn at the end of each advocate's time). - Each of V0-V4 in two forms: record normal scores (as P) and raw standardized differences (as B4). Selection: leave-one-term-out over the training terms (each of OT2015-2019 and OT2021 held out in turn; fit on the rest). The candidate with the lowest pooled out-of-term log loss is chosen; within 0.002 of it, the earlier and simpler candidate (lower V, raw before record) is chosen. Its fit on all training terms is then frozen (FROZEN), and only then are OT2024-OT2025 downloaded and measured. Test: the chosen model on OT2024-OT2025, once, with the phase 1 criterion unchanged (best of B0-B3 by at least 2.0 points with the 95% case-bootstrap interval above 0; better in each term; log loss below B1). OT2022-OT2023 are reported for the chosen model as already-seen data, not as a test. ## Result, phase 2, 2026-10-05 Selection (training terms, leave-one-term-out log loss): V0 raw 0.6323 / record 0.6308; V1 0.6375 / 0.6362; V2 0.6157 / 0.6169; V3 0.6125 / 0.6132; V4 0.6123 / 0.6135. Chosen by the rule: V3-raw (V4-raw within 0.002). Frozen in FROZEN2. Test, once, OT2024-OT2025 (104 cases, 783 votes): V3-raw 68.07%; best baseline B2 pitch 63.47%; B0 61.17%; B3 59.77%. Gain +4.60 points, 95% interval -0.26 to +9.30: criterion 1 fails. Criterion 2 passes (OT2024 68.3 vs 63.6; OT2025 67.8 vs 63.3). Criterion 3 passes (log loss 0.604 vs 0.664). Not a success under the rules. Already seen, not a test, OT2022-OT2023 (828 votes): V3-raw 71.4% against B0 66.4%, +4.95 points, interval +0.72 to +9.08. ## Test 1 (phase 3), 2026-10-05: does delivery add information beyond the case file? Written before any case-facts model was built or run. Question: on votes it has not seen, does a model with the case file and the frozen V3 delivery features predict better than the same kind of model with the case file alone? Case file (all known before argument): justice; SCDB issueArea, lcDispositionDirection, lcDisagreement, certReason, jurisdiction, caseSource and the petitioner and respondent types (codes seen fewer than 20 times in training grouped as "other"); justice x lcDispositionDirection (ideology against the direction of the ruling below); and the United States' side as amicus or party (petitioner / respondent / none), read from the Oyez advocate descriptions. Model family: logistic regression (L2, C = 1) or gradient-boosted trees (sklearn HistGradientBoostingClassifier, defaults, random_state 0), whichever has the lower leave-one-term-out log loss on the training terms for the case file alone. The combined model is the same family with the V3 delivery features added. Both are then fitted on all training terms and frozen (FROZEN3). Primary test: OT2024-OT2025 (delivery frozen before these terms were measured). Statistic: per case, the case-file model's summed log loss minus the combined model's; 2,000 bootstrap resamples of cases (seed 20261005). Delivery adds information if the one-sided 95% lower bound (5th percentile) is above 0. Reported alongside: both models' vote and case-outcome accuracy, and the same statistic on OT2022-OT2023 as already-seen data. ### Test 1, amendment, 2026-10-05 (before any test term was scored) The case-file model as specified overfits: its leave-one-term-out log loss on the training terms was 0.7318 (logistic) and 0.7748 (boosted), worse than a coin flip, so a comparison with it would be against a strawman. Amended: each of the two models (case file alone; case file + delivery) gets its own settings, chosen by the same leave-one-term-out log loss on the training terms, from the same grid: logistic regression with C in {0.003, 0.01, 0.03, 0.1, 0.3, 1}, and gradient-boosted trees with max_depth 3, learning_rate 0.05, max_iter 200, l2_regularization 1.0 (random_state 0). Everything else is unchanged. ## Test 2, 2026-10-05: the experts' read of the same argument Written before any recap was fetched or read. Expert read: SCOTUSblog's argument recap for the case, the first post linked from the case's SCOTUSblog page that was published on the argument day or up to 3 days after it and is not a preview, a morning roundup, or petitions coverage. Fetched politely (2 s between requests; pages robots.txt allows). Coding: Qwen3-8B (local, temperature 0) reads only the recap and the two party names and answers which side the justices appeared more likely to favour: petitioner, respondent, or unclear. I read 20 recaps myself and report how often I agree with the model's code. Questions, on OT2024-OT2025 (primary) and OT2022-OT2023 (also reported): 1. Head to head on case outcomes, cases with a clear expert read: the expert read against Phosphoros (case file + delivery, frozen in test 1, majority of predicted votes). Accuracy of each and the paired bootstrap interval of the difference (two-sided 95%). 2. Does delivery add to the experts? Logistic model, settings chosen as in test 1 on the training terms, with the case file + expert read (petitioner / respondent / unclear) against the case file + expert read + delivery. One-sided 95% lower bound of the log-loss gain per vote, case bootstrap, as in test 1. Expert reads for the training terms are coded the same way. ## Test 3, 2026-10-05: prediction markets Written before any market price was fetched. Cases: every argued case in OT2022-OT2025 with a Polymarket or Kalshi market on its outcome. For each, the market question is mapped by hand to "the petitioner wins" (the mapping is written down with the case before prices are read). Market probability: the last traded price at the end of the day after the argument (UTC), from the public price history. Phosphoros' probability that the petitioner wins: from the frozen test 1 model (case file + delivery), the chance that at least a majority of the participating justices vote for the petitioner, treating its vote probabilities as independent. Reported per case: market, Phosphoros, outcome; and for both, the Brier score and the number of cases called right. With this few cases the comparison is descriptive: no significance claim is made from it. Test 3 mapping: research/scotus/market_map.json (written before prices). ### Test 2, amendment, 2026-10-05 (before any recap was coded) The fetch at 2 s per request would have taken about 6 hours. Pause lowered to 1 s; posts whose address marks them as opinion analyses, symposiums, amicus round-ups, cert grants or announcements are skipped before being opened. The choice of recap (earliest qualifying post within 3 days) is unchanged. ## Test 4, 2026-10-05: the live term (OT2026) Written before any OT2026 argument was measured. OT2026 opens 2026-10-05. Model: V3-raw as frozen in phase 2 (FROZEN2), unchanged. (The test 1 case-file model needs SCDB codes that are only published after decisions, so it cannot forecast live.) Procedure: when Oyez publishes an OT2026 argument with a timed transcript, it is measured and each justice's vote probability is written to research/scotus/live_predictions.jsonl before the decision, one line per case, each line carrying the SHA-256 of the previous line (a chain, so no line can be changed or back-dated without breaking every later hash). The latest hash may be posted publicly as a timestamp. Statistic (the yardstick from the power analysis): per vote, the log loss of the justice-identity model (each justice's training share of votes for the petitioner) minus V3's; case bootstrap, 2,000 resamples, seed 20261005; delivery adds information if the one-sided 95% lower bound is above 0. Also reported: vote and case accuracy against always-petitioner. The test closes when 80 OT2026 cases are decided (power analysis on the training terms: about 73 cases for 80% power); interim results may be shown but are labelled interim. ### Test 2, amendment 2, 2026-10-05 (before any recap was coded) "A view from the courtroom" posts are columns about the scene in the courtroom, not analyses of how the argument went, and they are often published before the argument analysis on the same day. They are skipped like previews, so the recap is the earliest remaining post within 3 days; cases already fetched with one are fetched again. ### Test 2, amendment 3, 2026-10-05 (before any recap was coded) For many OT2023-OT2024 cases the SCOTUSblog case page does not link the argument analysis at all (e.g. Royal Canin v. Wullschleger lists four 2024 posts and nothing after argument). For a case whose page yields no qualifying post, the posts of the argument's month (and of the next month when the 3-day window crosses into it) are read from the site's sitemap, and the earliest post published within the window, not skipped by the same rules, whose title or text names the case (a distinctive word of either party's name, or the docket number) is taken. ### Test 2, amendment 4, 2026-10-05 (before any recap was coded) The month scan as written matched posts that mention a case in passing ("Argument transcripts" listings, unrelated news). Changed: among posts in the window, the one whose title and text mention the case's distinctive party words most often is taken, with at least 3 mentions; "Argument transcripts" listings are skipped. Every month-scan match is listed by title for inspection before coding, and a match whose title is plainly about another matter is dropped. ### Test 2, amendment 5, 2026-10-05 (before any analysis) The extracted recap text began with the site's navigation and page metadata (about 1,000 characters), which the coder read first. That prefix is stripped; all recaps are coded again from the article text. Codes from the first, interrupted pass are discarded. ### Test 2, amendment 6, 2026-10-05 (before any analysis) Coder check on 20 recaps (seed 20261005), read by me before seeing the model's codes: exact agreement 7/20; where both took a side, the model chose the opposite side in 3 of 9 (Murthy v. Missouri, Relentless, Hencely v. Fluor, all clear in the recap), and over all recaps it coded "respondent" twice as often as "petitioner". That points to the coder mixing up which named party is the petitioner, which would weaken the experts' side of the comparison. Changed: the coder reasons before answering (Qwen3 thinking on, up to 1,024 tokens) and is told to first work out which name in the report refers to the petitioner (the report may say "the government", "the challengers", a lawyer's name). The first 20 recaps were used to find the fault, so agreement is checked again on a fresh 20 drawn with seed 20261006, read by me before seeing the new codes. Bar, fixed now: if, where both take a side, the coder and I disagree in more than 20% of recaps, test 2 is reported as inconclusive and no comparison with the experts is claimed. Coder check, fresh 20 (seed 20261006), read by me before seeing the codes: exact 11/20; where both took a side (7), opposite in 1 (14%, GEO v. Menocal): under the 20% bar, so test 2 goes ahead. The coder is cautious: I took a side on 13 of the 20, it on 7 (it called e.g. "Affirmative action appears in jeopardy" unclear). Its sides are mostly right but it under-reports the experts' leans, which makes the expert side weaker than the recaps are; question 2 is read with that in mind. ## Result, test 2, 2026-10-05 Codes (reasoning coder): petitioner 136, respondent 69, unclear 346, no recap 61. OT2024-OT2025 (104 cases, 783 votes): Q1, cases with a clear expert read (30): experts 80.0%, Phosphoros 73.3%, difference interval -26.7 to +13.3 points (no clear difference; experts ahead). Q2: case file + expert read + delivery against case file + expert read: log-loss gain 0.038 per vote, one-sided 95% lower bound +0.018: delivery adds information to the experts' read; votes 71.1% vs 68.3%. Caveat: the coder leaves many clear recaps "unclear", so the expert input is weaker than the recaps themselves. Already seen, OT2022-OT2023: Q1 experts 75.9% vs Phosphoros 72.4% (29 cases); Q2 gain 0.050, lower bound +0.030. ## Test 2b, 2026-10-05: the experts' read coded by Claude Written before any recap was coded by Claude. Ihor judged the local 8B coder too weak a reader. Claude (Opus 5.5, this session) reads every recap that test 2 found (title, opening and closing of the text) and codes the lean the recap reports: petitioner, respondent or unclear, from the recap's words only. Known contamination: Claude knows the outcome of many of these cases, which can only make the experts look better than they were; a result in Phosphoros' favour is therefore conservative, an expert advantage may be inflated. Codes are written to media/scotus/experts/codes_claude.json, batch by batch, before any analysis. Analysis: identical to test 2 (Q1 head to head on clear reads; Q2 delivery added to case file + expert read, settings chosen on the training terms), frozen in FROZEN5. The prediction-market test (test 3) is dropped from the claims at Ihor's call: markets mix insider information and automated trading, and cover a handful of headline cases. ### Test 2b, amendment, 2026-10-05 (before any recap was coded for 2b) At Ihor's request the coder is six Claude Sonnet 5.5 subagents instead of Opus reading all 551 itself, each given the same written instructions and a sixth of the recaps, writing codes_sonnet_.json. Check: Opus' own reads of 65 recaps (the 40 spot-check recaps and the first 25 in reading order), written to check_codes_opus.json before any Sonnet code existed. Bar as before: where both take a side, opposite in at most 20%. The same outcome-knowledge caveat applies. ### Test 2b, correction, 2026-10-05 (before any 2b analysis) A coder flagged that Oyez lists the parties in reverse for some cases (Oyez's "first_party" labelled Petitioner is the respondent). The caption decides: the petitioner is the party named first in the case name. Where the caption's first party matches Oyez's second_party and not its first_party, the expert code is flipped (petitioner <-> respondent). Found by that rule: LabCorp v. Davis, Coinbase v. Suski, Garland v. Cargill. The same three affect test 2's codes; test 2 is superseded by 2b and is not re-run. ## Result, test 2b, 2026-10-05 Sonnet coders checked against Opus' 65 reads: 56 exact, 0 opposite of 33 where both took a side; overall codes petitioner 198, respondent 109, unclear 244. OT2024-OT2025 (104 cases, 783 votes): Q1, 46 cases with a clear expert read: experts 91.3%, Phosphoros 73.9%, difference interval -30.4 to -4.3 points: the experts are clearly better head to head (their accuracy may be inflated by the coder knowing outcomes). Q2, delivery added to case file + expert read: log-loss gain 0.035 per vote, one-sided 95% lower bound +0.008: delivery still adds information; votes 73.9% vs 71.4%. Already seen, OT2022-OT2023: Q1 88.1% vs 73.8% (42 cases, interval -28.6 to +2.4); Q2 gain 0.044, lower bound +0.022. ## Test 2c, 2026-10-05: the experts' read, uncontaminated Written before any recap was anonymised, read or coded for 2c. Why: the 2b coder (Sonnet 5.5) was trained on data through June 2026, so it may know the outcome of every case in OT2022-OT2025. Its 91.3% may be inflated. Coder: Claude Opus 4.8 (`claude-opus-4-8`), called headless through the Claude Code CLI with every tool turned off (`claude -p --tools ""`), default effort. Its published training-data cutoff is January 2026. Sonnet 5.5 and Opus 5.5 (cutoff June 2026) know every OT2025 decision; Opus 4.6 and Haiku 4.5 are older and weaker. Opus 4.8 is the strongest coder that cannot have seen these outcomes. Clean set: OT2025 cases with a recap, a usable SCDB vote and an SCDB decision date after 2026-01-31 (50 cases by the decision dates, counted before any coding). A case is also dropped if, in the knowledge probe below, Opus 4.8 states who won it (whether or not it is right). Every OT2026 case, coded before its decision, is added to the clean set as the term goes on. Anonymisation, by code, before any coder sees a recap: each word of the petitioner's and respondent's names (the petitioner is the party named first in the case caption, which handles the Oyez reversals) that is not a generic word is replaced by "Petitioner" or "Respondent"; when the United States, the President or a federal officer is a party, "United States", "the government", "the administration" and the officer's surname go to that side's label; the names of the advocates in the Oyez argument record become "counsel for Petitioner" / "counsel for Respondent" / "counsel for the amicus"; every date, month, weekday and four-digit year becomes "[date]"; sentences about when or how the Court will decide (e.g. "a decision is expected by summer") and any editor's update are removed. The justices' names stay. Every anonymised recap is scanned for leftover capitalised names before coding, and the scan is listed. Coding instruction (same for every coder, saved as `coder_prompt_2c.txt`): read the anonymised report of an oral argument and say which side the justices, as a bench, appeared more likely to favour, from the report's words only: petitioner, respondent or unclear. Knowledge probe, before coding: Opus 4.8 is asked, with the case name and no recap, whether it knows who won each clean-set case. Sonnet 5.5 gets the same probe, to show how much the 2b coder knew. Coder check, before Opus 4.8 codes anything: 20 anonymised clean-set recaps (seed 20261007) are read by me (Opus 5.5) and coded to `check_codes_2c.json`. Bar unchanged: where both take a side, opposite in at most 20%; otherwise test 2c is reported as inconclusive. Caveat, stated now: my own training runs through June 2026, so my reads may lean towards the outcomes I know; that can only add disagreement with a clean coder, not hide it. Also coded, to measure the inflation directly: Sonnet 5.5 on the same anonymised clean-set recaps, with the same instruction and settings. The 2b Sonnet codes (names not hidden) are already on file. Primary result, clean set only: on the cases where Opus 4.8 takes a side, the experts' accuracy on case outcomes against Phosphoros' (the frozen test 1 model, case file + delivery, majority of predicted votes), each with a 95% case-bootstrap interval and the paired interval of the difference (2,000 resamples, seed 20261005). Also reported: Phosphoros on all clean cases; the same head-to-head with Sonnet 5.5 (anonymised) and with the 2b Sonnet codes on the same cases; and, secondary, the test 2b log-loss gain (delivery added to case file + expert read, models as frozen in FROZEN5) with Opus 4.8 codes on the clean set. With about 50 cases these intervals will be wide; the result is reported whatever it shows. ### Correction, 2026-10-05: arguments lost to a docket-number mismatch Found while running test 2c, after its first report: some Oyez docket numbers carry a trailing space ("23-1197 ") and SCDB's do not, so the vote lookup missed those arguments and they were dropped as "no vote". Training terms: none lost. OT2022-OT2023: 5 of 112 cases lost. OT2024-OT2025: 11 of 115 lost. Every held-out result above was therefore computed on about 90% of the cases. The cases were dropped by a formatting quirk, not by anything about the case. Fix: the docket is stripped of spaces before the vote lookup (model.py, phase2.py). No frozen file changes (all were fitted on training terms, which were unaffected). The held-out runs are repeated with the fix: phase 1 (OT2022-OT2023), phase 2 (OT2024-OT2025), test 1, test 2b and test 2c. The original outputs are kept as *_before_docket_fix.json. The verdicts recorded above stand as the record of what was run; the corrected numbers are reported beside them, and where a verdict changes, both are stated. ## Test 5, 2026-10-05: words + voice against words alone Written before any transcript content was coded. Question: the experts read what the justices asked; Phosphoros reads how they spoke. Once a model has the content of the questions, does delivery still add information? Content channel. For each argument, each justice's own turns to each side (the same turns V3 keeps: 5 words or more, presiding procedure removed), anonymised by code: party names, the case name and docket number become "Petitioner"/"Respondent"/"this case"; advocates' names and forms of address ("Mr. Smith", "General Prelogar") become "counsel"; dates become "[date]". The justices are shown as "Justice A", "Justice B" ... in an order fixed by a hash of the argument, so the coder does not know who is speaking. One call per argument to Claude Opus 4.8 (headless CLI, tools off), instruction saved as `content_prompt_5.txt`. For each justice and each side the justice addressed, two ratings from 0 to 4: - skepticism: how skeptical of or hostile to that side's position the justice's questions and comments are (0 supportive, 2 neutral probing, 4 strongly hostile); - hypotheticals: how hard the justice pressed that side with difficult hypotheticals, slippery slopes or line-drawing problems (0 none, 4 many and pointed). Features per justice and argument: the side differences P - R of both ratings; a 0/1 input when the justice addressed only one side (differences then 0); and the bench, the mean of the other justices' differences. Case file, two versions. Full: as in test 1 (SCDB codes, known only after the decision, so usable on past terms only). Live: what is known on argument day: the justice, the United States' side (party or amicus, from the Oyez advocate descriptions) and justice x United States' side. Models, for each case-file version: A = case file + content; B = case file + content + delivery (V3 features as frozen in FROZEN2). Each model gets its own settings from the test 1 grid by leave-one-term-out log loss on the training terms (OT2015-OT2021, OT2020 excluded), then is fitted on all training terms. The live A and B, the prompt and the content codes of the training terms are frozen in FROZEN6 before any OT2026 argument is coded. Already seen, reported but not a test: OT2022-OT2025, B against A (log-loss gain per vote with the one-sided 95% case-bootstrap lower bound, vote and case accuracy), and A against the case file alone (does content add?). Also reported: the OT2025 cases decided after 2026-01-31, where the coder cannot know the outcome, as a check on whether outcome knowledge leaked into the training codes (if content predicts much better on seen terms than there, it may have). The test, OT2026 live: each argument is coded and both live models' vote probabilities are written to the hash-chained live log before the decision. Statistic as in test 4: per vote, log loss of A minus log loss of B; 2,000 case resamples, seed 20261005; delivery adds information beyond the words if the one-sided 95% lower bound is above 0. Closes at 80 decided OT2026 cases; interim results are labelled interim. The OT2026 SCOTUSblog recaps are coded by Opus 4.8 as in test 2c before each decision, so the clean expert comparison grows over the term. ## Result, test 2c, 2026-10-05: the experts' read, uncontaminated Anonymiser: three faults found while I read the check recaps, fixed before any coder ran: a state "Secretary of State" party was treated as the federal government; a headline was dropped when its first body sentence was removed; an advocate's surname ("Brown") was replaced inside "Ketanji Brown Jackson". The leftover-name scan (anon/_leftover_names.json) shows other litigants of consolidated cases, cited precedents and places; famous cases stay recognisable from their facts. Knowledge probe: Opus 4.8 said it knew the outcome of none of the 50 clean cases, so none was dropped. Sonnet 5.5 said it knew 3 (all right: 24-1287 tariffs, 24-557, 24-758). Coder check (my 20 blind reads; I first read 24-1260 without its headline because of the anonymiser fault, coded it after the fix): exact 18/20; where both took a side, opposite 0 of 15. Passes. Clean set, after the docket correction: 47 of the 50 cases have votes and a measured argument (24-1021, 24-1287 and 25-406 have no timed transcript). Opus 4.8 takes a side on 39: experts right on 34, 87.2% (95% interval 76.9 to 97.4); Phosphoros on the same 39 cases 64.1% (48.7 to 79.5); difference -23.1 points, paired interval -43.6 to -2.6. The experts are clearly better head to head, also when uncontaminated. Phosphoros on all 47: 63.8% (51.1 to 76.6). Inflation, measured directly: on the 32 cases all three coders called, Opus 4.8 (clean) 31/32, Sonnet 5.5 anonymised 31/32, Sonnet 5.5 2b (names shown) 30/32. Opus 4.8 and anonymised Sonnet never took opposite sides (38 cases); Opus 4.8 and 2b once (24-345). So outcome knowledge did not inflate the 2b experts. The 2b figure (91.3%) was high partly because that coder took a side only on clearer recaps: where it said "unclear" and Opus 4.8 took a side (7), Opus 4.8 was right 3 times. Secondary, delivery added to case file + expert read (FROZEN5 models, Opus 4.8 codes), 47 clean cases, 361 votes: log-loss gain 0.0055 per vote, one-sided 95% lower bound -0.031: not shown on this small clean set; votes 73.1% vs 77.0%. (Before the docket correction, 40 cases: experts 84.4% vs Phosphoros 65.6% on 32, interval -40.6 to +3.1; secondary gain 0.0125, lower bound -0.027. Kept in experts2c_clean_2025_before_docket_fix.json.) ## Result, docket correction, 2026-10-05: the held-out runs repeated | Run | Before (cases) | Corrected (cases) | |---|---|---| | Phase 1, OT2022-23: P vs best baseline | +2.65 pts, -0.72 to +5.93; failed (107) | +2.76 pts, -0.35 to +6.19; fails criterion 1 (112) | | Phase 2, OT2024-25: V3-raw vs best baseline | +4.60 pts, -0.26 to +9.30; failed (104) | +4.72 pts, +0.11 to +9.32; **passes all three criteria** (115) | | Phase 2, OT2022-23 seen | +4.95 pts, +0.72 to +9.08 (107) | +5.20 pts, +1.05 to +9.16 (112) | | Test 1, OT2024-25: delivery beyond case file | gain 0.041, lower +0.017 (104) | gain 0.044, lower +0.023 (115) | | Test 1, OT2022-23 seen | gain 0.048, lower +0.026 | gain 0.049, lower +0.028 | | Test 2b, OT2024-25 Q1 experts vs Phosphoros | 91.3% vs 73.9% (46) | 90.9% vs 70.9%, -32.7 to -7.3 (55) | | Test 2b, OT2024-25 Q2 delivery beyond experts | gain 0.035, lower +0.008 | gain 0.031, lower +0.007 | The phase 2 verdict changes. As run on 2026-10-05 it failed criterion 1; on all 115 cases its interval clears 0 by 0.11 points, so it passes, narrowly. Both are the record: the failure was the run as specified, the pass is the same run without the data bug. The bug was found in test 2c, not by looking for a pass, but a reader should weigh how thin the margin is. Phase 1 still fails. ### Test 5 implementation correction, 2026-10-05 (before full coding or model selection) The two-argument pilot hid speaker headings but left references such as "Justice Kagan's question" inside the excerpts. Those references can identify speakers by elimination. All justice names inside excerpts now use the same argument-specific Justice A/B labels as the headings (a justice absent from the excerpts becomes "another justice"). The old pilot is preserved separately and is not used in the fit. This implements the already specified identity masking; ratings, features, model grid and success criterion stay unchanged. Each coding response must contain every justice and exactly the sides shown, with both integer ratings in [0, 4], and must identify the requested coding model in CLI usage metadata. Missing or malformed outputs are recorded as failures and never silently filled as neutral. Inputs, prompt and successful outputs carry hashes; resumed runs must match them. Selection requires complete valid coding for its input arguments. Full coding cost is authorized by Ihor's request to continue this handoff; no recurring job is authorized or scheduled. Test 5 reporting detail fixed before fitting: the full case-file baseline is the already frozen test 1 case-only model. The live case-only baseline, needed to report whether content adds to the argument-day case file, uses the same training-only LOTO grid as A and B. This baseline is descriptive; the OT2026 criterion remains B against A. Bench content means include all other speaking justices, including those without a usable outcome label. A one-side rating contributes zero differences, as specified, and is counted explicitly. ### Test 5 extraction correction, 2026-10-05 (before full coding or selection) The V3 measurement CSV stores only the first 300 characters of each turn for inspection, although its word counts and delivery features use the full turn. The pilot content extractor incorrectly treated that preview as the question. Content now retrieves the complete turn from the cached Oyez timed transcript, using media ID, section and turn position and checking its timestamps. V3's turn selection stays unchanged. Pilot codes are preserved and invalidated by the changed input hashes. No past delivery measurements or frozen fits change. Test 5 live integrity detail, fixed before live coding: each run refreshes Oyez's case list, case status and timed transcripts; a failed refresh aborts rather than trusting a stale undecided status. Cases with a decision event or published decision are refused before coding and again before appending. Both A/B probabilities, coding hashes and frozen-model hashes are logged together. Recaps can appear after the forecast; their clean expert codes are appended as separate chained records before the decision, without changing a forecast. The entire existing chain is verified before any append. A local hash chain checks internal consistency; without an externally recorded head it cannot prove that its whole history was not rewritten. No public timestamp is claimed. ## Test 5 continuation status, 2026-10-05: coding blocked by Claude usage limit Opus 4.8's January 2026 reliable and training-data cutoffs were checked against Anthropic's current model documentation: https://platform.claude.com/docs/en/models/opus-4-8/overview . Successful CLI responses identify `claude-opus-4-8` in model usage metadata. Sandbox calls reported no login; the existing login works outside the sandbox. There are 367 available training arguments, 113 OT2022-23 arguments and 116 OT2024-25 arguments (596 total, before vote-based model exclusions). Complete retained questions contain 14,312,521 characters. The corrected full-question pilot succeeded, then the training batch reached HTTP 429: "You've hit your session limit ยท resets 7pm (Asia/Tbilisi)". Seven valid full-question training codes are saved, with $0.64493 reported list-price cost for those retained codes (excludes superseded pilots and the availability check; not an invoice). The batch was terminated; failed records remain explicit and resumable. No failed or missing rating was treated as neutral. The wrapper now stops retrying a login/usage-limit error and the batch stops new submissions, while saving any successful calls already in flight. `content_models.py` implements the prescribed full/live A/B fits, training-only LOTO grid, FROZEN6 and retrospective reports. It refuses incomplete or stale coding. It has not fitted anything: FROZEN6 does not exist and there are no new comparisons or test 5 conclusions. `live.py` now supports both live models and separately appended clean expert reads, refreshed Oyez status and verified hash chains. Its read-only check found no live records and no test 5 freeze. No OT2026 content has been coded, no forecast was written, and no recurring job was scheduled. No commits were made. Offline validation: `test5_checks.py` passes 10 checks of full-turn recovery, identity masking, invalid/missing ratings, side/bench differences, argument-day features, log tampering, post-decision refusal, majority probabilities, bounded stopping on session limits, and incomplete-coding refusal. These are tooling checks, not scientific results; fitting and live operation remain unverified. ## Test 5-Luna variant, 2026-10-05 (before any Luna coding or model fitting) Ihor explicitly requested Luna subagents after the Claude session limit. This supersedes the earlier instruction not to switch coders for this separate variant. GPT-6 Luna subagents now code the complete historical set using the unchanged `content_prompt_5.txt`, complete anonymised questions, opaque argument identifiers and exactly the same rating schema. No case outcomes, real speaker names, docket IDs, argument dates or existing Opus codes are supplied to coders. Each subagent receives a disjoint queue and writes a separate ratings file. All outputs are checked for full label/side coverage, integer ratings and unchanged input/prompt hashes before fitting. No missing rating is imputed. The seven valid Opus codes are preserved and are not mixed into Luna's fit. Luna's published knowledge cutoff is 2026-05-18: https://developers.openai.com/api/docs/models/gpt-6-luna . The old January-cutoff subset is therefore not a clean test for Luna. Historical results, including that subset, are descriptive only and carry no outcome-blind claim. A separate subset decided strictly after 2026-05-18 is reported as a post-cutoff sensitivity check, with no guarantee that outcome knowledge is absent. OT2026 before-decision coding remains the only prospective words-versus-delivery test. Feature definitions, training terms, LOTO settings grid, A/B comparison, bootstrap seed and 80-case live stopping rule are unchanged. Luna outputs, model files, reports and freeze are separate from Opus (`FROZEN6_LUNA`). No historical result is used to choose a coder or alter the model grid. Tools used by coder subagents are restricted by instruction to reading their assigned blinded texts and saving ratings; they must not browse or inspect other files. No code or model change is chosen on historical outcome performance. No commits, public timestamp or recurring job is authorized. ### Justice-name join correction, 2026-10-05 (before Luna fitting or reporting) Found while checking the new variant's justice mapping: SCDB's modern dataset contains both `RHJackson` (Robert H. Jackson) and `KBJackson` (Ketanji Brown Jackson). `model.scdb_name` previously joined on the surname alone and rejected ambiguous surnames. It therefore returned None for every Ketanji Brown Jackson historical row. Those rows were silently excluded from every past held-out comparison that uses this lookup. The live lookup, restricted to current justices, was unaffected. The training terms end at OT2021, before Ketanji Brown Jackson joined, so this does not change a training fit or its frozen settings. Fix: retain unique-surname matches; if several identifiers share a surname, require their first/middle initials to match the Oyez full name. Ambiguous names without enough initials remain missing. Before changing any report, save its current version as *_before_justice_fix.json, then repeat the frozen held-out runs (phase 1, phase 2, test 1, test 2b, test 2c). Keep both sets of results and report every changed verdict. Luna coding inputs are unchanged; the Luna variant uses the corrected vote rows when fitted/reported. This is a join correction, not a retuning or selection on held-out performance. ## Result, justice-name join correction, 2026-10-05 Only Jackson was ambiguous among the justices in the study. No training-term justice was ambiguous, and all training fits and frozen settings remain unchanged. The correction restores 88 eligible phase 1 votes in OT2022-23, 110 phase 2 votes in OT2022-23 and 115 votes in OT2024-25. Case counts stay 112 and 115. Current outputs and *_before_justice_fix.json preserve both versions. | Frozen comparison | Before correction | With Jackson restored | |---|---|---| | Phase 1, OT2022-23 | +2.76 points, CI -0.35 to +6.19; failed (869 votes) | +3.66 points, CI +0.42 to +7.02; passes all three criteria (957 votes) | | Phase 2, OT2024-25 | +4.72 points, CI +0.11 to +9.32; passed (868 votes) | +6.21 points, CI +2.04 to +10.44; passes all three criteria (983 votes) | | Test 1, OT2024-25 | gain 0.044, lower +0.023 | gain 0.0660, lower +0.0451 | | Test 2b head to head, OT2024-25 | experts 90.9%, Phosphoros 70.9%; paired CI -32.7 to -7.3 points | experts 89.1%, Phosphoros 72.7%; paired CI -30.9 to -3.6 (55 cases) | | Test 2b delivery beyond expert reads, OT2024-25 | gain 0.031, lower +0.007 | gain 0.0577, lower +0.0330 | | Test 2c clean head to head | experts 87.2%, Phosphoros 64.1%; CI -43.6 to -2.6 | experts 84.6%, Phosphoros 66.7%; CI -38.5 to +2.6 (39 cases) | Phase 1's verdict changes to a pass after correcting the missing votes. Its record model reaches 69.38% vote accuracy, while raw B4 reaches 70.11%: this does not show that the within-justice record beats raw differences. Phase 2's V3 model reaches 69.58% against the best baseline's 63.38%; its original failed run, docket-corrected narrow pass and this corrected pass all remain part of the record. The clean expert point estimate still leads, but its paired interval now crosses zero. The earlier claim that clean experts are clearly better is not supported by the fully joined data. No recap was recoded; including Jackson changes some majority-of-votes case labels as well as predictions. The coder check is unchanged (0 opposite among 15 both-side reads). Clean secondary comparison now uses 408 votes across 47 cases: gain 0.0400, one-sided lower +0.0046. Delivery beyond clean expert reads is shown on this subset. ## Result, test 5-Luna, 2026-10-05: retrospective content and delivery All 596 anonymised argument inputs have valid Luna ratings: 367 training, 113 OT2022-23 and 116 OT2024-25. Exact justice/side coverage, integer 0-4 scores and copied input/prompt hashes validate. Three Luna subagents coded disjoint inputs; the finished worker 1 took the final 30 unstarted worker 0 inputs. The dispatch records this reassignment. An account usage limit interrupted coding; the user authorised continuation, and saved ratings were preserved. No Opus ratings were mixed into this variant. Selection used 2,658 eligible training votes in 358 cases. The unchanged training LOTO grid selected logistic C=0.03 for both full A and B; live A selected the fixed boosted model, live B logistic C=0.03, and the live case-only baseline logistic C=0.01. The full case-only baseline remains the previous frozen logistic C=0.03. FROZEN6_LUNA now records the fitted models, training codes, prompt, coding provenance and feature source files, including the imported phase2.py and case3.py implementations. No held-out setting was changed after these results. The table gives log loss of A minus log loss of B per vote and the registered one-sided 95% case-bootstrap lower bound (2,000 resamples, seed 20261005). Positive gain favours delivery. These are already-seen historical reports, not the prospective test. | Historical subset | Cases / votes | Full gain / lower | Live-compatible gain / lower | |---|---|---|---| | OT2022-23 | 112 / 976 | +0.0501 / +0.0312 | +0.0620 / +0.0353 | | OT2024-25 | 115 / 983 | +0.0322 / +0.0148 | +0.0827 / +0.0527 | | OT2025 decisions after 2026-01-31 | 48 / 417 | +0.0300 / +0.0046 | +0.0698 / +0.0278 | | OT2025 decisions after Luna's 2026-05-18 cutoff | 32 / 284 | +0.0244 / -0.0015 | +0.0603 / +0.0290 | The full model does **not** clear the delivery criterion on the post-May-18 subset. The live-compatible model does. Neither subset establishes an outcome-blind historical experiment; OT2026 remains the prospective test. On OT2024-25, full A/B vote accuracy is 69.9%/72.3%, but case accuracy falls from 73.9% to 73.0%. Live A/B vote accuracy is 64.9%/71.4% and case accuracy 66.1%/72.2%. Log-loss improvement must not be described as improvement on every accuracy measure. Content beyond the case file (case-only minus A) also clears the lower-bound criterion for both versions on OT2022-23: full gain +0.0289, lower +0.0104; live +0.0324, lower +0.0052. On OT2024-25 full clears it (+0.0663, lower +0.0445), while live **does not** (+0.0336, lower -0.0015). On the post-May-18 subset both clear it: full +0.0783, lower +0.0413; live +0.0685, lower +0.0364. All unrounded losses and vote/case accuracies, including the January subset, are in content5_luna_seen_2022_2023.json and content5_luna_seen_2024_2025.json. This comparison measures frozen V3 features beyond two coded content ratings and their bench aggregates. V3 includes question/word counts as well as timing, interruptions and acoustic measurements. It does not establish a pure voice effect or improvement beyond every possible representation of the words. Integrity checks: 15 offline tests pass; git diff --check passes. The frozen live check validates the empty chain (genesis) and the Luna freeze. No recurring job or public chain timestamp has been installed. Live run after freezing: the sandbox could not refresh Oyez, so the authorised network run repeated the refresh successfully. Oyez lists 32 OT2026 cases and no argument audio or timed transcripts. The runner wrote no content codes or forecasts. The log remains empty; absence of available argument data is not a prospective test result. ## Closing note, 2026-10-05 A parallel Claude session resumed Opus 4.8 content coding before it learned of the Luna variant and was stopped after 3 more arguments: the Opus file (media/scotus/content/codes_5_2015_2021.json) holds 10 valid codes, not 7. They are not used by the Luna fit or any reported result. The test is concluded for now at Ihor's call: no live job is scheduled, the live log is empty, and OT2026 remains the only prospective test if it is resumed. ## Test 6, 2026-10-06: can Phosphoros match the experts? (before any coding or fitting) Ihor asked to build three improvements, freeze them and test against the experts. Written before any test 6 coding, fitting or result. **Coder.** `claude-opus-4-8` (training cutoff January 2026) through `claude -p --tools ""`, the same coder as the clean expert reads of test 2c. One call per argument with `content_prompt_6.txt`. Input: every justice turn of 5+ words that V3 keeps, in time order, numbered, under headers saying which side's counsel is at the lectern. Party, counsel, docket and date names are anonymised as in test 5 (`content.anonymise_turn`); justices keep their names, as in the expert recaps. Output per argument: for each justice a lean (integer 0-100, probability of a vote for the Petitioner) and a reason; for every turn a stance toward the side at the lectern (-1 challenges it, 0 neutral or procedural, +1 helps it); and the case name if the coder recognises it. Invalid outputs are retried up to twice and otherwise left missing; nothing is imputed except the documented zero-plus-flag for missing features. **Coded sets.** All training arguments 2015-2021 (selection only) and the OT2025 arguments of the 50-case test 2c clean set (decided after 2026-01-31, the coder's cutoff). OT2026 arguments are coded live, before decision. **Features per justice vote** (all available on argument day): - Case (live-compatible): justice identity, side of the United States. - 1, words: logit of the justice's lean. - 2, friendly vs hostile delivery: for each side, hostile turns per minute, friendly turns per minute, hostile words per minute, cut-ins on hostile turns per minute; mean pitch (semitones) of hostile turns. Features are the Petitioner-minus-Respondent differences; missing pitch is 0 with a flag. - 3, the court: mean lean logit of the other justices; mean lean logit of the other justices in the same bloc; the same two means of the net hostile-minus-friendly rate. Blocs fixed a priori: liberal (Ginsburg, Breyer, Sotomayor, Kagan, Jackson), conservative (Scalia, Thomas, Alito, Gorsuch), centre (Kennedy, Roberts, Kavanaugh, Barrett). - Delivery V3: the frozen phase 2 raw channels, flags and bench counts. **Models.** R: the coder's lean alone (no fit). A: case + 1. B: A + V3. C: B + 2. **D (primary, "Phosphoros 6"): C + 3.** Each A-D is a standardised logistic regression; C chosen from {0.01, 0.03, 0.1, 0.3, 1} by leave-one-term-out log loss over the training terms (2015-2019, 2021) only. Training arguments whose case the coder names are excluded from fitting and selection (recall would inflate the lean's weight); their count is reported. Case outcome = majority of predicted votes, as before. Settings and models are frozen in `FROZEN7` before any test-set prediction is computed. **Clean test (once, after freezing).** The 2c clean set (OT2025 decided after the coder's cutoff), cases with votes and measures: 1. Head to head on the cases where the clean expert read (Opus 4.8, anonymised recaps, existing codes) takes a side: D's case accuracy minus the experts', paired case bootstrap (2,000, seed 20261005). D **beats** the experts if the 95% interval lies above 0; **is worse** if it lies below 0; **matches** only if the point estimate is within 5 points of the experts (or above) and the interval includes 0. With about 40 cases the test is underpowered; a tie is weak evidence, not proof of equality. 2. Delivery beyond the words: log loss(A) minus log loss(D) per vote, one-sided 95% case-bootstrap lower bound above 0. 3. Reported, not criteria: R, A, B, C; D on all clean cases; experts where sided with D filling the rest; per-justice accuracy; recognised cases. Any clean case the coder recognises is reported; it is not excluded. **Live.** OT2026 arguments get the same frozen D before decision, logged to a separate hash chain `live6_predictions.jsonl`; the comparison with live expert reads closes at 80 decided cases with criteria 1 and 2. No commits, schedule or public timestamp are authorised by this test. ### Test 6, amendment, 2026-10-06 (after a 3-argument pilot, before any fitting) The pilot coder named the real case for all three arguments, including obscure ones (Hawkins v. Community Bank of Raymore, OBB Personenverkehr v. Sachs, Ocasio v. United States). Excluding recognised training arguments would leave almost no training data. Amendment: fit and select on all training arguments; report how many the coder recognised. Recall in training can only make the frozen model trust the coder's lean too much; it cannot inflate the clean test (decided after the coder's cutoff) or the live test. Pilot cost was about $0.25 per argument at list price. ### Test 6, Luna coder amendment, 2026-10-06 (before any Luna coding or fit) The test 6 Opus pilot and partial training coding are preserved in their separate profile. The Luna profile uses the same test 6 prompt, integer 0-100 justice lean, per-turn -1/0/+1 stance, features, candidate models, selection and criteria. Its inputs and outputs are isolated from Opus under `media/scotus/content/luna6/`, with independent `FROZEN7_LUNA`, model artifact and live log. Coders receive only the prompt, an opaque queue entry and that argument's anonymised transcript input. They receive no case IDs, dates, outcomes, manifests or historical results. Every argument must validate against its source input and prompt hashes, model, justice and turn coverage, and integer output contract before assembly or fitting; incomplete or stale coding blocks both. Luna's published training-data cutoff is 2026-05-18. The January 2026 clean subset was outcome-blind for Opus but is not outcome-blind for Luna. All Luna historical clean results are descriptive. OT2025 arguments decided strictly after May 18 may be shown as an optional post-cutoff sensitivity analysis, with no guarantee that outcome knowledge is absent. OT2026 coded before decisions remains the only prospective Luna test. No sensitivity analysis changes the primary criteria or permits refitting after the live test begins. ### Test 6 Luna result, 2026-10-06 (after the frozen historical comparison) All 414 opaque inputs validated: 367 training arguments and 47 historical evaluation arguments. No Opus codes were mixed in. Settings selected on training terms only were C=0.01 for all A-D variants; 2,658 training votes. The model, prompt, training codes, selection/validation code and dispatch were frozen in FROZEN7_LUNA before evaluation. On the 39 expert-sided cases, D was right on 28 (71.8%) and the experts on 33 (84.6%): difference -12.8 points, paired 95% interval -33.3 to +7.7. Matching or beating the experts is not demonstrated. D called 34/47 cases (72.3%) and 300/408 votes (73.5%) correctly. Experts first, D for uncalled cases, called 39/47 (83.0%). A-to-D log-loss gain per vote was 0.0941, one-sided 95% lower bound 0.0556. This passes the delivery-beyond-words criterion historically; D includes turn counts/timing and court features as well as acoustics. It does not establish a pure voice effect. All results remain descriptive because the January-cutoff evaluation set is not outcome-blind for Luna. No prospective forecasts or public timestamps were produced. ### Test 6, Opus subset amendment, 2026-10-06 (before any Opus fit or clean coding) Ihor asked for a cheaper Opus run. The Opus profile stopped at a usage limit after 121 valid training codes (all 68 OT2015 arguments, the first 53 OT2016 arguments in date order; two others failed validation and stay excluded). Amendment: the Opus profile is fitted on exactly those 121 arguments. Settings are chosen by the same grid and log loss, leave-one-term-out over the terms present (2015, 2016). The selection is `test6_opus121.py`. It is frozen in `FROZEN7` with the model, prompt, training codes, test6.py and itself, before any clean argument is coded. Then the 47 clean OT2025 arguments are coded blind (no outcomes, scores or case lists in the coder's input) and `test6.py clean` runs once. Clean scoring requires all 47 codes to be valid; failed codes are retried; nothing is scored on a subset. Criteria 1 and 2 are unchanged. Training on a third of the data makes D noisier. R, the coder's read alone, needs no fit and is the main reported comparison with the experts, whose recaps were also coded by Opus 4.8. Further Opus training codes, if ever added, form a new fit for the live test only, never a second clean score. ### Test 6 Opus result, 2026-10-06 (clean set, run once after FROZEN7) All 47 clean OT2025 arguments coded validly (Opus recognised 22 by name; all were decided after its January 2026 cutoff). Settings chosen on the 121 training codes: A C=0.03, B C=0.03, C C=0.1, D C=0.03. Criterion 1 (primary, D): on the 39 expert-sided cases D was right on 30 (76.9%), the experts on 33 (84.6%): difference -7.7 points, paired 95% interval -20.5 to +5.1. Verdict **not shown** (point estimate more than 5 points behind). Criterion 2 **fails**: A-to-D log-loss gain 0.0485 per vote, one-sided lower bound -0.0032. Reported, not criteria: R, the coder's read with no fit, was right on 34 of the 39 expert-sided cases (87.2%) and 39 of 47 overall (83.0%). R minus experts +2.6 points, paired 95% interval -5.1 to +12.8 (computed after the run; not a registered criterion). R and the experts agree on 36 of 39. R is more accurate on cases it did not recognise (88.0%, 25 cases) than on those it did (77.3%, 22). Case accuracy on all 47: R 83.0%, D 78.7%, C 76.6%, B 78.7%, A 74.5%; vote log loss D 0.459, R 0.472, A 0.507. D called the petitioner in 57% of cases and caught 14 of 18 respondent wins (the Luna profile: 6 of 18). Experts first, D for the 8 uncalled cases: 85.1%. Reading: the reading of the words, not the delivery, closes the gap to the experts here. On this set the fitted model (trained on 117 arguments) loses case accuracy against the raw read, and delivery beyond the words is not shown. A single 39-case retrospective set cannot establish equality with the experts. OT2026 remains the prospective test. ## Reporting note, 2026-10-07: case calls on the public page Not a criterion, not a refit, no new model. For phosphoros.xyz/tests, phase 2 case calls are scored against the actual winner in SCDB (`partyWinning`), not against the majority of measured votes. The two rules differ in two OT2024 cases where unmeasured votes broke a tie among the measured ones: 23-1002 (Hewitt v. United States) and 23-1270 (Riley v. Bondi), both won by the petitioner 5-4. A case is called for the petitioner when more than half of the measured justices are predicted to vote for the petitioner; a tie is a respondent call. V3-raw (FROZEN2), OT2024-OT2025: 84 of 115 cases (73.0%, 95% case-bootstrap interval 65.2 to 80.9); always the petitioner 77 of 115 (67.0%). The paired difference is +6.1 points, interval -0.9 to +13.0: a case-level gain is not shown; the vote-level result above is the claim. Scored by the measured-vote majority instead: 82 of 115 (71.3%) against 75 of 115 (65.2%). An earlier public figure, 72.6% of 113, left the two tied cases out. Test 2c clean-set figures are the same under either rule (no ties): expert read 33 of 39; Phosphoros (frozen test 1 model) 26 of 39, 32 of 47 overall, and 6 of the 8 cases the expert read left unclear. Every public figure is produced by `report/public_numbers.py` (read-only; it checks the freezes and refits nothing) into `report/public_numbers.json`.