Online communities pilot

Pilot report: r/singularity, vanguard sub-series, 2024Q4 to 2026Q3

A working document. This is one of the documents the site is built from, published as it is written, in the project's own terms, for anyone checking the work. It assumes the reader knows the method. The plain account of the same material is on the home page.

Pilot report: r/singularity, vanguard sub-series, 2024Q4 to 2026Q3

Model-coded, no human coder. Updated 2026-10-08: the figure and the results now come from the majority codes of the full-sample run (all 390 posts, three models, final codebook). Updated 2026-10-04 with Gemini's trial results (all three models complete on both trials). Updated 2026-10-03 after John's rulings: first on the first version of this report, then on the reliability design, which he revised twice that day. Coding is fully model-based (amendment section 9; log entries of 2026-10-03). Run record: docs/vanguard/pilot-run-r-singularity-2026-10.md; codebook: docs/vanguard/codebook.md.

How coding was checked. Three models from different companies each coded all 390 screened posts blind, under the final codebook (rules R2 and R3 adopted 2026-10-03, R1a adopted 2026-10-04):

The final code for each field is the majority vote: the value at least two models gave (data/vanguard/r-singularity-pilot-codes-majority-full-final.csv); where no two agree the field is "unresolved". Models from different companies can still share blind spots; this is a disclosed limitation. An earlier check on a 140-post subset, under the codebook before the rules, is kept under Agreement for the record.

Figure

site/vanguard-pilot-r-singularity.png (counts in site/vanguard-pilot-r-singularity.csv), labelled "model-coded" and captioned "Framing in sampled high-score posts from selected AI enthusiast venues; not representative of participants or the public."

pilot figure

Results (majority codes, all 390 posts, final codebook)

Written 2026-10-08. The figure is now drawn from the majority codes (scripts/vanguard_figure.py), not from Claude's first pass. Three things about how the numbers are counted:

Primary measures. Share of eligible posts by what the main claim is about, all quarters together. Denominators differ because of the unresolved posts above; 156 top posts and 104 random posts are counted in all.

measure top posts random posts
main claim is a specific capability (a model, benchmark, robot or demo) 89 of 153 (58%) 52 of 103 (50%)
main claim is a consequence (jobs, society, the economy) 14 of 153 (9%) 16 of 103 (16%)
main claim is industry and governance (AI company and policy drama) 43 of 153 (28%) 28 of 103 (27%)
capability posts that extrapolate to a timeline or consequence 7 of 89 (8%) 8 of 51 (16%)
headline claims (poster affirms; ironic posts left out) 0 of 143 (0%) 1 of 96 (1%)

Leaving out posts coded from the title alone (59 of the 156 top posts, 37 of the 104 random posts): capability 40% and 42%; consequence 13% and 20%; industry and governance 41% and 32%; extrapolation 8% and 30%; headline claims 0 and 0.

By quarter, capability main claims range from 39% (2025Q1) to about 80% (2025Q3, 2025Q4) of top posts, and industry-and-governance claims peak in 2025Q1 (11 of 18; DeepSeek, Stargate, the Musk–Altman fight) and 2026Q1 (10 of 20; Anthropic and the Pentagon). Extrapolation from a capability to a timeline or consequence is uncommon throughout: at most 3 of 11 capability posts in any quarter of top posts.

Headline claims, near-absent (poster affirms; any of AGI already achieved, AGI by a stated horizon, self-improvement under way, singularity framing): no top post in the whole window and 1 random post (2024Q4, an AGI forecast with a stated horizon). Three top posts affirm "AGI achieved" or singularity framing by majority but are coded ironic by majority (the titles "AGI 🚀", "AGI achieved" and "Post-Singularity Free Healthcare"), so they are left out; counting them would give 3 of 143 (2%). One more "AGI achieved" top post (2025Q2) has no majority on whether the poster agrees. AGI main claims of any kind are 4% of top posts (6 of 156) and 6% of random posts (6 of 104); self-improvement main claims 0 and 1.

Samples. Top posts: the 230 posts screened in the first pass were coded by all three models; 161 are eligible by majority. The 20-eligible cut-off was re-applied per quarter from the majority codes: rank 26 in 2024Q4, 26 in 2025Q2, 27 in 2025Q3, 30 in 2026Q1 and 28 in 2026Q2, leaving 5 eligible posts below those ranks out. Three quarters fall short of 20: 2025Q1 has 19 eligible among ranks 1 to 29, 2025Q4 has 19 among ranks 1 to 32 and 2026Q3 has 18 among ranks 1 to 26; every screened post counts there, and reaching 20 would need the three models to code further down each quarter's ranked reserve. 156 top posts are counted. Random posts: 160 screened, 104 eligible by majority, all counted. Exclusion reasons are each model's own and are not voted; the majority file records only whether the post counts.

Agreement

Three models, all 390 posts, final codebook (2026-10-07)

This is the check the figure rests on. Cohen's kappa (agreement corrected for chance) for each pair of models; n is the number of posts both coders counted. "Whether the post counts" is compared on all 390 posts; the other fields on the posts both models in the pair found eligible. "No majority" is the share of the 265 posts the majority finds eligible where all three gave a different answer. Full table: data/vanguard/three-model-reliability-full-final.csv.

What is coded Claude–Gemini Claude–GPT Gemini–GPT No majority Script's verdict
Whether the post counts 0.91 (390) 0.91 (390) 0.94 (390) 0% reliable
What the claim is about 0.77 (259) 0.83 (254) 0.76 (258) 2% reliable
Timing (already here, forecast, …) 0.65 (259) 0.63 (254) 0.63 (258) 5% reliable
Whether the poster agrees 0.75 (259) 0.63 (254) 0.56 (258) 3% reliable
Hedging 0.73 (259) 0.66 (254) 0.54 (258) 0% reliable
Tone of the post 0.74 (259) 0.75 (254) 0.79 (258) 0% reliable
Whether the poster wants it faster or slower 0.46 (259) 0.59 (254) 0.45 (258) 0% unreliable (every pair below 0.6; ruled unreliable by John, 2026-10-03)
Whose claim it is 0.72 (259) 0.71 (254) 0.71 (258) 1% reliable
Coded from the title alone 0.99 (259) 0.99 (254) 1.00 (258) 0% reliable
Projects to a time horizon 0.75 (259) 0.72 (254) 0.73 (258) 1% reliable

Under John's test (unresolved on more than 15% of posts, or every pairwise kappa below 0.6), every field passes except whether the poster wants AI sped up or slowed down, which is kept out of figures. The figure's five measures use whether the post counts, what the claim is about, timing, whether the poster agrees, tone (for the irony exclusion), projects to a time horizon and coded from the title alone; all pass, so all five panels are drawn. Gemini and GPT agree less with each other than either does with Claude on whether the poster agrees (0.56) and on hedging (0.54); both fields still pass because the other two pairs clear 0.6.

Earlier check: the 140-post subset, before the rules (2026-10-03)

Kept for the record; superseded by the table above. Cohen's kappa for each pair of models. "Whether the post counts" is compared on all 140 posts; the other fields on the posts both models in the pair found eligible (78 to 85). "Unresolved" is the share of the 88 posts the majority finds eligible where no two models agree. Full table: data/vanguard/three-model-reliability.csv.

field Claude–Gemini Claude–GPT Gemini–GPT unresolved verdict
whether the post counts 0.82 0.82 0.91 0% reliable
what the main claim is about 0.57 0.62 0.75 6% (5 posts) reliable
forecast or already happened 0.66 0.66 0.77 5% reliable
whether the poster agrees 0.38 0.36 0.69 1% reliable
hedging 0.22 0.18 0.56 2% unreliable
irony 0.60 (0.595) 0.57 0.58 3% unreliable
whether the poster wants AI sped up or slowed down 0.55 0.43 0.65 0% reliable
whose claim it is 0.67 0.64 0.78 1% reliable
coded from the title only 0.97 0.97 1.00 0% reliable
draws a timeline or consequence from a capability 0.71 0.64 0.65 0% reliable

On this subset, under John's test, hedging and irony were unreliable. Both pass on the full sample under the final codebook (table above). No field came close to the 15% unresolved limit.

Claude was the outlier on this subset. Gemini and GPT agreed with each other more than either agreed with Claude on almost every field. Whether the poster agrees passes only because the Gemini–GPT pair reaches 0.69, and what the main claim is about passes because two pairs reach 0.6. The same-family check (two copies of Claude, before the rulings: kappa 0.87 on what the claim is about) badly overstated reliability.

Codebook rules R1 to R3, before and after

Where things stand (2026-10-04). John adopted R2 and R3 on 2026-10-03 (now in docs/vanguard/codebook.md) and split R1 into two rules to be trialled separately: R1a for whether the poster agrees and R1b for whose claim it is. Their results are in the next section. On 2026-10-04 he adopted R1a, with a sentence on jokes and reaction titles added, and rejected R1b. The first trial below is now complete on all 140 posts.

The rules are in docs/vanguard/proposed-rules-2026-10-03.md:

All three models recoded the 140 posts with the three rules added. Each rule targets a different field, so they were tested together. Claude and GPT coded all 140 posts on 2026-10-03. Gemini coded 112 that day, when its daily request limit ran out, and the last 28 on 2026-10-04 under the same frozen codebook (docs/vanguard/codebook-frozen-2026-10-03-before-rules.md). Kappa before and after on all 140 posts:

field rule Claude–Gemini Claude–GPT Gemini–GPT
what the main claim is about R2 0.57 → 0.82 0.62 → 0.85 0.75 → 0.79
whether the poster agrees R1 0.38 → 0.75 0.36 → 0.74 0.68 → 0.63
hedging R3 0.21 → 0.77 0.18 → 0.53 0.56 → 0.56
whose claim it is (R1 side effect) 0.66 → 0.69 0.64 → 0.81 0.78 → 0.57
whether the poster wants AI sped up or slowed down (none targeted) 0.55 → 0.52 0.43 → 0.65 0.65 → 0.47
irony (none targeted) 0.60 → 0.71 0.57 → 0.75 0.58 → 0.74
forecast or already happened (none targeted) 0.66 → 0.75 0.65 → 0.74 0.77 → 0.71

With the rules, hedging and irony pass the test. On the first 112 posts (the figures John ruled on) the rows read much the same, with two exceptions: hedging between Gemini and GPT was 0.55 → 0.62 and is now 0.56 → 0.56, and whose claim it is between Gemini and GPT was 0.80 → 0.62 and is now 0.78 → 0.57.

How to read this:

  1. The Gemini–GPT column is the cleanest test. Between Gemini's and GPT's two runs only the rules changed. The Claude trial codes came from fresh coders, so the Claude pairs mix the rules' effect with the change of coder.
  2. R2 helps in every pair, including Gemini–GPT (0.75 → 0.79). R3 helps a lot in the Claude pairs but not between Gemini and GPT (0.56 → 0.56 on all 140 posts). Hedging still passes the test through Claude–Gemini (0.77); the fresh-sample check should look at it closely.
  3. R1 brings Claude into line but slightly lowers Gemini–GPT agreement, both on whether the poster agrees (0.68 → 0.63) and on whose claim it is (0.78 → 0.57). The trial coders found that R1 has no default when a post names no source, and no answer for a model's own output in a screenshot. Refinement for John to consider: when no source is named, attribution = original; a model's output = lab_or_company.
  4. Optimistic by construction. The rules were written from disagreements on these same posts, so they will look better here than on new posts. Any rule John adopts is checked again on a fresh sample (section 9).
  5. Other refinements the trial coders raised:
    • R3: an "if" that sets a condition is not a hedge.
    • R2: a launch or release decision is the company's act, but what the product does is a capability claim. When a title makes both claims, the main claim needs a tie-break.
    • R1: "denies" means rejecting that the claim is true. Disliking what is reported goes in development preference.

The trial codes (*-rules-trial.csv) aren't used in any figure or count.

R1a and R1b, each tried alone (ruled 2026-10-04: R1a adopted with the reactions sentence, R1b rejected)

Run 2026-10-03 (Claude, GPT) and 2026-10-04 (Gemini) on the same 112 posts (data/vanguard/r1-split-trial-posts.csv), with today's codebook (R2 and R3 adopted) plus one rule each. Each run serves as the other's comparison, so "without" means the same codebook, day and models with only the rule under test missing (vanguard_three_way.py --r1-split). The 90% ranges come from resampling the posts 2,000 times (scripts/vanguard_r1_bootstrap.py).

field rule Claude–Gemini Claude–GPT Gemini–GPT
whether the poster agrees R1a 0.45 → 0.78 0.50 → 0.74 0.53 → 0.65
whose claim it is R1b 0.77 → 0.83 0.75 → 0.72 0.78 → 0.75

Decision packet (recommendation followed by John's ruling: refine and adopt R1a; reject R1b): docs/vanguard/decision-packet-r1a-r1b-2026-10-04.md. The split-trial codes (*-r1a-trial.csv, *-r1b-trial.csv) aren't used in any figure or count.

Edge calls still open

Decisions for John

  1. Rules R1a and R1b ruled 2026-10-04: R1a adopted with the reactions sentence, R1b rejected. (R2 and R3 adopted 2026-10-03.)
  2. Where the figure's numbers come from done 2026-10-08: the full-sample run finished on 2026-10-07 and the figure is drawn from the majority codes (Results above).
  3. Two agent readings to confirm or overrule (made 2026-10-08 while redrawing the figure; run record of that date): posts coded ironic do not count as affirming a headline claim (3 top posts affected); and the claim-kind shares count the main claim only, because secondary claims were never compared across the models.
  4. The three short quarters. 2025Q1, 2025Q4 and 2026Q3 have 19, 19 and 18 eligible top posts by majority, not 20. Reaching 20 means having all three models code the next few posts down each quarter's ranked reserve (a handful of posts, a few cents each for Gemini and GPT); or the shortfall can stand and be disclosed, as it is now.

Not done

The full-sample run is complete (2026-10-07) and the figure uses it (2026-10-08). Not done: the three short quarters above; the cross-era sample and the comment sample (deferred); and the section 12 preconditions (venue inventory, X decision). No other venue was collected.

Rendered from docs/vanguard/pilot-report-r-singularity-2026-10.md at commit 5cc5451.