Document

Amendments

Amendments

Dated changes to the method, the panel or the codes, with the reason. Nothing in the method changes without an entry here. Reviews referred to below are saved verbatim in docs/reviews/.

2026-09-30: reconciliation of the two outside reviews (Gemini 3.1 Pro, GPT-6 Astra Medium)

Both reviews were read against the kickoff plan (docs/2026-09-30-kickoff.md). The result is docs/methodology-v2.md, which supersedes the method sections of the kickoff plan. The kickoff plan stays as the record of where the project started. John ruled on 2026-09-30 that no third review would be sought and that these two would be worked with.

A1. The headline claim is replaced by three separately registered outcomes (GPT)

Was: one claim, "the central public question moved from whether, to when, to now". Now: three outcomes measured and reported separately. What surveyed populations believed (surveys). What published outlets argued (news and opinion in fixed outlets). What specialist communities discussed (forums and the named expert panel). The whether, when, now story is tested in each, with a written failure condition, and the essay reports where the three agree and where they diverge. Why: "central public question" had no observable definition, no specified public, and no way for a different question (jobs, deepfakes, reliability) to win. GPT's point that the old structure would read changes in attention, vocabulary and membership as changes in opinion was correct. This also absorbs Gemini's "insider projection" objection, which becomes the expected divergence between the specialist and population outcomes rather than a flaw.

A2. The single stance code is split into endorsement, affect and agency (GPT; Gemini's disruption profile folded in)

Was: stance with six values from dismissive to resigned. Now: endorsement (accepts, rejects, uncertain, conditional, unexpressed), affect (enthusiasm, fear, resignation, neutral, mixed, unexpressed), agency (can influence the outcome, cannot, unexpressed), consequence_type (existential, economic, social, mixed, unspecified). Why: the old scale mixed a judgment (dismissive) with an emotion plus a view about control (resigned). Someone who expects AGI in five years can be thrilled, frightened, resigned or organizing. Gemini's "disruption profile" (existential, economic, social) is the consequence_type field. Old seed codes are recoded, not mapped, because a mapping would launder the old scale's ambiguity into the new fields.

A3. A target-capability field is added, with self-improvement split four ways (GPT)

Was: statements were about "AGI" as one thing. Now: target_capability says which capability the statement is about: passing as human in conversation, a specific task, broad competence at most economically valuable work, exceeding humans broadly, or one of four self-improvement levels (AI helps humans build AI; AI does parts of AI research on its own; AI changes its own code or weights; AI sustains successive improvements with little human involvement). Why: Pew found in 2010 that 81% of US adults expected a computer to converse indistinguishably from a human by 2050. That is early acceptance of one capability, not of AGI, and the old scheme could not keep the two apart. The four self-improvement levels get different answers from the same person.

A4. question_frame is kept but widened so other questions can win (GPT item 3; Gemini's capability-versus-impact test)

Was: six frames, all about AGI itself. Now: the six remain, plus consequences, governance and other_concern with a sub-tag (jobs, reliability, privacy, bias, misinformation, military, consciousness, ownership, other). Each unit records the primary frame and every frame present. Question presence is coded separately from the answer given. Why: if jobs or deepfakes was the leading question in mainstream outlets after 2023, the data must be able to say so. Gemini's capability-versus-impact comparison is now a standing report, not a special test.

A5. Expert panel eligibility is made position-independent and datable (GPT)

Was: statements before 2017 and after 2024, plus "prominent enough to have shaped public discussion". Now: statements before 2017 and after 2024, plus named in at least three distinct national-outlet pieces about AI futures dated 2017 or earlier (the prominence test is dated and checkable, not retrospective). Quiet years are recorded as rows in the search log. Within-person change is reported separately from change in who is speaking. Why: choosing today's prominent thinkers selects for the arguments that survived. The seed session's panel-verification step is amended to collect the three-piece evidence.

A6. Forum sampling: random samples for prevalence, top posts kept as a separate series (both reviews)

Was: the 20 top-scoring r/singularity posts per quarter. Now: for each community and quarter, a seeded random sample of posts and, independently, of comments, with inclusion probabilities recorded, drawn without regard to score. The top-20 series is kept and labeled "what the community rewarded". Account concentration (share of sampled units from the ten most prolific accounts) is reported, and a stable-account subset is analyzed to separate changed speech from membership turnover. Communities are reported separately and never pooled as if they were independent samples of the public. Why: scores measure virality, and Arctic Shift changed how it captures score fields in November 2023, right on the proposed period boundary.

A7. Vocabulary: piloted across the whole period before freezing, with era terms (both reviews)

Was: a list of current terms frozen before the pull. Now: a retrieval vocabulary built from a pilot sample spanning 2005 to 2026, including older terms (thinking machines, strong AI, human-level machine intelligence, machine intelligence, intelligence explosion, technological singularity, Skynet, robot uprising) alongside current ones. Version frozen before the main pull; precision and recall checked per era on a held-out sample; any later revision reruns every year. The 1% threshold in the vocabulary ledger is replaced by a published numerator, denominator, distinct-account count and single-thread share, with a minimum volume of 500 comments in the month and a persistence rule of three consecutive months. Why: a list chosen today would make it look as if nobody discussed this before 2015 because they used different words.

A8. Media sample: sized from the smallest change worth detecting, stratified, news and opinion coded separately (GPT; Gemini)

Was: 15 opinion pieces per year. Now: outlets fixed to those with documented public archives (New York Times Archive API for metadata from 1851; Guardian Open Platform for full text from 1999; Media Cloud collections only after an outlet-by-month coverage audit). The eligible frame is enumerated per outlet and period before sampling. Periods are three years, not one. Sample size per outlet-period is set to distinguish a 20-point change in a share at 95% confidence (about 50 pieces); where fewer are eligible, all are taken. News and opinion are coded separately, and the author's own endorsement is coded separately from claims the piece quotes. Why: with 15 pieces the margin of error on a share near 50% is about 25 points, which cannot support an annual claim. Op-eds measure what editors chose to publish, so this outcome is labeled "published argument", never reader belief.

A9. Attention series kept as context only, never combined (GPT)

Each series (Wikipedia page views from 2015-07 with the 2007 to 2015 legacy counts marked as a different measurement; Google Trends with export settings recorded; GDELT optional pending a retrieval pilot) is normalized within itself and shown separately. No composite "public acceptance" index. Google Books Ngram moves to the vocabulary appendix, unsmoothed, because Google drops phrases found in fewer than 40 books, which makes it unusable for first appearances.

A10. Surveys become the backbone of the population outcome, with specific early anchors (both reviews)

Added: Pew/Smithsonian 2010 and 2014 (capability expectations), Eurobarometer 2012, 2014 and 2017 (robots, multi-country), Zhang and Dafoe 2018 (GovAI, high-level machine intelligence timelines), Lloyd's Register Foundation World Risk Poll 2019, 2021 and 2023, Pew's 2025 public-versus-expert comparison (parallel questions, the model for comparability), Monmouth 2015 and 2023, YouGov 2023 onward, Gallup, Ipsos with its own note that several national samples skew urban and educated. Every row dated by fieldwork, with publication date separate, exact wording, and a construct label (concern, feasibility, timeline, impact). The absence of a comparable population series for 2005 to 2015 is stated, not papered over.

A11. Prior work to build on, not redo (both reviews)

Fast and Horvitz 2017 (New York Times coverage of AI, 1986 to 2016; its categories are the starting point for media coding), Zhang and Dafoe 2019, Pew 2025, the AI Impacts survey series, the Stanford AI Index public-opinion chapter, Gnambs and Appel on Eurobarometer waves, Gaffney and Matias 2018 on Reddit archive gaps (reason for the coverage audit), GCRI's survey of AGI projects.

A12. Reliability and reproducibility rules (GPT)

Double-coding of a stratified subset with dates and identities hidden where practical; disagreements and per-category agreement published; ambiguous and unexpressed kept as codes; any automated coding validated separately by era and source. Published with the data: sampling frames, queries, code, random seeds, source versions, retrieval dates, document identifiers, exclusions, missingness audits, and snapshots with checksums where the source permits.

A13. Expert panel expanded into three tiers with written quotas (John, 2026-09-30, after discussing panel size)

Was: a core panel of 15 to 20 plus an unsized 2020s cohort. Now: three tiers.

Declined, with reasons

2026-09-30: implementation decisions for version 2 (decided by: agent)

Made while carrying out the version 2 migration. John can reverse any of them.

2026-10-01: John's rulings on the fifteen open decisions

Decided by: john. The decisions were listed in docs/status-2026-10-01.md; the rulings below were sent by John on 2026-10-01. Where an agent had to interpret a ruling to apply it, the interpretation is marked as such.

A14. Rulings on panel eligibility, the codebook, and process

Panel

  1. Was: unclear whether a piece the person wrote counts as "naming" them for the three-national-pieces test. Now: it counts, but for at most one of the three.
  2. Was: the skeptic tag needed statements "pre-2017". Now: "2017 or earlier", matching the earliest-statement test (C1).
  3. Was: unclear whether co-written statements count for the 2024-or-later test (C2). Now: they count when the person is a named author.
  4. Daniel Dennett is dropped from the core panel; his statements go to the wide list.
  5. frontier_lab is time-varying and applies if it was true at the time of any statement in the record, with the dates in the justification. philosopher means the person's primary role, so Hinton and Hassabis are not tagged. Chalmers is non_us_uk. Suleyman is not a skeptic.
  6. For the 2020s cohort, the three-national-pieces test is information only, not a condition.
  7. Archive checks for prominence cases blocked by unreachable news sites will run later through the Guardian and New York Times APIs (application programming interfaces: their official data services) once John has keys. Not attempted now.
  8. The wide list keeps all 184 names; its target is "whatever the rule yields", not 60 to 100.

Codebook

  1. Endorsement is about whether the capability will exist, ever or by the horizon the speaker names. Denying that it exists now is present_capability: disputed; endorsement is then coded from what the speaker says about the future, or unexpressed if they say nothing about it.
  2. "May", "might" or "could" with a date is uncertain, with the timeline fields filled and the hedge kept in timeline_text. accepts requires "will", "expect", "I think", a stated probability of 50% or more, or a bet. Agent interpretation: when a hedge and an acceptance word govern the same date ("I think it may be 20 years or less"), the hedge on the date decides and the code is uncertain.
  3. "Decades away" is accepts with a qualitative timeline.
  4. "Takeoff" and "self-improvement" take the lowest self-improvement level the words support, marked ambiguous, with any speed the speaker gives noted in timeline_text. Agent interpretation: when the words say nothing about who does the improving, the lowest level is si_assist.

Process

  1. John double-codes a blind sample himself (data/double-coding/).
  2. John merges the branch himself.
  3. Licence: Creative Commons Attribution 4.0 (CC BY 4.0) for the data, MIT for the scripts.

Why: these were the open questions found while building the panel table and while recoding the seed under version 2; recoders were applying 9 to 12 inconsistently. Rows recoded under rulings 9 to 12 carry coder: claude-agent-v2.1 and a note naming the ruling.

2026-10-01: John's rulings on the eight calls raised after A14

Decided by: john. The eight calls were listed in docs/status-2026-10-01.md (section "New questions for John").

A15. Firm dates, routes, minimum years, split statements and time-varying tags

  1. Firm, unhedged first-person dates ("I'm still saying 2029") are accepts. Added to the list of what accepts requires.
  2. A hedge and an acceptance word on the same date: the hedge decides (uncertain). Confirmed.
  3. A bare "takeoff" is si_assist. Confirmed.
  4. Was: rejecting a route ("language models are not a road to AGI") kept as rejects. Now: a route claim says nothing about whether the capability will exist, so endorsement is unexpressed, with the route claim kept in claim_summary and notes. Applies to LeCun's four route rows.
  5. Was: "at least a decade" and "at least thirty or fifty years" kept as rejects with a minimum year. Now: they assert arrival with a minimum, so under A14 ruling 11 they are accepts, timeline_type: range with the low bound filled, the high bound left open, and "at least" kept in timeline_text.
  6. "Predict", "almost certainly" and "going to" count as acceptance words. Confirmed.
  7. New ruling 13: one row per capability. When one statement makes distinct claims about distinct target capabilities, it gets one row per capability, with ids suffixed a, b, and the shared quote and source. Applied to marcus-2026-04 (split into a si_assist row and a si_sustained row) and altman-2025-06 (a self-improvement row and a superintelligence row).
  8. frontier_lab is time-varying (A14 ruling 5), so Kurzweil, Chollet, Kai-Fu Lee and Karnofsky are tagged, with the dates in the justification. Kokotajlo is not tagged philosopher, since the tag means primary role.

Why: these were the cases where applying A14 needed a reading John hadn't given. Validator changes that follow: statement ids may carry a one-letter suffix, and a range may leave the high bound open when the row carries the lower_bound tag.

2026-10-02: the newspaper-prominence rule loosened

Decided by: mixed. The seed-phase agent raised the question on 2026-10-01 and offered the looser wording; John chose it on 2026-10-02 when asked in the management session.

A16. What counts as "named in" a national piece

Was: an archive hit counted toward the three national pieces only if the person's name was in the headline or abstract (the agent's rule for the 2026-10-01 archive checks). Now: a piece counts if the person is named anywhere in its text and its headline or abstract shows the piece is about where AI is going. Business news, obituaries and pieces on other subjects do not count, on a reading of the headline and abstract.

Why: the strict rule excluded 1,037 of 1,184 archive hits and moved seven candidates to fail, among them people quoted at length on AI's future (Shanahan in twelve Guardian pieces). Being quoted in the text is what "named in" meant in the original rule.

Effect: Legg, Cowen and Yudkowsky pass (all thin); Shanahan, Bengio and Floridi are unclear, each on one remaining point; Pinker, Sutskever, Wooldridge and Goertzel still fail. Core passes go from 19 to 22. Details and the agent's three judgment calls are in docs/panel-verification-2026-10.md, section "Re-marking under A16".

2026-10-06: John's rulings on the A16 readings, Sutskever, the thin readings and the freeze

Decided by: john (rulings 1, 2, 3 and 5); mixed (ruling 4: recommended by the coordinating session, chosen by John). Sent by John on 2026-10-06 in the management session, in answer to the open points in docs/panel-verification-2026-10.md, and relayed to this branch by the coordinating session.

A17. Panel rulings before the freeze

  1. Legg, Cowen and Yudkowsky stay in the core. John confirmed the verifier's three judgement calls: "Legg named in passing, Cowen on jobs-and-automation pieces, and Yudkowsky on two edge outlets". They stay marked thin.
  2. Sutskever: "Not a timeline statement". His co-signed 2015 "hard to predict when" is not an AGI timeline statement from before 2018, so he meets the cohort's first test and passes the 2020s cohort. His 2024-or-later statement was checked against the record: the 2025-11-25 Dwarkesh Patel interview, "I think like 5 to 20" (years), in his own words. Correction recorded: John's first answer the same day said the 2015 line does count as a timeline statement and that Sutskever passes. That answer was given to a question the coordinating session had framed backwards; read literally, it would have failed him (as the agent flagged, the same reason failed Suleyman and Karnofsky). This ruling is John's answer to the question framed correctly, and replaces the first.
  3. Freeze after the three gaps are closed. Don't freeze now. The three unclear core cases (Shanahan, Bengio, Floridi) were closed the same day; all three pass thin.
  4. Both thin readings count. Shanahan's 2025 podcast remark that it is hard to tell "whether we really are on the road to producing general intelligence that is comparable to human general intelligence" counts as his 2024-or-later statement on whether AGI arrives. Floridi's 2017 Financial Times column on robot rights counts as a piece about where AI is going under A16. Both stay marked thin; the core stays at 25.
  5. "Search first, then freeze". Before the freeze, search for one to three more 2020s cohort members at no cost, with the same method and the rules exactly as written, drawing on the wide list and other prominent voices of the 2020s, preferring people who broaden the cohort. Done the same day: Liang Wenfeng, Arthur Mensch and Arvind Narayanan pass (all thin), taking the cohort to 12. The freeze was drafted as A18 and first marked PROPOSED; John then said "Freeze and merge" (A18, below).

Why: these were the last open points before the freeze. Effect: core passes 22 to 25; cohort 8 to 12; the three new cohort members leave the wide list (308 names). Details in docs/panel-verification-2026-10.md, sections "Closing the three unclear cases (2026-10-06)", "Cohort search (2026-10-06)" and "Freeze candidates as of 2026-10-06"; evidence in docs/panel-evidence-2026-10/gaps-2026-10-06.md and docs/panel-evidence-2026-10/cohort-search-2026-10-06.md.

2026-10-06: the expert panel is frozen

Decided by: john ("Freeze and merge", 2026-10-06). The three readings under "Readings chosen at the freeze" are decided by: mixed (recommended by Claude, chosen by John). Drafted by claude-agent on 2026-10-06 under A17 rulings 3 and 5 and first recorded as a proposal; John's yes was relayed to this branch by the coordinating session the same day. The tiers and quota tags are written into data/speakers.csv.

A18. The panel is frozen

Was: candidates checked in docs/panel-verification-2026-10.md; no one frozen. Now: the three tiers below are fixed for full collection. Any later addition, removal or tag change needs its own amendment entry (methodology-v2 section 3).

Core panel (25). Kurzweil, Bostrom, Hanson, Hinton, LeCun, Hassabis, Altman, Marcus, Ng, Russell, Brooks, Musk, Chalmers, Brynjolfsson, Schmidhuber, Kai-Fu Lee, Walsh, Tegmark, Suleyman, Legg, Cowen, Yudkowsky, Shanahan, Bengio, Floridi. Fifteen of them pass on thin evidence, named in the verification table; thin passes are frozen as thin and the reason travels with the row.

Quota tags (minimums met, all four):

quota minimum count who
skeptic 6 8 Hanson, Hinton, LeCun, Marcus, Ng (medium confidence), Brooks, Kai-Fu Lee, Floridi
non_us_uk 3 6 Hinton, Chalmers, Schmidhuber, Kai-Fu Lee, Walsh, Bengio
philosopher 2 4 Bostrom, Marcus, Chalmers, Floridi
frontier_lab 3 11 Hinton, LeCun, Hassabis, Altman, Ng, Musk, Suleyman, Kurzweil, Kai-Fu Lee, Legg, Shanahan

2020s cohort (12). Amodei, Aschenbrenner, Kokotajlo, Cotra, Huang, Leike, Mollick, Chollet, Sutskever, Liang Wenfeng, Arthur Mensch, Arvind Narayanan. At the top of the ten-to-twelve range. Six pass thin (Huang, Leike, Mollick, Liang, Mensch, Narayanan). Tags: frontier_lab 5 (Amodei, Leike, Chollet, Sutskever, Liang); non_us_uk 2 (Liang, Mensch). No cohort member can carry skeptic as defined, since the tag rests on statements from 2017 or earlier; Mensch and Narayanan are skeptics in outlook and that is recorded in their rows.

Wide list (308). Every name in data/wide_list_candidates.csv as of 2026-10-06: the MIRI and AI Impacts predictions dataset and the full TIME100 AI lists 2023 to 2025, less core and cohort members, plus Dennett (A14 ruling 4). Names quoted on AGI timelines in the stream M media sample are added by rule when that stream runs, without a new amendment, because the rule already names that source.

Not on the panel: Pinker, Wooldridge, Goertzel, Mitchell, Etzioni, Davis, Pearl and Karnofsky (each fails a core test as searched); Dennett (dropped, A14 ruling 4; on the wide list); from the cohort search, Sayash Kapoor (meets the rule on the essay co-written with Narayanan, not added: no independent statement, and the cohort would exceed 12), Abeba Birhane and Katja Grace (near misses; on the wide list). A later find that would pass one of them is recorded, not acted on, unless John amends the panel.

Readings chosen at the freeze (decided by: mixed, recommended by Claude, chosen by John, 2026-10-06):

  1. frontier_lab: Liang yes, Mensch no. DeepSeek counts as a company building the most capable models of its time; Mensch's research-scientist role at DeepMind (2020 to 2023) and Mistral do not meet the tag.
  2. Mensch's April 2024 "I don't believe in AGI" counts as his 2024-or-later statement, consistent with the Shanahan reading in A17 ruling 4 (a statement on whether AGI arrives, with no date, counts).
  3. Narayanan, not Kapoor, for the essay they wrote together. Kapoor stays on the wide list.

Why: A17 rulings 3 and 5 asked for the freeze once the three gaps were closed and the cohort search was done; both were done, and John said "Freeze and merge".

Effect: data/speakers.csv gains 32 rows (20 core, 12 cohort) beside the five seed speakers, each with tier, quota tags (justification, date, link) and an inclusion note naming the evidence and any reason the pass is thin. The wide list's frozen record is data/wide_list_candidates.csv; methodology-v2 section 3 puts tier: wide on the light-capture rows, so wide-list people enter data/speakers.csv when their statement is collected, not before (many are listed by surname only in the MIRI dataset, so a speaker row could not yet be filled in honestly). Next step: full collection for the 37 core and cohort members.

2026-10-08: the singularity-entry definition adopted, with a fourth leg

Decided by: john, on four questions the agent put to him with recommendations; John followed the recommendation on the anchor, the thresholds and the backward test, and overruled it on the economic leg. Full text: docs/singularity-entry-definition.md; frozen values: docs/singularity-entry-thresholds.json.

A19. Singularity entry: anchor, thresholds, leg D, and publication

Was: a draft (docs/singularity-entry-definition-draft-2026-10-07.md) with three legs (pace accelerating; forecasts failing; AI doing the improving), placeholder thresholds and no gauge published. Now:

  1. Anchor. METR's 50% task-completion time horizon anchors legs A and B. A successor must be named before it saturates, with the overlap published.
  2. Thresholds frozen as the draft wrote them: doubling time at or below one month; forecast horizon at or below one year; AI share of frontier research work at or above one half; reference year 2019; four consecutive quarters to count as entered.
  3. Backward test owed, not met. The definition must return "not entered" for 1995 to 2019; the series to run it do not exist yet. Adopted anyway, with the gap stated on the page. If the test later fires on those years, the thresholds are amended.
  4. Leg D, the economy, added (John's ruling against the agent's recommendation). World output doubling time from World Bank annual world growth. Threshold: growth at or above 30 percent a year, Davidson's published "explosive growth" line (2021). Reference: the 1961 to 2019 mean of the same series. One annual reading covers its four quarters; revisions reset the count. The agent had recommended publishing the economic reading beside the mark, not inside it, because output lags capability by years and the progress number (the smallest gauge) will track the economy until it moves. John accepted that cost for the sake of a leg measured outside the AI field.
  5. Gauges and status are published at agizeitgeist.com from the frozen file. Leg B is shown as not yet measurable until the forecast-versus-outcome ledger exists; progress is then reported as "at most" the smallest measured gauge.

Agent decisions inside the adoption (John can reverse any):

Rendered from docs/amendments.md at commit 1c1b2cb.