Document

Codebook, version 2

Codebook, version 2

How each coded unit is coded. Written 2026-09-30 and revised 2026-10-01 (amendment A14) from docs/methodology-v2.md section 2, which is the authority; this file turns its rules into decisions with examples. Changes go in docs/amendments.md first.

Version 1 (the single stance code and the six-value question_frame) is retired. Rows coded under version 1 were recoded from their sources, not mapped (amendment A2).

About the examples. Examples marked constructed are invented to show a rule and must never be cited as quotes.


0. One unit, and how to read it

A unit is one statement by one speaker on one occasion. Code only what the quoted words and their immediate context say. Don't use what the speaker is known for, what they said elsewhere, or what the article around the quote implies.

Read the source, not just the stored quote. The stored quote (60 words or fewer) is evidence for the codes. If a code depends on a sentence outside the stored quote, either change the quote to include it or code unexpressed / ambiguous.


1. target_capability: what capability is the unit about?

code the capability typical words
conversation passes as human in conversation Turing test, can't tell it from a person
narrow_task one specific task or domain chess, driving, radiology, coding contests
broad_competence most economically valuable or intellectual work; "AGI", "human-level AI" when used this way AGI, human-level, do any job a person can
superintelligence exceeds humans broadly superintelligence, smarter than us at everything, the Singularity (when it means AI beyond humans)
si_assist AI helps humans build AI AI speeds up our researchers, copilots for ML engineers
si_research AI does parts of AI research on its own AI runs experiments, writes the training code, automated AI researcher for some tasks
si_modify AI changes its own code or weights rewrites itself, modifies its own weights
si_sustained successive improvements with little human involvement intelligence explosion, recursive self-improvement, "takeoff", each generation builds the next
none no capability is claimed a statement about the word "AGI", about hype, about policy with no capability in view

Choosing among them. Pick the capability the speaker's claim is about. If they name two (for example "human-level by 2029, Singularity by 2045"), code the one the unit leads with and record the other in definition_note.

The four self-improvement levels. Ask two questions in order:

  1. Who does the improving? Humans with AI help → si_assist. The AI itself → go on.
  2. How far does it go?
    • A part of the research pipeline, with humans still steering → si_research.
    • The system edits itself (its own code or weights) → si_modify.
    • Repeated rounds, each better system making the next, with little human role → si_sustained.

"Takeoff" and loose self-improvement words (ruling 12, A14). Code the lowest level the words support and set ambiguous: yes; put any speed the speaker gives ("slow", "fast", "gentle") in timeline_text. When the words don't say who does the improving (a bare "takeoff", "the AI will improve itself"), the lowest level is si_assist. Code si_sustained only when the words themselves describe repeated rounds with little human role ("each generation builds the next", "an intelligence explosion with no one in the loop"). The labels "recursive self-improvement" and "intelligence explosion" alone don't do that.


2. frame_primary and frames_present: which questions does the unit raise?

code the question
possibility Can it be built at all, in principle?
timeline When will it arrive?
capability_now Is a current system already it, or close?
self_improvement Can current or imminent systems improve themselves?
definition What would count?
consequences What happens if or when it arrives (risk, economy, society)?
governance What should be done about it (regulation, pauses, safety work, coordination)?
other_concern A different AI question leads: jobs, reliability, privacy, bias, misinformation, military, consciousness, ownership

Test for possibility against timeline: delete every date from the unit. If the claim still stands and is about whether it can be done, it's possibility.


3. endorsement: does the speaker endorse that the target capability will exist?

Endorsement is about the future: whether the capability will exist, ever or by the horizon the speaker names (ruling 9, A14). Whether it exists now is a different field, present_capability.

code meaning
accepts says it will happen (or has already happened). Needs "will", "expect", "I think", "predict", "going to", "almost certainly", a stated probability of 50% or more, a bet, "decades away", "at least N years", or a firm, unhedged first-person date ("I'm still saying 2029") (A14 rulings 10 and 11; A15 rulings 1, 5 and 6)
rejects says it won't happen, ever or by the horizon given. "It isn't here yet" is not a rejection (ruling 9)
uncertain explicitly says they don't know, gives a probability below 50%, or hedges the date with "may", "might" or "could" (ruling 10)
conditional endorses it only if a stated condition holds ("if scaling continues", "if we solve X")
unexpressed discusses it without taking a position

Silence is never acceptance. Endorsement is coded only from words the speaker actually said. Every code other than unexpressed needs an anchor: the exact words in the stored quote that carry the position, recorded in notes as basis: "<words copied from the quote>". scripts/validate.py checks that the anchor appears in the quote. If you can't point to such words, the code is unexpressed.

Rulings of 2026-10-01 (A14), with examples from the seed data:

Hard cases:

One row per capability (A15 ruling 13)

When one statement makes distinct claims about distinct target capabilities, it gets one row per capability. The rows share the quote, source and date; their ids carry a one-letter suffix (marcus-2026-04a, marcus-2026-04b); each row is coded for its own capability.

Seed example 1, marcus-2026-04: "it's great for example at (some) code optimizations, but not even (as the blog makes clear) actual RSI, which remains speculative." Two claims:

Seed example 2, altman-2025-06: "We are past the event horizon; the takeoff has started. Humanity is close to building digital superintelligence…"

4. present_capability: does it exist now, in the speaker's view?

code meaning
demonstrated the speaker says it already exists and points to it
claimed the speaker says it exists, without pointing to evidence
disputed the speaker says current systems do not have it
hypothetical discussed as a future or possible capability
na no capability in view (target none)

na exactly when target_capability is none.

5. Timeline fields

As in version 1, with three tags (in notes) for claims that aren't arrival dates:

field rule
timeline_type point_year, range, prob_by_year, qualitative, none
timeline_year an absolute year; for "in 10 years" add it to the statement year and record computed: 2023 + 10 in notes
timeline_year_high only for range
timeline_prob 0 to 1, only when the source gives a probability. "Soon" is never converted to a number
timeline_text the speaker's own words for the timing, always filled when a hedge is present

6. Consequences, affect and agency

field codes rule
consequence_type existential, economic, social, mixed, unspecified the kind of consequence the unit expects. unspecified when none is named
consequence_valence beneficial, harmful, mixed, unspecified whether those consequences are good or bad in the speaker's words
affect enthusiasm, fear, resignation, neutral, mixed, unexpressed emotional posture, only from explicit cues ("exciting", "scary", "I'm worried", "nothing we can do"). A plain forecast is unexpressed, not neutral; use neutral only when the speaker signals detachment ("I'm just reporting the trend")
agency can_influence, cannot_influence, unexpressed does the speaker say the outcome can be shaped? "We must regulate" → can_influence; "it's coming whatever we do" → cannot_influence

These are separate on purpose. Someone who expects AGI in five years can be thrilled, frightened, resigned or organizing.

7. elicitation: how did the statement come about?

code meaning
unprompted the speaker chose to say it: their own blog, post, book, essay, talk
interview answering a question in an interview, podcast or hearing
survey answering a survey item
quoted the unit is a claim quoted inside someone else's article, without the full exchange available

Secondhand rows are usually quoted.

8. Other fields

9. Notes tags

Notes start with any of these tags, separated by ; , then free text:

tag meaning
secondhand see section 8
incidental found while searching for a different speaker
sampled kept under the every-third rule for a busy year
computed: <working> how an absolute year was worked out
basis: "<words>" the words carrying the endorsement (required unless unexpressed)
not_by: YYYY "won't happen by YYYY" (section 5)
lower_bound the timeline year is a minimum (section 5)

10. Mapping to Fast and Horvitz 2017

Methodology-v2 section 5 asks for the categories of Fast and Horvitz's 2017 study of New York Times AI coverage to be reconciled with these codes before the media stream starts. That reconciliation belongs to the media stream and has not been done yet.

Rendered from docs/codebook.md at commit 1c1b2cb.