A dated record and a scored test

AGI Zeitgeist Tracker

The spirit of the age about artificial general intelligence, as a dated and sourced record since 2005, beside a test, with its rules fixed in advance, of whether the world has entered the singularity. Every number on this page links to the row it came from.

AGI here means AI that can do most intellectual work a person can. The working question: did the main public argument move from whether it is possible, to when it arrives, to whether today's systems are already close or improving themselves? The method is built so the data can say no. The seed phase is complete and the expert panel is frozen; the full panel traces and the public streams have not started.

Singularity entry status: Not entered. Scored on 2026-10-08 under the definition adopted 2026-10-08, with its thresholds frozen before any gauge was published. Progress toward the mark is at most 0.00 on a scale of 0 to 1, set by the weakest leg (D). Leg B cannot be scored until its forecast ledger exists, so the number is an upper bound. Entry requires all four gauges at 1 for 4 consecutive quarters.

What people believe

The tracker proper. One row per public statement, survey result, market reading or poll, coded under methodology version 2.

Share of coded statements by primary frame, by year, seed panel of five people
Seed phase: what each statement was mainly about, by year, for five traced people (Kurzweil, Hinton, LeCun, Marcus, Altman). Counts in chart-1-frames.csv. Five people are not a panel; this chart exists to test the pipeline.
Framing in sampled high-score posts from r/singularity, 2024Q4 to 2026Q3, model-coded
Pilot, model-coded: framing in sampled high-score posts from r/singularity, 2024Q4 to 2026Q3. Not representative of participants or the public. Three models from different companies (Claude, Gemini and GPT) each coded all 390 sampled posts blind under the final codebook, and the majority vote is the final code; where no two agree, the post is left out of the affected measure. Counts in vanguard-pilot-r-singularity.csv; pilot report.

Entering the singularity: the test and the score

"Any prediction made now about the next five years will look stupid in retrospect" is one of the standard definitions of the singularity, and each definition alone has a known false alarm. The adopted definition requires four things to be true at once, for four consecutive quarters. Each gauge runs from 0 (the 2019 reference) to 1 (the frozen threshold); the smallest gauge is the honest distance remaining.

A. The pace is accelerating
Kurzweil's school. The doubling time of a frozen capability measure is itself shrinking.
121 days
per doubling, trailing twelve months to 2026-02-05. Threshold: 30 days and still falling. Reference: 188 days.
0.43
B. Forecasts have stopped working
Vinge's school. The horizon over which forecasts beat the trend is collapsing. Confirmed only in arrears.
+6.9 yr/yr
stand-in only: excess pull-in of the expert consensus date, 2022 to 2023-10-01. The gauge needs a forecast ledger that is not built yet. Threshold: a one-year horizon.
not yet measurable
C. AI is doing the improving
Good's school. The share of the work that produces frontier progress that is supplied by AI.
26%
of Anthropic's AI research work "led" by its own model, August 2026, by its own measurement. Threshold: one half. One company, self-reported.
0.52
D. The economy has changed mode
Hanson's school. A new growth mode doubles world output a hundred times faster than the last.
2.9% a year
world output growth in 2025, doubling every 24 years. Threshold: 30% a year (2.6-year doubling). Reference: 3.5% (20-year doubling).
0.00

Leg A. How long a task the best AI can finish

The anchor measure is METR's time horizon: the length of software task, in minutes an expert needs for it, that a model finishes half the time. It is chosen because it has not yet saturated the way benchmarks do. The lower panel is the definition's quantity: the doubling time of the frontier over a trailing twelve months. A flat line would mean steady exponential growth, which is Kurzweil's world but not yet a singularity; a falling line is the signature the definition looks for. The gauge is scored on the last window with no reading past 16 hours, and the threshold also requires the doubling time to have fallen over the last two windows (not met today).

METR 50% time horizon by model release date on a log scale, with the trailing twelve-month doubling time of the frontier below
Hollow points were not the frontier at release. METR treats estimates past 16 hours as beyond its task suite's reliable range, and its own doubling-time fit excludes them; windows that include such a reading are hollow in the lower panel and are not scored. The latest frontier reading is Claude Mythos (early preview) at 17 hours on 2026-04-07. Doubling-time series in leg-a-doubling-time.csv.
Every reading (26 models) and every doubling-time window
releasemodel50% horizon95% rangefrontierrow
2019-02-14GPT-20.1 min0.0 to 0.1 minyesmetr-th11-gpt2
2020-05-28davinci-0020.1 min0.1 to 0.2 minyesmetr-th11-davinci-002
2022-03-15GPT-3.5 Turbo Instruct0.6 min0.3 to 1.1 minyesmetr-th11-gpt-3-5-turbo-instruct
2023-03-14GPT-4 (0314)4.0 min1.9 to 8.0 minyesmetr-th11-gpt-4
2023-11-06GPT-4 (1106)4.0 min1.9 to 8.4 minyesmetr-th11-gpt-4-1106
2024-03-04Claude 3 Opus4.0 min1.7 to 8.8 minmetr-th11-claude-3-opus
2024-04-09GPT-4 Turbo3.7 min2.0 to 6.7 minmetr-th11-gpt-4-turbo
2024-05-13GPT-4o7.0 min4.0 to 12.9 minyesmetr-th11-gpt-4o
2024-06-20Claude 3.5 Sonnet (June 2024)11.4 min5.5 to 22.4 minyesmetr-th11-claude-3-5-sonnet-20240620
2024-09-12o1-preview20.3 min11.7 to 33.4 minyesmetr-th11-o1-preview
2024-10-22Claude 3.5 Sonnet (October 2024)20.5 min10.1 to 40.8 minyesmetr-th11-claude-3-5-sonnet-20241022
2024-12-05o138.8 min0.4 to 1.1 hyesmetr-th11-o1
2025-02-24Claude 3.7 Sonnet1.0 h0.6 to 1.7 hyesmetr-th11-claude-3-7-sonnet
2025-04-16o32.0 h1.2 to 3.2 hyesmetr-th11-o3
2025-05-22Claude Opus 41.7 h1.0 to 2.7 hmetr-th11-claude-4-opus
2025-08-05Claude Opus 4.11.7 h1.0 to 2.7 hmetr-th11-claude-4-1-opus
2025-08-07GPT-53.4 h1.9 to 6.8 hyesmetr-th11-gpt-5-2025-08-07
2025-11-18Gemini 3 Pro3.7 h2.3 to 6.3 hyesmetr-th11-gemini-3-pro
2025-11-19GPT-5.1-Codex-Max3.7 h2.2 to 6.6 hmetr-th11-gpt-5-1-codex-max
2025-11-24Claude Opus 4.54.9 h2.7 to 10.4 hyesmetr-th11-claude-opus-4-5
2025-12-11GPT-5.25.9 h3.3 to 13.6 hyesmetr-th11-gpt-5-2
2026-02-05Claude Opus 4.612.0 h5.3 to 60.6 hyesmetr-th11-claude-opus-4-6
2026-02-05GPT-5.3-Codex5.8 h3.2 to 13.6 hmetr-th11-gpt-5-3-codex
2026-02-19Gemini 3.1 Pro6.4 h3.9 to 11.6 hmetr-th11-gemini-3-1-pro
2026-03-05GPT-5.45.7 h3.1 to 12.8 hmetr-th11-gpt-5-4
2026-04-07Claude Mythos (early preview)17.4 h8.5 to 55.1 hyespast 16 hmetr-th11-claude-mythos-preview-early
window enddays per doublingfrontier pointsincludes a reading past 16 h
2024-06-201723
2024-09-121384
2024-10-221405
2024-12-05935
2025-02-24956
2025-04-16897
2025-08-07916
2025-11-181315
2025-11-241316
2025-12-111396
2026-02-051217
2026-04-071177yes

Leg B. Are forecasts holding up?

The real test scores old forecasts of the measure above against simple trend extrapolation and asks how far ahead they still win. That ledger does not exist yet and can only be scored once the forecast period has passed, so leg B is shown as not yet measurable and the progress number above is an upper bound. The stand-in, available now, is the expert consensus date for human-level AI. If forecasters had the trend right, the date would stand still and the years remaining would fall by one each year, the dashed line. Instead the three AI Impacts surveys show the remaining distance collapsing: the step from 2022 to 2023 removed about 14 years of distance in 1.7 calendar years. That is forecasts failing in the slow direction, which is the direction the definition records.

Years remaining until the surveyed 50% date for human-level AI, by survey year, against a one-year-per-year reference line
Surveys use different definitions and aggregation, so the 2013 survey is shown but not joined to the series. Steps in leg-b-pull-in.csv.
The survey readings, the steps between them, and the market reading for context
survey50% yearyears remainingrow
2016-01-01206145.0aiimpacts-2016-hlmi50
2022-01-01205937.0aiimpacts-2022-hlmi50
2023-10-01204723.3aiimpacts-2023-hlmi50
stepcalendar yearschange in years remainingexcess pull-in per year
2016 to 2022-01-016.0-8.0+0.3
2022 to 2023-10-011.7-13.7+6.9

Manifold market, "AGI created and publicly announced before 2030", for context only (a market price, not a survey):

dateprobabilityrow
2024-01-0736%manifold-agi2030-2024-01-07
2024-12-0747%manifold-agi2030-2024-12-07
2025-12-0736%manifold-agi2030-2025-12-07
2026-09-3053%manifold-agi2030-2026-09-30

Leg C. Is AI doing the improving?

The definition wants the share of inputs to frontier AI progress supplied by AI. Nobody outside the labs can measure that, so what exists is self-report. Two kinds are recorded: a company's share of all new code written by AI, which is a proxy at best and is shown but not scored, and Anthropic's self-measured share of its own AI research work that its model "leads", which is closer to the definition's measure but is judged by the model itself and is the figure scored. The figures do not measure the same thing and should not be compared with each other.

Bar chart of company self-reported shares of AI-done work, with the one-half threshold marked
Every bar is a company statement about itself. Anthropic's row for February 2026 is "under 1%", recorded as 1. Its "above 90%" row is recorded as 90.
The self-reports
datemeasurevaluerow
2026-02-01share of Anthropic's AI research and development work that Claude 'leads' (R&D Automation Index, prototype)1% (upper bound)anthropic_institute-2026-02-rd-leads
2026-04-24share of Google's new code that is AI-generated and approved by engineers (company self-report)75%google-2026-04-code-share
2026-08-01share of Anthropic's AI research and development work that Claude 'leads' (R&D Automation Index, prototype)26%anthropic_institute-2026-08-rd-leads
2026-08-01share of Anthropic's AI research and development work at or above the level where Claude 'collaborates' (R&D Automation Index, prototype)90% (lower bound)anthropic_institute-2026-08-rd-collaborates

Leg D. Has the economy changed mode?

Robin Hanson's reading of the long record is that the world economy has had three growth modes, hunting, farming and industry, each doubling output about a hundred times faster than the one before: roughly 230,000 years, 860 years and 15 years per doubling. A new mode would double output in years or less. Leg D asks whether that has started, using world real output growth from the World Bank. The threshold, 30 percent a year, is Tom Davidson's published line for "explosive growth" (Open Philanthropy, 2021), chosen so the number is not this project's own invention: far above anything in the modern record, well short of Hanson's "days, not years". This leg is measured by people outside the AI field, which is why it is in the definition; the cost is that output lags capability by years, so it will usually be the weakest leg.

World output growth by year since 1961, with the 30 percent threshold line far above the series
Annual readings; the series is revised as countries revise their accounts, and a revision that pulls a year back below the line resets the persistence count. Doubling times in leg-d-world-growth.csv.
Every year (65 readings)
yeargrowthdoubling time (years)row
20252.92%24.1worldbank-gdp-growth-2025
20242.90%24.2worldbank-gdp-growth-2024
20232.86%24.6worldbank-gdp-growth-2023
20223.44%20.5worldbank-gdp-growth-2022
20216.49%11.0worldbank-gdp-growth-2021
2020-2.89%none (output fell)worldbank-gdp-growth-2020
20192.68%26.2worldbank-gdp-growth-2019
20183.29%21.4worldbank-gdp-growth-2018
20173.45%20.4worldbank-gdp-growth-2017
20162.79%25.2worldbank-gdp-growth-2016
20153.13%22.5worldbank-gdp-growth-2015
20143.17%22.2worldbank-gdp-growth-2014
20132.89%24.3worldbank-gdp-growth-2013
20122.75%25.6worldbank-gdp-growth-2012
20113.31%21.3worldbank-gdp-growth-2011
20104.52%15.7worldbank-gdp-growth-2010
2009-1.33%none (output fell)worldbank-gdp-growth-2009
20082.09%33.5worldbank-gdp-growth-2008
20074.43%16.0worldbank-gdp-growth-2007
20064.49%15.8worldbank-gdp-growth-2006
20054.06%17.4worldbank-gdp-growth-2005
20044.50%15.7worldbank-gdp-growth-2004
20033.07%22.9worldbank-gdp-growth-2003
20022.32%30.2worldbank-gdp-growth-2002
20012.03%34.5worldbank-gdp-growth-2001
20004.57%15.5worldbank-gdp-growth-2000
19993.58%19.7worldbank-gdp-growth-1999
19982.78%25.3worldbank-gdp-growth-1998
19974.00%17.7worldbank-gdp-growth-1997
19963.59%19.7worldbank-gdp-growth-1996
19953.18%22.1worldbank-gdp-growth-1995
19943.42%20.6worldbank-gdp-growth-1994
19931.86%37.6worldbank-gdp-growth-1993
19922.06%34.0worldbank-gdp-growth-1992
19911.23%56.7worldbank-gdp-growth-1991
19902.71%25.9worldbank-gdp-growth-1990
19893.64%19.4worldbank-gdp-growth-1989
19884.53%15.6worldbank-gdp-growth-1988
19873.75%18.8worldbank-gdp-growth-1987
19863.31%21.3worldbank-gdp-growth-1986
19853.66%19.3worldbank-gdp-growth-1985
19844.73%15.0worldbank-gdp-growth-1984
19832.58%27.2worldbank-gdp-growth-1983
19820.41%169.4worldbank-gdp-growth-1982
19811.89%37.0worldbank-gdp-growth-1981
19801.82%38.4worldbank-gdp-growth-1980
19794.14%17.1worldbank-gdp-growth-1979
19784.14%17.1worldbank-gdp-growth-1978
19773.94%17.9worldbank-gdp-growth-1977
19765.22%13.6worldbank-gdp-growth-1976
19750.81%85.9worldbank-gdp-growth-1975
19741.99%35.2worldbank-gdp-growth-1974
19736.46%11.1worldbank-gdp-growth-1973
19725.54%12.9worldbank-gdp-growth-1972
19714.11%17.2worldbank-gdp-growth-1971
19703.77%18.7worldbank-gdp-growth-1970
19695.96%12.0worldbank-gdp-growth-1969
19685.95%12.0worldbank-gdp-growth-1968
19673.74%18.9worldbank-gdp-growth-1967
19665.42%13.1worldbank-gdp-growth-1966
19655.63%12.7worldbank-gdp-growth-1965
19646.63%10.8worldbank-gdp-growth-1964
19635.02%14.2worldbank-gdp-growth-1963
19625.32%13.4worldbank-gdp-growth-1962
19613.91%18.1worldbank-gdp-growth-1961

What the definition still owes

  1. The backward test. Run on 1995 to 2019 the definition must return "not entered", with the internet years firing only leg B and the deep-learning years firing only leg A. The leg A series starts in 2019 and leg C in 2026, so the test cannot run yet. The definition was adopted without it, with this gap stated here; if the test later fires on those years, the thresholds are amended in the open.
  2. The leg B ledger. Old forecasts of the anchor measure (starting from the 2016 expert survey's milestone dates) scored against what happened, by lead time. Until it exists, leg B is unmeasured.
  3. A second source for leg C that is not a company's report on itself.

Method and data

Belief measures and world measures live in separate files and are never combined. statements.csv, aggregates.csv and the community samples hold what people said and believed. world_measures.csv holds capability readings, company self-reports and world output, with the same columns and the same rule that every row carries its source and the question wording. Every search that produced a row is in search_log.csv. The method, the codebook, the definition and every dated change are published here; the scripts that rebuild this page from the data are in the source repository.

Cite as: Fredrickson, J. (2026). AGI Zeitgeist Tracker [dataset]. Version of 2026-10-08, commit c63f3d0. https://agizeitgeist.com. Data under CC BY 4.0; scripts under MIT. Figures are static images; the CSV beside each one holds the plotted numbers. Thresholds frozen 2026-10-08 in amendment A19.