Epoch AI’s cover photo
Epoch AI

Epoch AI

Research Services

San Francisco, California 12,471 followers

Research institute investigating the trajectory of AI

About us

Epoch AI is a multidisciplinary research institute investigating the trajectory of Artificial Intelligence (AI). We scrutinize the driving forces behind AI and forecast its ramifications on the economy and society. We emphasize making our research accessible through our reports, models and visualizations to help ground the discussion of AI on a solid empirical footing. Our goal is to create a healthy scientific environment, where claims about AI are discussed with the rigor they merit.

Website
https://epoch.ai
Industry
Research Services
Company size
11-50 employees
Headquarters
San Francisco, California
Type
Nonprofit
Founded
2022
Specialties
AI Governance, AI Forecasting, and Consultations

Locations

  • Primary

    28 Geary St

    Ste 650 #1917

    San Francisco, California 94108, US

    Get directions

Employees at Epoch AI

Updates

  • We estimate AI infrastructure could soon support hundreds of millions or even billions of AI agents. At the high end, that’s enough to rival the working hours of the global human workforce. The exact number of agents depends on the efficiency of the models they run on and whether soaring demand for AI continues. We analyzed different scenarios To estimate the size of the agent population, we first determine what an hour of continuous agent work costs, comparable to human wages. In the agent sessions we analyzed, Codex workloads averaged ~$16–$18 per agent-hour at API prices, compared with ~$24–$50 for Claude Code. Next, we use 2025–27 HBM shipments to estimate available hardware for running agents. HBM gives a common basis for comparing capacity across different vendors and chips. How many agents that hardware can support depends on the model and workload. Future agents could use more compute for harder tasks, while more efficient models could support far more agents on the same hardware. Using all this capacity would require a massive increase in global demand for AI. Consider the relatively conservative scenario of 20% utilization of total capacity. Under our reference assumptions, that would correspond to $2.6–$5.3T in annual spending. Even if leading labs' revenues keep growing 5×/year, their combined annualized revenue would reach only ~$1T by the end of 2027. The AI hardware buildout is a bet that people will pay for AI to perform valuable tasks throughout the economy, at hourly costs roughly comparable to human wages. Analysis by Jason Li. Full report and code: https://lnkd.in/e3BcYdks

    • Chart: potential concurrent agents from 2025–27 memory shipments by model, from 30–60 million (Claude Fable 5) and 97–195 million (GPT-5.6 Sol) to 240 million (Kimi K3), 571 million (GLM-5.2) and 1.9 billion (DeepSeek V4 Pro).
    • Chart: potential concurrent agents versus reference serving cost per agent-hour; lower serving costs allow the same hardware to support more agents, from 1.9B (DeepSeek V4 Pro) down to 30–60M (Claude Fable 5).
  • We’ve collected metadata on the usage patterns of 5,000 ChatGPT users and made it available in our new ChatGPT usage explorer. The histories span Dec 2022–Dec 2025 and show the cohort using ChatGPT more intensively over time. More on this below. Metadata was provided by @YouGov, which randomly selected ChatGPT users from its US panel. All message text was removed before sharing. Two key caveats for interpretation: users opted in to share their histories, and had used ChatGPT at least once in 2026, when the sample was drawn. - Frequency of use increased. Among active users in our dataset, the share using ChatGPT 21+ days/month quadrupled in two years, rising from 2.6% in Dec 2023 to 10.5% in Dec 2025. - Message volume grew too. The median active panelist sent or received 14 messages in Jan 2023, and 36 in Dec 2025. The mean is much higher, at 158, because a small group of heavy users sends a large share of all messages. - The average conversation got longer, although the typical one didn't. The median conversation had two prompts in all but one month, while the mean rose from ~4 to 6. These trends come from a self-selected panel of users. The data has not been weighted to match the demographics of US adults or of ChatGPT users as a whole. Read our full write-up: https://lnkd.in/eDq4xD3S Explore the data yourself: https://lnkd.in/eMrzcnCU

    • Stacked area chart of active ChatGPT panelists by days active per month, Dec 2022 to Dec 2025, showing the share active on 21 or more days quadrupling while the share active on just one day shrank.
  • Claude Opus 5.5 has taken the top spot on the Epoch Capabilities Index (ECI) with a score of 167, narrowly ahead of GPT-6 Astra. Claude Sonnet 5.5 has roughly matched Claude Fable 5.1 (165). For both Opus and Sonnet 5.5, the gap between ECI and SWE ECI is minimal, less than one ECI point. In comparison, GPT-6 Astra’s SWE ECI is around two ECI points lower than its general ECI. Compare ECI across models in our explorer: https://lnkd.in/epWxkjhT

    • No alternative text description for this image
    • No alternative text description for this image
  • Can AI tell if you've built your IKEA furniture wrong? Our new benchmark, the Furniture Assembly Benchmark (FAB), gives models the manual and a photo of a half-completed piece of furniture and asks them to spot the mistake. The top score has gone from 28% to 80% in just 10 months. On FAB, AI must assess builds based on photos like the one shown below. The benchmark consists of 60 photos from 3 different furniture builds. There are two ways to get a question wrong: miss a real mistake, or flag a build that's fine. Many recent frontier models favor one failure mode over another. Gemini 3.1 Pro, for example, passes almost no correct builds, whereas GPT-5.4 catches almost no mistakes. For practical utility, speed matters as well as accuracy. Here GPT-6 Astra is a clear outlier. Not only is it the most accurate, but, of all the models we tested, it is also the fastest. We think FAB is a good proxy for a number of economically important tasks that require similar visual reasoning, like fixing a car or repairing household appliances. We hope that benchmarks like this can help better track AI’s progress in this domain. Read our full write-up of the results: https://lnkd.in/ei9zNUT7

    • No alternative text description for this image
    • No alternative text description for this image
    • No alternative text description for this image
    • No alternative text description for this image
  • AI is getting cheaper more quickly than any other transformative tech in history. At a given level of performance, cost has fallen ~47%/quarter since 2023. That’s 4× faster than DNA sequencing, 6× faster than compute, 18× faster than lithium batteries, and (up to 1973) 54× faster than electricity. For one example of the plunging “cost of thought,” take OpenAI’s o3 model. Released in January 2025, it was the first LLM to score >25% on FrontierMath Tiers 1-3. The bill for achieving this mark: $0.55. This summer, GPT-5.6 Luna did the same for $0.0015. A 377-fold drop in 18 months. Our estimates come from an expanded Epoch dataset on the performance of 222 AI models across up to 11 benchmarks. The dataset is built using a method borrowed from CAISI, which lets us estimate how the same model would perform under various budget constraints. The speed of price drop varies by domain: slower on game-based puzzles, at 39–43% per quarter, faster on math problems, at 50–52% per quarter. Costs for a given level of performance often fall the fastest when that level is state-of-the-art. Averaging across five benchmarks, cost falls 66% per quarter for SOTA performance. Two years later, the price drop slows to 32% per quarter. The trillion-dollar question: If AI companies collectively slow the development of AI, will prices also fall more slowly? Read the full report: https://lnkd.in/eeCrBKhE

    • No alternative text description for this image
  • The share of math preprints on arXiv that acknowledge AI has risen rapidly, from 4% in April to 25% in August. This increase holds even when filtering to papers with at least one author who published regularly before 2023. AI is acknowledged for many things, including literature review and coding, but it also generates substantial research ideas, with 6% of papers in August acknowledging this kind of contribution. Explore the data yourself on our new AI Use in Math Research explorer: https://lnkd.in/e2Qv7AQN Learn more about this Data Insight: https://lnkd.in/ea5w3fWq

    • No alternative text description for this image
  • Another problem from FrontierMath: Open Problems has been solved! The solution was elicited by Becker, Greger, and Peters in an interactive session with GPT-6 Astra. Peters originally suggested the problem for the benchmark. He had this to say. We’ve marked this problem as solved by humans + AI, reflecting that humans were actively engaged in elicitation but that the core ideas came from AI. For analyses where a binary human vs. AI distinction is required, we suggest considering these to be AI solutions. This is the first problem to be solved that we marked as a “Major Advance”. Our standards for this are given below, and in the FAQ on our website. Learn more about FrontierMath: Open Problems on our website! https://lnkd.in/d8Rgd6Pt

    • Quote card from Dominik Peters (CNRS Research Scientist, Université Paris Dauphine - PSL) on GPT-6 Astra designing a new core-stable multi-winner voting rule based on maximizing harmonic entropy.
    • No alternative text description for this image
  • Introducing Benchmark Reviews: our new initiative to audit AI benchmarks. We are launching with 15 benchmarks: 4 Verified, 9 Flawed, and 2 with not enough information for a review. Benchmarks assess AI capabilities, but the benchmarks themselves vary substantially in quality. We hope to be a consistent source of information on benchmark quality. See our reviews here: https://lnkd.in/e9PYSh6t We assign each benchmark a verdict based on our rubric: Flawed, Verified, or Not Enough Info. To avoid conflicts of interest we do not review Epoch-created benchmarks, but welcome external reviews. https://lnkd.in/eRTgaHzN Verified benchmarks can broadly be interpreted as described, and any errors that exist do not substantially affect the results. Alongside any verified benchmark we publish a full review and assessment of the benchmark, including any weaknesses and limitations we think it has. Flawed benchmarks have one or more substantive flaws we believe users need to be aware of to accurately interpret results, most commonly that >20% of the tasks have accuracy-impacting errors. In this case, we publish a limited writeup of the flaws we found. If we aren’t able to access enough information to review a benchmark, we’ll designate it ‘Not Enough Info.’ We will try to work with the creators of private benchmarks to conduct reviews while keeping the questions/tasks outside of public knowledge. Each review is of a specific benchmark version, and if a benchmark is updated to fix errors we may review the new version. We share findings with developers and will publish a link to their response, if they’d like. Read our full methodology here: https://lnkd.in/er7eh_vi While we are launching with 15 benchmarks, we plan to continue reviewing benchmarks on an ongoing basis, prioritizing those with the most impact and reach, as well as highlighting benchmarks we think are high quality that might be overlooked.

    • No alternative text description for this image
  • Trade data is consistent with more than $3B of chips smuggled into China via Malaysia. Between April 2024 and June 2025, China recorded $3.8 billion in server imports from Malaysia, averaging $106,000 each. Malaysia declared the same shipments at $17,000 each. The two countries agree on the number of machines traded, but they disagree by 6× on their value. Furthermore, other countries importing Malaysian-origin servers over the same period paid 5 to 70 times less per unit. This flow began a few months after the US banned exports of Nvidia’s A800/H800 GPUs to China (October 2023) and then collapsed within weeks of Malaysia requiring a permit to export or transship high-performance AI chips (July 2025). This doesn’t prove diversion, but it independently matches established cases of chip smuggling into China, where intermediaries are alleged to have routed AI servers through Malaysia. Learn more about the methodology behind this Data insight: https://lnkd.in/epgdjF5K For a sense of scale, this trade could represent 150,000 H100e of compute, about a quarter of our earlier modeled estimate of all chips smuggled into China through 2025. Read our previous full report on our modeled estimates: https://lnkd.in/e-QxWpUm

    • No alternative text description for this image
  • GPT-6 Astra leads the Epoch Capabilities Index (ECI), ahead of competing models like Claude Fable 5.1, and its Math-ECI also sets a new record. However, on software engineering benchmarks, Fable 5.1 remains state-of-the-art. Math-ECI and SWE-ECI are domain-specific scores, refit on only one domain’s benchmarks. They measure a model’s abilities in one single domain relative to other models. Based on pre-release evals, we gave GPT-6 Astra an overall ECI of 169 at launch. Only one of its benchmark results at the time was in software engineering (MirrorCode). But since then, three more SWE results have since been added, and its ECI has subsequently fallen to 166. Note that each domain-specific ECI is calculated from 3 to 6 benchmark results, so our confidence intervals are wide: GPT-6 Astra’s SWE-ECI of 164 (90% CI: 160 to 170) overlaps with Fable 5.1’s 167 (90% CI: 162 to 179). GPT-5.6 Sol (ECI 162) and Kimi K3 (158) are shown for comparison. These results come from our Domain-specific ECI Explorer, where you can view Math and SWE ECIs or design your own variant: https://epoch.ai/eci Full data and methodology in our latest Data Insight: https://lnkd.in/e85Tdgza

    • No alternative text description for this image

Similar pages

Browse jobs