1: AI is advancing fast

AI has existed since the 1950s, but in the 2010s and especially 2020s, “deep learning” systems have made great strides in capability. The most prominent deep learning systems are “large language models” (LLMs) trained on completing human text, such as GPT-5 (which powers ChatGPT). These models end up being able to do a surprisingly wide range of tasks.

::: carousel

  • Recent AI performance on several benchmarks historically considered difficult for AI. You can see the acceleration in AI progress as the curves get steeper and bunch closer together over time. Note that single models can achieve state-of-the-art performance across a wide range of tasks, without a separate model being required for each. (Normalized with initial level at -100, human level at 0.) Source: Kiela et al. (2023), with minor processing by Our World in Data.
  • Trends in computing power used for AI, not including algorithmic progress. Because more computation produces more powerful AI, the exponential growth depicted in these charts indicates rapid improvement. Source: Epoch AI (2024).
  • Estimates of algorithmic progress in various domains. “Effective compute doubling time” refers to the time it takes for algorithmic progress to halve the computing power required to achieve a given level of performance. Note the log scale. Source: Epoch AI (2024).

:::

What's driving this progress? A few different factors are scaling up exponentially:[1]

  • The amount of money spent (2.6x growth per year), which is now on the order of a hundred million for the largest training runs.
  • The quality of computing hardware technology (1.35x growth per year), measured in computations available per dollar — see also “Moore’s Law.”
  • The quality of software algorithms (3x growth per year), measured according to how much computing power it takes to reach a given level of performance.
  • The amount of data used to train language models (3x growth per year).

The total computing power applied to AI increases exponentially because of the first two points, and once you include software improvement, what you might call the “effective computing power” increases at an even faster exponential rate.

This growth compounds quickly. The computing power used to train the largest models in 2024 was greater than in 2010 by a factor of billions[2] — and that’s before taking into account better software.

As a result of this scaling, AI has become smarter. Various AI systems can now:

  • Generate realistic images, speech, music, and video from a short text prompt.
  • Operate robots in factories and on construction sites.
  • Beat the best humans at many board and video games.
  • Predict the structures of proteins.

Scaling has also made some AI systems more general. A single language model can:

  • Hold a conversation like a human over text or audio.
  • Use its reasoning to play video games like Pokemon without specific training — badly, and with scaffolding, but it can do it.
  • Play chess at the level of an intermediate human player.[3]
  • Give sensible advice on a wide range of topics.
  • Tutor users competently.
  • Translate between most languages.[4]

To get a sense of the capabilities of current language models, you can try them for yourself — you may notice that current freely-available systems are much better at reliably answering complex questions than anything that existed two or three years ago.

You can also see AI progress on quantitative benchmarks. AI has rapidly improved on measures of its ability to:

  • Understand written text.[5]
  • Answer science questions suited to graduate students.[6]
  • Solve coding competition problems.[7]
  • Solve Olympiad-level problems in mathematics.[8]

The “horizon length” of AI has also doubled every seven months.[9] That is: newer AI systems are increasingly able to remain coherent in their thinking while doing coding tasks, succeeding at new types of tasks that take humans longer and longer amounts of time.

In fact, researchers are struggling to keep up and create new benchmarks. One benchmark, called “Humanity’s Last Exam,”[10] is a collection of problems from various fields that are hard even for experts. As of April 2025, AI progress on solving it has been only modest — although we haven’t checked in the past hour.

Given that progress is so fast, it’s natural to ask: will such trends take AI all the way to the human level?


  1. These numbers are from Epoch AI (retrieved 2025-04-03). ↩︎

  2. https://epoch.ai/blog/training-compute-of-frontier-ai-models-grows-by-4-5x-per-year ↩︎

  3. With the right prompting/examples/fine-tuning: https://dynomight.net/more-chess/ ↩︎

  4. https://arxiv.org/abs/2407.03658 ↩︎

  5. https://paperswithcode.com/dataset/mmlu ↩︎

  6. https://paperswithcode.com/dataset/gpqa ↩︎

  7. https://paperswithcode.com/dataset/apps ↩︎

  8. https://paperswithcode.com/dataset/math ↩︎

  9. https://metr.org/blog/2025-03-19-measuring-ai-ability-to-complete-long-tasks/ ↩︎

  10. https://paperswithcode.com/sota/humanity-s-last-exam-on-humanity-s-last-exam ↩︎



AISafety.info

AISafety.info is a project founded by Rob Miles. The website is maintained by a global team of specialists and volunteers from various backgrounds who want to ensure that the effects of future AI are beneficial rather than catastrophic.

© AISafety.info, 2022—2026

Aisafety.info is an Ashgro Inc Project. Ashgro Inc (EIN: 88-4232889) is a 501(c)(3) Public Charity incorporated in Delaware.