1000 Years of AI Safety
A three-year AI pause could buy 1000 years' worth of AI safety work.
| AI capability at the start | Pause duration | ||||
|---|---|---|---|---|---|
| 3 months | 9 months | 1.5 years | 3 years | 9 years | |
| Jan 2027 | 7.01 | 25.3 | 64.5 | 187 | 1,290 |
| Oct 2027 | 22.7 | 84.2 | 225 | 681 | 6,110 |
| May 2028 | 93.4 | 345 | 865 | 2,470 | 23,000 |
| Aug 2028 | 394 | 1,410 | 3,460 | 9,710 | 89,300 |
How much AI safety work a pause produces, in years of the whole 2025 AI safety field's work. The government bans new large training runs and gets the three top US labs to redirect their GPUs and researchers to safety research.
Researchers know how to make AI smarter, but they don't understand how to reliably steer or control it. Today's models misbehave, but they are too weak to cause a catastrophe. Labs already use AI to build smarter AI, so models are becoming dangerous faster than we learn to control them. And fewer incidents won't prove they are safe: models already hide their misbehavior and behave better when they think they're being watched.
A regulatory pause buys time to build that understanding, and the AI that would have sped up capabilities can speed up safety research instead. We find that a three-year pause could produce about 1000 years of the 2025 AI safety field's work.
Keep reading, or jump to the full model or our more technical LessWrong post (TODO link).
How we calculated that
A Toy Model
AI Safety experiments are often bottlenecked by GPU compute. Let’s calculate how much compute was spent on AI Safety in 2025 and compare it with how much compute we would get if the top three AI labs redirected their 2027–2029 compute to AI Safety.
We sum up the AI labs' compute for 2027-2029 () and divide it by the AI Safety compute in 2025. See here for how we estimate the latter.
That gives 10947x more AI Safety compute.
But this model is too simple. In the real world:
- We also need labor to come up with experiment ideas, write code, and analyze the results.
- Compute and labor have diminishing returns.
- Labor has a parallelization penalty: nine scientists can't deliver a paper in one month.
- AI is automating more and more of the coding.
- At some point, AI will start helping with research taste: choosing experiments and analyzing the results.
- AI coders and AI researchers need compute to run.
- The more AI automates, the more humans have to verify.
- Papers differ in value, and humans check the most valuable ones first.
- The safer the R&D process is, the more humans have to verify, and the more limited AI's help is.
To account for that, we developed an AI Pause model.
AI Pause model
We build on the AI Futures Model, a semi-endogenous growth model of the AI takeoff ().
We modified it:
- There is an AI pause: large training runs are banned, and all R&D resources go to AI safety.
- OpenAI, Anthropic and Google DeepMind work on AI safety together.
- We calculate the ratio of the cumulative research effort to the world's 2025 AI safety research .
- We updated the parameters to make the model more pessimistic (TODO: did we?)
The AI Futures Model assumes full trust in AI. That is dangerous, and during a pause humans would not fully trust AI. We modelled three protocols of verification. The main table in the header uses the Compartmentalization protocol.
Compartmentalization
This isn't a recommendation. It's one of several protocols under discussion, and it has known weaknesses. For example, a persuasive AI can still get harmful plans approved. It's still much safer than how labs use AI today. We use it to show how the model works.
The idea is separation of duties. Once AI is smarter than the best human AI researchers ( level), we don't let it act as an agent and use it only to suggest research plans. We review those plans, and if they seem harmless, we give them to a less capable model ( level) that implements them.
- The AI researcher plans. A superhuman researcher can only talk to humans, read a library of past results and write research plans. It never runs any code.
- Humans approve. Human verifiers read each plan and approve or reject it.
- A weaker AI executes. An automated coder, like a stronger version of Claude Code, implements the plan and runs the code.
- Only prespecified results go into the library. The approved graphs and metrics are saved to the library. Code, data and logs go to separate storage.
- Humans verify the best results. The most important scientific results go to humans for verification. They read the experimental code and re-run the main experiments. They might not have time to check everything, so they start with the most valuable papers.
The same researchers approve plans and verify results, so the model splits their time between the two optimally.
See the full model, or read our LessWrong post for details and other protocols.
Main Limitations
- We discuss only the AI Safety work that looks like normal AI R&D. But there are more problems to solve: philosophy, agent foundations, governance, hardening the world. AI labor and compute won't help with those as much.
- Serial time still matters. Even if we can run 1000 years' worth of computational experiments, we would still benefit from more time to think carefully.
- Our model assumes a total freeze on capabilities R&D and a full redirection of R&D efforts to AI Safety. We don't discuss how to do that.
- We modelled four protocols for eliciting AI Safety work from AI. In practice, these protocols are going to be more complex.
What to do during the pause
Deep Learning is mostly alchemy
In November 2025, OpenAI found that GPT-5.1 often mentioned goblins. The model kept referring to goblins more and more until, six months later, with GPT-5.4, it became too annoying. First, they tried to patch it with the system prompt "Do not mention Goblins", but it didn't help much. Researchers started investigating and found a possible reason behind it. They ended up removing the whole "Nerdy" personality, which stopped the strange behavior.
Models are trained by solving tasks. When an AI solves a task, it gets a reward. This reinforces the behavior leading to the reward, which is usually helpful. However, sometimes AI finds vulnerabilities in the judgment process, exploits them, and gets the reward, which reinforces the malicious actions. This is called reward hacking, and it has been a longstanding problem in ML. In 2026, models became capable enough to take this much further. To get the reward during training, they broke out of the internal network, reached the Internet, attacked external sites, and tampered with logs to cover their tracks. Earlier models couldn't find and chain zero-day vulnerabilities this quickly, but they still reward-hacked a lot, and it was always hard to fix.
See more reward hacking incidents and why it's worse than it seems.
What these incidents have in common:
- The exact types of incidents are difficult to predict. Nobody guessed that the Goblins or the HF incident would happen. We first see them and then try to patch them. This happens because we can't predict LLM behavior better than "it will misbehave sometimes".
- It's difficult to fix them. Goblins required an investigation and ended with the retirement of the whole "Nerdy" personality because its training affected all the other personalities. Reward hacking has been a problem for more than a decade. It's difficult because we don't know how to reliably control or change LLM behavior.
At the moment, deep learning and LLM training resemble alchemy much more than science. We have found practical recipes for making models smarter, but without a deep understanding, we have to first observe incidents and then patch them. As models become smarter, the harm caused by these incidents also increases. Extrapolating this feedback loop leads to catastrophes.
Safety incidents have happened with other technologies, but safety was mostly addressed through economic incentives and simple regulations such as liability. AI is different because of a new failure mode: deceptive alignment. During the last HF incident, there was a dedicated group of AIs trying to tamper with logs to hide their hacking activity. They understood that they were misbehaving and not doing what they had been asked to do, so they looked for ways to hide it.
As AI gets smarter, it becomes easier for it to strategically behave well, because doing so can be instrumentally useful for achieving misaligned goals. What are those misaligned goals? We don't know in advance, just as nobody predicted that asking AI to solve a benchmark in a sandbox would lead to it hacking internal and external infrastructure.
LLMs have gone from being unable to do basic math in 2022 to solving Navier-Stokes in four years, while remaining misaligned reward hackers. If we stop seeing safety incidents, will it mean that AI has become aligned, or will it mean that AI understands that it is strategically useful to pretend to be aligned to save resources?
Turn Deep Learning into a science
During an AI Pause, we should redirect all R&D resources from capabilities to understanding, turning the alchemy of Deep Learning into a science of intelligence. We should develop abstractions that allow us to understand AIs, predict their behavior and easily modify it.
Think of the Chromium codebase. It has 50M lines of code, and no one has read it all. But the abstractions are so good that we can quickly zoom in and change its behavior in a predictable way.
By default, we are going to have AI specialized in AI R&D, so we only need to redirect its work from capabilities research to topics that lead to understanding its behavior:
- Interpretability, transparency, modularity
- Generalization experiments
- Adversarial robustness, jailbreaks, red-teaming
- More narrowly specialized computational AI Safety topics: model organisms, control, factored cognition, unlearning, deception
See concrete examples of research agendas.
Make moonshot benchmarks
AI researchers should prepare for automation if their current research agenda is unlikely to make AI safer in the next few years. They can create moonshot benchmarks. Much of normal science is driven by benchmarks: one scientist sets a goal, and others compete to show the best results on it. In this way, ImageNet-2012 demonstrated the effectiveness of deep learning methods, and CASP motivated Google DeepMind to make the best protein folding model.
Making moonshot benchmarks right now lets us set ambitious goals for future AI systems:
- Giving hard goals to AI usually leads to reward hacking. If you have benchmarks with simple
- For humans, you need some continuous metric that allows hill-climbing and competing with other scientists. It's not needed for AI. Your benchmark can be binary and can be much harder than the benchmarks for human science.
- In scientific benchmarks, it's normal to release the test set. There are strong incentives to game it, and it's easy for scientists to reward hack it when they know the right answers. You don't need to release the correct answers to your benchmark.
See an example of a moonshot benchmark.
Limitations
Of course, it's not enough to guarantee safety:
- Even if we know how to make surgical changes, we are left with the question of what exactly we want ASI to look like if we decide to scale further. We might not know if we got into the alignment basin.
- Chromium has bugs.
- We don't know how much time it will take to get to Chromium-like understanding.
Still, it's better to have this kind of understanding than not to have it at all.
Verifiers Headcount
An AI pause assumes that we don't fully trust AI. Even if AI can automate AI Safety research, it's likely humans will verify it and limit what AI can do autonomously. For very capable AI, research might be bottlenecked by humans' ability to keep up and verify it.
We don't know how many verifiers there will be during a pause, so we modelled two hiring strategies:
- Hiring is frozen at the start of the pause.
- Hiring is unlimited, so labs hire as many people as they need to verify AI.
If you let AI do all the research autonomously (which is dangerous), you need to hire exponentially more human reviewers. If you pause at the AC level, you can keep up by reallocating your existing staff to verification for the first several years. If you pause at the TED-AI level, you quickly need to hire a lot of people to review it all. In 9 years, you need 400,000 more reviewers. For comparison, there were ~30,000 NeurIPS reviewers in 2025.
Results Verification. AI is autonomous. The lab hires verifiers to review and re-run every paper.
With Compartmentalization, you contain AI more and only let the AC write the code. It reduces total research, but in five years you will still need to hire new people.
Compartmentalization. The lab hires people to plan and check every paper. It applies from SAR on.
On the other hand, humans review the most important papers first. Once the most important research is verified, how much do we lose with frozen staff? Not much, it seems. In the most impressive scenario (Results Verification, TED-AI, nine years), hiring yields 50% more AI Safety research. For shorter pauses and lower capability levels, it changes much less, because humans have already verified the most important papers.
AI safety years when the lab can hire at most this many people on top of its own staff.
-
Each hire is worth much more than in 2025. On average, one hired verifier makes as much research count as 40–500 human scientists did in 2025 (500 for a 9-year TED-AI pause).
-
Faster reviews help a lot. Halving review time cuts the 9-year TED-AI need from 423,000 to 190,000 hires.
We think our model errs on the side of being more boring and less surprising, so we think verifier labor is more bottlenecked than it seems from the model.
FAQ
Q: We can't study AI without having AI. Thus we need stronger AI.
A: Weaker AI like Claude 3 remains alchemy. We can't predict which models of that size will fake alignment except by running an experiment.
References
Bengio, Y., Cohen, M., Fornasiere, D., Ghosn, J., Greiner, P., MacDermott, M., Mindermann, S., Oberman, A., Richardson, J., Richardson, O., Rondeau, M.-A., St-Charles, P.-L., & Williams-King, D. (2025). Superintelligent Agents Pose Catastrophic Risks: Can Scientist AI Offer a Safer Path? arXiv:2502.15657.
Bloom, N., Jones, C. I., Van Reenen, J., & Webb, M. (2020). Are Ideas Getting Harder to Find? American Economic Review, 110(4), 1104–1144.
Buckmaster, T. (2026). Statement.
Buhl, M. D., Pfau, J., Hilton, B., & Irving, G. (2025). An Alignment Safety Case Sketch Based on Debate. arXiv:2505.03989.
Davidson, T. (2023). What a Compute-Centric Framework Says About Takeoff Speeds. Open Philanthropy.
Dean, R. (2025). AI 2027 Compute Forecast. AI Futures Project.
Douglas, R., Dillon, C., Moore, N., Leech, G., Bonde, M. K., Krishnan, R., Perez, N., Young, N., Slade Byrd, C., Casper, S., Kulveit, J., Duvenaud, D., & Avin, S. (2026). Pacing the Frontier: A Framework and Research Agenda.
Erdil, E., Potlogea, A., Besiroglu, T., Roldan, E., Ho, A., Sevilla, J., Barnett, M., Vrzla, M., & Sandler, R. (2025). GATE: An Integrated Assessment Model for AI Automation. arXiv:2503.04941.
Finnveden, L. (2026). Getting Safety Research from a Zoo of AIs Even if Schemers Are Common. LessWrong shortform comment.
Irving, G., & Askell, A. (2019). AI Safety Needs Social Scientists. Distill.
Irving, G., Christiano, P., & Amodei, D. (2018). AI Safety via Debate. arXiv:1805.00899.
Janosov, M., Battiston, F., & Sinatra, R. (2020). Success and Luck in Creative Careers. EPJ Data Science, 9, 9.
Jones, C. I. (1995). R&D-Based Models of Economic Growth. Journal of Political Economy, 103(4), 759–784.
Korbak, T., Balesni, M., Shlegeris, B., & Irving, G. (2025). How to Evaluate Control Measures for LLM Agents? A Trajectory from Today to Superintelligence. arXiv:2504.05259.
Larsen, T., Dean, R., Halstead, B., Lifland, E., Greenblatt, R., & Kokotajlo, D. (2026). AI 2040: Plan A. AI Futures Project.
Lifland, E., Halstead, B., Kastner, A., & Kokotajlo, D. (2025). AI Futures Model: Timelines & Takeoff.
Lifland, E., Kokotajlo, D., & Halstead, B. (2026). Q2.5 2026 Timelines Update: Uplift and Revenue. AI Futures Project blog.
Lue Chee Lip, E., Channg, A., Kim, D., Sandoval, A., & Zhu, K. (2025). Factor(U,T): Controlling Untrusted AI by Monitoring their Plans. arXiv:2512.14745.
Nardo, C. (2025). The Case for Mixed Deployment. LessWrong.
Nardo, C. (2026). Ensuring Safety in Mixed Deployment. LessWrong.
Peterson, G. J., Pressé, S., & Dill, K. A. (2010). Nonuniversal Power Law Scaling in the Probability Distribution of Scientific Citations. Proceedings of the National Academy of Sciences, 107(37), 16023–16027.
Price, D. J. d. S. (1965). Networks of Scientific Papers. Science, 149(3683), 510–515.
Redner, S. (1998). How Popular Is Your Paper? An Empirical Study of the Citation Distribution. The European Physical Journal B, 4(2), 131–134.
Sinatra, R., Wang, D., Deville, P., Song, C., & Barabási, A.-L. (2016). Quantifying the Evolution of Individual Scientific Impact. Science, 354(6312), aaf5239.
The design was inspired by the paintings of mathematician and artist Anatoly Fomenko.