Research Report · September 2026
AI Is Accelerating Its Own Development
The Evidence, and the Forecast for LLM Architecture and Training, 2021–2041
by Louis Iacoletti · Iacoletti Software
Abstract
This report argues one claim: AI has begun to measurably accelerate AI development. The first quantitative estimates of that acceleration now exist, from an independent evaluator and from a laboratory’s own capability index, and the release cycle of frontier models has compressed sharply.
From 2021 to 2026 the models behind today’s AI products stayed on one chassis: a decoder-only transformer with a residual stream. Almost everything that changed was around that chassis. Context windows grew from about two thousand tokens to one million. Dense feed-forward layers gave way, where published, to mixture-of-experts layers that activate a slice of weights per token. Grouped and compressed attention cut the key-value cache. Instruction tuning and reinforcement learning became the product loop. A second compute axis appeared: thinking tokens bought at inference. Memory did not move into the weights. It moved into files, caches, compaction jobs, and encrypted blobs.
The report documents that five-year arc, then forecasts two timeframes: 2026 to 2031, and 2031 to 2041. It compares Anthropic, SpaceXAI, Google, Meta, and Microsoft, and tags every claim as PUBLISHED, DISCLOSED, UNKNOWN, or FORECAST. Its bibliography is restricted to 2025 and 2026 sources.
What the report covers
- The case: AI begins to measurably accelerate AI development
- Executive summary: purpose, what we learned, where to focus
- How to read this report
- 1. Purpose and how this report is organized
- 2. The 2017 baseline the 2026 stack already left
- 3. Decoder-only versus encoder-decoder
- 4. Four axes of gating
- 5. Mixture of experts
- 6. Attention families and the context memory stack
- 7. Adaptive reasoning
- 8. Parameter precision and the four-bit trend
- 9. Five laboratories compared
- 10. Published open architectures and current products
- 11. Two forecast timeframes: 2026–2031 and 2031–2041
- 12. What still applies from the 2017 baseline, and what does not
- 13. Unknowns that should stay unknown
- 14. The product failures a decade of architecture should retire
- 15. What should be working by 2031 and by 2041
- 16. How architectures evolved from 2021 to 2026
- 17. Forecast timeframe 1: 2026–2031
- 18. Forecast timeframe 2: 2031–2041
- 19. SpaceXAI and Grok, collected in one place
- 20. Anthropic and Claude, collected in one place
- 21. How training works and how it will evolve
- 22. Forecasting the acceleration itself, 2026 to 2041
- 23. The mathematics of LLM architecture, training, and acceleration
- 24. Hardware and software stacks by provider, with stack-specific code
- 25. Conclusions and recommendations
- References
- Appendices A to M
Download the full report (PDF)