by Haytham ElFadeel - hfadeelm@gmail.com
2024
Recent discussions suggest that scaling laws in AI might be slowing. But does this mean innovation is hitting a ceiling?
My take: the underlying scaling relationships haven’t “broken”—what’s changing is where the best ROI is. Some popular benchmarks are saturating and pretraining gains can look incremental, so the field is reallocating effort toward dimensions that return more capability per unit compute. Right now, the lowest-hanging fruit looks a lot like scaling reinforcement learning / post-training optimization (especially for reasoning and tool-use), not just “more pretraining FLOPs”.
This article explores: (1) what “slowing” actually means, (2) the general pattern of technological progress, and (3) what’s next for language models.
Part 1 — Is it really slowing, and if so why?
Advancement in math, reasoning, and language has felt slower in 2024 compared to the rapid progress of 2022–2023. At the same time, we’ve also seen major progress on new frontiers (e.g., multimodality and improved real-time interaction).
A key nuance: “progress” is not a single scalar.
If you track a saturated benchmark, year-over-year gains will naturally flatten even if the underlying model quality is improving. Engineers also shift attention to new capability (e.g. vision, audio, planning, tool use, long-horizon tasks), so the headline benchmark deltas can understate real innovation.
Human nature and the scaling laws…
Some argue that scaling laws are slowing or even breaking. This is not true—at least, not yet.
Interestingly, this narrative echoes similar "doom and gloom" predictions in other fields. Take Moore’s Law, which predicts that the number of transistors on integrated circuits doubles roughly every two years. Critics have predicted its demise for over 20 years, yet Moore's Law remains alive, supported by robust roadmaps from industry leaders like IMEC, TSMC, and Intel [ref]. As Peter Lee, Microsoft Research VP, once joked: "The number of people predicting the death of Moore’s Law doubles every two years".
This phenomenon isn’t unique to technology. The U.S. Bureau of Mines in 1919 and M. King Hubbert in 1956 predicted that we will run out of Oil soon. Thomas Malthus in 1798 and Paul Ehrlich in 1968 predicted that we will run out of food "The battle to feed all of humanity is over. In the 1970s hundreds of millions of people will starve to death in spite of any crash programs embarked upon now".
What all those predictions missed is human ingenuity, our ability to think creatively, to innovate and find solutions.
What the scaling law tell us…
The scaling law [ref] predicts a ~20% reduction in loss for every order-of-magnitude increase in compute (i.e. increase in compute here means a combination of model size, amount of training data, training time/compute). It’s important to remember, the scaling law doesn’t guarantee uniform improvements across all downstream tasks. To put that in perspective, if we make the simplistic assumption it translates directly for a given benchmark with 80% accuracy, with the order of magnitude increase of compute the new accuracy will be 84%.
Scaling also faces practical challenges, Software moves much faster than Hardware. While Nvidia's latest GPUs (Blackwell) offer 3x-6x performance improvements over their predecessors (Hopper), training larger models is hard and constrained by hardware, power, and economics. This has spurred innovation in areas like parallelism (to maximize the cluster utilization and efficiency), reward optimization (RLHF, DPO), synthetic training data (to maximize training signal and compute), and inference-time techniques like (Chain of Thought - CoT, Tree of Thought - ToT, and Reflection),
“Slowing” as an ROI story: we’re optimizing the best axis
The more interesting question is not “is scaling slowing?” but:
Which axis of scaling gives the highest return per dollar right now?
Today, pretraining while still very important is increasingly constrained by:
- Hardware availability and cluster economics.
- Training throughput / parallelism complexity.
- and (in some domains) data quality and domain coverage.
As a result, the field is leaning into high-ROI scaling dimensions, including:
- Better training signal: synthetic data, filtering, curriculum, domain mixing.
- Inference-time compute: more thinking at test time (search, deliberation, self-consistency).
- Post-training scaling: RL-training.
Part 2 — The pattern of technological progress
Advancement in technology doesn’t follow a linear trajectory. It often follows two macro-patterns:
- Law of diminishing returns (S-curve): long gestation → rapid takeoff → plateau.
- Law of accelerating returns: when one S-curve saturates, a new breakthrough starts another S-curve that stacks on top of the previous one.
When one S-curve plateaus, the next “engine” of progress tends to be a new paradigm, a new data source, a new objective, or a new interface layer.
Stacked S-curves show up throughout history:
Machine Learning
- 1980s: k-NN, rule-based systems, backprop
- 1990s: decision trees, boosting, SVM
- 2000s: kernel methods, LSTM, random forests
- 2010s: deep learning, CNNs, autoencoders, pre-training
- 2020s: web-scale transformers, diffusion, multimodality, post-training/RL scaling
Phones
- 1920s: rotary phones
- 1960s: dial pad phones
- 1970s: cordless phones
- 1980s: cell phones
- 1990s: texting
- 2000s: smartphones (internet, apps, cameras)
- 2010s: better smartphones (connectivity, cameras, sensors)
The real challenge for leaders and innovators is (a) when to jump to the next curve and (b) how to tell a real curve from a dead end.
Part 3 — What’s next for language models?
1) Iterative improvements: better economics, especially for smaller models
Besides pushing bigger GPU clusters for the next generation of foundation models, there is an equally significant push to improve the economics of training and inference.
Smaller models (SLMs) have gotten substantially better thanks to:
- higher-quality and more targeted data mixtures.
- better optimization and architecture tweaks.
- improved objectives and training recipes.
- post-training advances (e.g. RLVR).
This trend should continue: smaller, cheaper models are easier to deploy at scale, so any per-parameter capability gain has outsized practical value.
2) More applications: agents and workflow automation
Workflow automation (UI automation + tool use) is a major frontier, especially with multimodal systems integrating vision and audio.
Agents aim to execute multi-step tasks on a user’s behalf: navigating UIs, calling APIs, planning, verifying, retrying, and summarizing outcomes. The hard part isn’t “one action” — it’s reliable long-horizon behavior under uncertainty, with good security boundaries.
3) Searching for the next big thing
Here are a few obvious “next dimensions” (and how I’d reframe them with an ROI lens):
- Scaling
- Pretraining scale will continue, but it’s increasingly expensive and operationally complex.
- The more immediate ROI is often in scaling post-training compute: RL rollouts, verifiers, search, and evaluation-driven learning loops.
- Planning
- LLMs can plan, but robust hierarchical planning remains limited. Better planning likely requires richer world models, better memory, and training signals that reward long-horizon success (again: an RL-shaped problem).
- Robotics
- End-to-end “language → grounded action” remains open: mapping natural language goals into reliable sequences of motor actions in messy, partially observed environments.
- Hallucinations / reliability
- Reducing hallucinations by a few percent is easy; eliminating them is hard.
- Reliability likely improves via verification loops, tool grounding, and training on correctness signals—again strongly aligned with post-training/RL scaling.
- More data and more modalities
- Text isn’t the only way to scale. Multimodal training helps anchor representations and improve generalization.
- A likely path forward includes more video/audio and more structured interaction traces (tool calls, UI trajectories), not just web text.
Conclusion
While scaling laws may appear to slow, the broader trajectory of AI development continues its march forward with new capabilities (e.g. multimodality - vision, audio, multi-step reasoning).
History shows that innovation thrives on challenges, and new breakthroughs will pave the way for the next era of technological progress — the next S-curve. The question isn’t whether we’ll overcome these hurdles—but how and when.