By testing combinations of year-specific model recipes and data corpuses, the researchers determined that data improvements were over three times more effective than model improvements at a specific training scale.
These findings highlight a shift in the AI infrastructure landscape toward sophisticated data engineering.
While early datasets like 2019’s OpenWebText relied on simple Reddit links, 2025-era corpuses like UltraFineWeb utilize massive web scrapes and AI classifiers to filter for high-quality information.
The source notes that while model research has introduced critical stability and scaling innovations that allow for larger clusters and more parameters, the actual "intelligence" gains at smaller scales are largely a result of better curation and extraction of the information being fed into the system.
As the industry moves forward, the reliance on high-quality web data may face a "data wall" as the supply of human-generated internet content is exhausted.
The study suggests that while small models require strict data filtering to perform well, larger frontier models may eventually become "container ships" capable of processing less refined data through sheer capacity.
Future progress may depend on whether synthetic data—information generated by AI rather than humans—can successfully expand training corpuses without degrading performance, or if automated research can further accelerate data curation techniques.