While benchmarking its ChatGPT LLMs internally, OpenAI researchers noticed
that after a significant amount of computational effort on a given size LLM, there was no improvement in Test Loss error reduction. Instead, the loss curve takes on a reverse-S shape or sigmoidal characteristic where the loss eventually remains constant at the foot of the S.
To combat this limitation, larger LLM instances—containing decades more neu-
ral net connections—were invoked to further reduce the Test Loss. However, that resulted in a sequence of successively lower sigmoidal curves that appear to fall on a descending line, which they called the ”compute-efficient frontier” (CEF). The CEF bound raised concerns that it might constitute a universal constraint on the scalability of all LLM systems.
This talk presents a model of LLM computational dynamics, based on the Univer-
sal Scalability Law, that provides a framework for understanding the CEF. The USL defines bistable minima in the tokenized neural net landscape.
This talk explains:
1. Why all training loss curves have a sigmoidal characteristic
2. Why the observed CEF constraint exists
3. What the CEF correlation means for LLM scalability
More details have been published in reference [1].