example.com/path/to/article
000 points · username · 0 hours ago
example.com28 points · 7 comments · 1 day ago · ronfriedhaber
FailMore
simonw
We match DeepSeek V4 Pro Base using ~50x fewer FLOPs – that’s around half of GPT3’s pretraining compute, or ~$0.5M on GB200.
If this holds up that's a really big deal.
ismael_rr
monneyboi
vatsachak
Basically, you can think of a LLM as a function which generates a probability distribution of words. If the next word in a series is "they", and one model predicts that word 40% of the time, and another model predicts that word 1% of the time, the latter model is worse as it is more surprised by the true distribution.
You can convert these probabilities into "bits":
surprise in bits = −log₂(probability of the actual token)
This is then normalised by text length:Bits per byte = total next-token surprise in bits / number of bytes in the evaluated text
So the lower you go on the charts, the less surprises in the LLMs distribution (a better model).
For more info: https://smalldocs.org/s/DfvdGuFsiR3LlzXw1H5J0K#k=AJ8V1AQECYj...