example.com/path/to/article
000 points · username · 0 hours ago
example.com133 points · 45 comments · 7 days ago · tkgally
Nine months have passed and much stronger models have been released, so I tried the experiment again today. The linked site shows the results.
Running ten of the prompts through six models at OpenRouter cost about twenty dollars, so I stopped there for now.
svcrunch
vova_hn2
I think they all just added a bunch of similar tasks to their training sets, so we cannot judge true emergent capabilities of the models anymore.
samayashar
I think that intertwining two entities (living/non-living) is still challenging but overall they're pretty sound.
outlore
kennywinker
ianberdin
MacBook Pro 3D in SVG for me the most helpful one.
meta-level
steinvakt2
sajithdilshan
theshrike79
[deleted]
zaphar
andy_ppp
miohtama
mock-possum
sceptic123
An elephant typing on a typewriter
A monkey, surely?
eddytrex_
BrokenCogs
pohl
neilellis
dcreater
dustfinger
qiine
water-drummer
[deleted]
simonw
villish
Qwen3.8 is very clearly distilled from Claude models.
1. It tests visual reasoning and structured output in a single task.
2. It seems to sort correctly on advancing general intelligence. As a counterexample, if I'm not misremembering, artificialanalysis.ai made some changes to their benchmark recently after Astra ranked below several older models.
3. While models have gotten significantly better in the past 2 years, the top model is still at 0.78 F1, so the test is not yet saturated. As a reference point, when I started, the top models were in the [0.1, 0.2] range.
[1] https://dorrit.pairsys.ai/