Arena.ai's new Alignment Index evaluates 27 models across 90,000 real-world agent sessions, ranking GPT-6.1-Sol at 87.9, Claude Opus 5.5 at 83.2, and Grok 4.7 at 82.7. But the $3.1B startup behind it ...
A two-phase study of over 2,500 Jordanian children finds that carefully calibrated simple regression models rival or beat ...
A new regional climate modeling study shows that replacing the fixed one-day aging timescale for black carbon with a dynamic, ...
Some results have been hidden because they may be inaccessible to you
Show inaccessible results