Best AI News — Updated Every 3 Hours
Story Page
← All Stories
Home Community Story
Community

[R] GPT-5.4-mini regressed 22pp on vanilla prompting vs GPT-5-mini. Nobody noticed because benchmarks don't test this. Recursive Language Models solved it.

Via r/MachineLearning
Sunday, Mar 29, 2026 · 9:43AM
Summary

GPT-5.4-mini produces shorter, terser outputs by default. Vanilla accuracy dropped from 69.5% to 47.2% across 12 tasks (1,800 evals). The official RLM implementation dropped too (69.7% to 50.2%). Our implementation - where the model writes Python to query data instead of attending to all of it with

Continue reading the full article
Read at r/MachineLearning
www.reddit.com
Back to all stories