Verify evals on Papers with Code

Author: NielsRoggeCreated Sep 14, 2026Updated Sep 14, 2026

Hi,

Niels here from the open-source team at Hugging Face.

I've made the following papers and their evaluation results available on Papers with Code:

The Llama 3.1 (405B, Instruct) results currently rank first on BFCLv4, MGS, and ZeroSCROLLS/QuALITY.

The Llama 3.1 (8B, Instruct) result currently ranks first on NIH/Multi-needle.

Would it be possible to verify these results and let me know if any score, model name, benchmark protocol, or openness metadata should be corrected?

You can also edit the task, methods, project page, and GitHub URL directly from each paper page using your Hugging Face account.

If you'd like to showcase the results in your repository README, you can copy these live leaderboard badges (or use the “Copy PwC badge” button in the Results section):

Papers with Code: SOTA on BFCLv4 Papers with Code: SOTA on MGS Papers with Code: SOTA on MGSM Papers with Code: SOTA on NIH/Multi-needle Papers with Code: SOTA on ZeroSCROLLS/QuALITY Papers with Code: #2 on ARC

Kind regards,

Niels