The source page loads here directly. Results follow Artificial Analysis' updates and caching. Scroll within the chart area to explore; if it is blank or difficult to use on your screen, open the full source above.
Apply the comparison
Match the benchmark to the deliverable.
These are HRI examples for your own pilot. Define what a usable result looks like before you compare models.
Excel / Analysis
A workbook you can audit.
Give two models the same monthly sales files. Ask for a variance analysis with formulas, a reconciliation to the source totals, and an exception sheet.
Check: Do the totals reconcile? Are formulas intact? Can a colleague change an assumption without rebuilding the workbook?
Supply a set of operating reports and ask for a two-page decision memo. Require a source for each material claim and a clear list of unresolved questions.
Check: Are the citations accurate? Does the recommendation follow from the evidence? How much rewriting does the reviewer need to do?
Turn the same approved report into a short management presentation. Specify the audience and the decision, and require editable slides with traceable figures.
Check: Does the story support the decision? Do chart labels match the source? Open the file to inspect layout, readability, and editability.
Choose one recurring deliverable. Use representative files you are authorized to share with the tools you are testing.
Hold the brief constant. Give each candidate the same inputs, output format, and definition of done. Record the model version and settings.
Review the actual artifact. Open the workbook, document, or deck. Check the content and whether the file works as intended.
Count the whole job. Record time to an accepted output, corrections required, and total cost. Repeat on a few different examples before standardizing.
Read with judgment
What to check before relying on a ranking.
Check the model settings and the index version.
Reasoning effort, tools, and the evaluation setup can change a result. A model score from an older index version may not be comparable with a current score. Use the version and labels shown in the source chart. Read the methodology.
Business-file scores are a closer lens, not a workflow guarantee.
AA-Briefcase combines correctness, analytical quality, and presentation in its overall score. Its file-type correctness chart is available; at our September 26 review, the separate analytical-quality and presentation file-type charts were marked "Coming soon". A polished deck still needs an accuracy check. Explore AA-Briefcase and example submissions.
Keep capability separate from product fit.
Also test the application your team will actually use. Check file access, integration with your existing systems, and the review process. A model leaderboard does not tell you whether a particular product fits those requirements.
How does this page stay current?
The embedded pages are served by Artificial Analysis when you open a view or reload the source. HRI does not copy scores into a static chart. New results appear when the publisher releases them and its cache refreshes; there is no guaranteed daily schedule.
HRI guidance reviewed September 26, 2026 against Intelligence Index v4.3.2 and AA-Briefcase v1.1. That date describes this guide, not the freshness of every source result. Charts and benchmark methodology belong to Artificial Analysis; HRI provides the interpretation and use cases.
When to return
A new model has shipped. Your volume has grown. A workbook now takes too long to review. Use a change in the work as the reason to revisit your shortlist.
Keep the comparison consistent
Save the view you use most. On your next visit, check whether the benchmark version has changed, then rerun one familiar task before switching your team's default.
Keep learning
Bring the next workflow to class.
Use the AI Practicum to turn a promising model into a repeatable way of working.