YOUR MODEL WATCHLIST

For your work.

Loading your workspace…

Your use case

PUBLISHED EVIDENCE

Your updates

Your workspace

Use cases save to this browser with a private cookie. Keep a recovery code to reopen them on another device.



Receive benchmark updates in your RSS reader, even when this page is closed.

Keep codes and feed links private. A new code or feed link replaces the previous one. Restoring a workspace replaces the workspace open in this browser.

A personal benchmark index

Your priorities weight published benchmark scores. Each benchmark is normalized against a fixed reference set: 50 is its reference mean, and 15 points is one reference standard deviation. The weighted average is your index. It isn't a percentage or a prediction of your success rate.

We use each named model's best published result for each benchmark. Reasoning settings and evaluation tools can differ across tests. Open a model to see the exact tested versions. Scores reflect those published setups, not necessarily the settings you use.

A model needs results for every selected capability to rank or trigger an update. Missing results appear separately. Small differences can be evaluation noise. An update is a reason to test a model on your work.

Task suggestions run on our own model server. It only suggests capabilities, never scores or models. Your description stays in your private workspace. Don't include passwords or confidential documents.

Evidence: Epoch AI (CC BY 4.0) and Vectara (Apache 2.0). Vectara measures document summarization, not a complete RAG pipeline. Its answer rate is shown alongside factual consistency.

Method cadence-personal-1. Reference population frozen September 12, 2026. Evidence refreshes every two hours; publisher updates can be less frequent. Price, latency, retrieval quality, and results on your own data aren't part of this index.