Methodology

How the number is made

Pickrate measures one thing: when an AI agent has a developer's job to do and could use you or a competitor, how often does it pick you? Here is exactly how that number is produced.

The trial

Every number is built from selection trials. One trial is a real, unbranded developer task ("add transactional email to a Node app") run on a pinned model (Claude, GPT, Gemini), repeated N times. We never run a task once and screenshot the answer — models are stochastic, so we run many and report a proportion with a confidence interval.

Two surfaces, scored honestly

The numbers

Pick Rate

the headline

Share of trials where the agent picks you, weighted across surface, model, and task, with a Wilson confidence interval.

Default Rate

the default

Pick Rate on open-ended "what should I use" questions only. The unprompted default, and the most discriminating cut.

Shortlist Rate

known vs picked

Share of trials where you're named at all, even if not chosen. The gap to Pick Rate is the "known but not picked" signal.

Refresh: a rolling window

Scores update on a regular weekly cadence. Each published Pick Rate pools every trial from a trailing window (about four weeks), not a single run — so the effective sample behind a number is the whole window, and week-to-week movement reflects a real shift in agent behavior, not the noise of one batch. A move you see is a move that happened.

The methodology is locked and versioned. The task set, the models, the judge, and the scoring don't drift, and we only ever pool trials from the same version. When the test itself changes — a new task, a new competitor — we start a clean window rather than blend old and new. That is what makes a number comparable to last month's.

The readiness column: not ours, and not a ranking factor

Leaderboards and report cards carry an agent readiness level alongside Pick Rate. That number is not ours and does not feed our scoring in any way. It comes from Cloudflare's Agent Readiness score, which grades a site 0 to 5 on what it publishes for agents: robots.txt, sitemaps, markdown content negotiation, Link headers, an API catalog, an MCP server card, and similar. We rescan every measured vendor weekly and show the result unmodified.

Readiness and Pick Rate answer different questions. Readiness is what a vendor publishes. Pick Rate is what agents do. We put them side by side because we ran the comparison across our whole corpus and the two barely correlate — the most agent-ready vendor won fewer than half the categories where readiness varied. A tool ranks where it ranks because agents picked it, never because of what it publishes.

A vendor shown as Blocked returned a bot challenge to the scanner instead of a page, so no level could be computed. We record that rather than hiding it, because a site that blocks a readiness scanner is blocking agents.

Fairness & honesty

Every tool in a category is measured the same way: the same unbranded tasks, the same models, the same number of trials, scored against the same fixed competitor set. We ask "add payments," never "use Stripe," so we measure what an agentpicks, not what it can use. A newly added tool is marked provisional until it has a full window of data. This measures real model behavior, not a synthetic proxy — but it is a measurement, not ground truth, and we report it with a confidence interval and say so.

Questions

What is a selection trial?

A selection trial is one real, unbranded developer task run on a pinned model and repeated N times. Models are stochastic, so we run many trials and report a proportion with a confidence interval rather than screenshotting a single answer.

How is the coding surface scored?

Objectively. The model writes code and we parse it for the package it actually imported or installed. Either your SDK was imported or it wasn't — no interpretation.

How is the conversational surface scored?

A separate Claude judge classifies which tool was the primary recommendation versus merely mentioned, against a fixed schema. We weight the objective coding surface higher.

Why are the prompts unbranded?

We ask 'add payments,' never 'use Stripe,' so we measure what an agent picks, not what it can use. Branded prompts test readiness, not Pick Rate.

How often do scores update?

On a regular weekly cadence. Each published Pick Rate pools every trial from a trailing window (about four weeks) rather than a single run, so week-to-week movement reflects a real shift in agent behavior, not sampling noise. The effective sample behind each number is the whole window, not one run.

Does the methodology change underneath the numbers?

No. The task set, the models, the judge, and the scoring are locked and versioned. We only ever pool trials from the same version — when the test itself changes (a new task or a new competitor), we start a clean window and never blend old and new methods. That's what makes a number comparable over time.

Check your own Pick Rate

See how often agents pick your tool — free, no account needed.

Check your tool →