RESOURCES · FIELD GUIDE

Why a Docs Vendor Can't Run the Docs Benchmark

A benchmark run by a vendor selling the remedy carries a structural conflict. What an independent API benchmark buys, and the test any score must meet.

Read
7 min
Updated
2026-08-06

Agent-readiness scores are starting to appear across the industry. Documentation platforms, infrastructure providers, and API tooling companies have all noticed the same shift Discry measures: AI agents now read public documentation and decide, in seconds, which APIs they can operate. Scoring that gap is a reasonable instinct, and more measurement in this space is broadly good news. Whether any particular score deserves your trust turns on an older question — who runs the measurement, and what they stand to gain from the result.

This piece makes that argument in general form, because it applies in general form. It applies to any vendor whose product is the remedy for the failures its benchmark finds. It also applies to Discry, which is why the last section turns the same test on us.

The structural conflict in vendor-run scoring

When the organization grading your documentation also sells the product positioned to fix it, the measurement carries a built-in lean. Every F the benchmark hands out is a qualified sales lead for the grader's own remedy. The people running it can be honest, careful, and technically excellent; the structure still points every incentive in one direction. Finance learned this expensively. Credit ratings paid for by the issuers being rated produced systematically inflated grades in the run-up to 2008 — the conflict was structural, and it survived the good intentions of everyone inside it. Accounting learned the same lesson earlier, and codified auditor independence as a requirement rather than a virtue.

For a documentation benchmark, the lean shows up in three specific places.

Methodology drifts toward what the product remediates. A rubric designed inside a vendor tends, over time, to weight the checks the vendor's tooling passes by default and to underweight the failure modes that tooling produces. The drift is rarely deliberate. It happens through a hundred small design decisions, each locally defensible, each made by people whose product roadmap is the fix.

The failing grade doubles as sales collateral. A benchmark whose worst results feed the grader's pipeline has a reason to find failure, and a reason to describe failure in terms its product resolves. The score stops being a measurement you can act on independently and becomes the top of a funnel.

Customer status pressures the grade. When a graded API becomes a paying customer, a low score becomes an account-management problem. The pressure to soften a finding for a signed customer, or to time a rescan generously, exists whether or not anyone yields to it — and an outside reader has no way to verify which happened.

Goodhart's law describes what happens when a measure becomes a target: it stops being a good measure. A vendor-run score arrives in the world already a target — designed, weighted, and maintained by the party selling the path to a better number.

What independence buys

Independence is a property of how the measurement is constructed, and it purchases concrete things.

Findings with no stake in the stack that produced them. An independent API benchmark writes no documentation and hosts none. It can grade a platform-generated docs site and a hand-rolled one by the same rule — what an agent can actually fetch and understand — because it earns nothing from the answer either way. That neutrality is what makes results comparable across the market: an index ranking APIs built on different tooling means something only when the ruler has no position in the tooling.

A methodology free to penalize what tooling produces. Documentation generators have house styles, and some of those styles fail agents — pages that render empty without JavaScript, entry-point files that route to marketing pages, reference docs split into fragments no context window assembles cheaply. An independent instrument can name those patterns and score them wherever they appear, including in the output of the most popular tools on the market. A benchmark whose parent company ships one of those tools carries an obvious reason to look elsewhere.

Grades that hold still when a customer signs. Under an independent instrument, the same scan runs for a prospect, a customer, and an API that will never spend a dollar. The commercial relationship has no input into the number. Producers can cite the grade upward to a board, and agent builders can rely on it downward in a build decision, precisely because it moves for one reason only: the documentation changed.

An audience beyond the sales funnel. Agent builders use readiness data as pre-flight ground truth — which APIs their agent can find, understand, and afford. That audience needs the data to be true more than it needs the data to be flattering, and serving it disciplines the whole instrument. A benchmark that exists to generate remediation leads has no such audience keeping it honest.

The standard any credible benchmark must meet

Independence makes credibility easier. Construction makes it real. Whoever runs a score — a vendor, a foundation, or Discry — you can apply the same test before accepting a single number.

Published methodology, end to end. You should be able to read exactly how the score is produced before you see your result. Discry's is public at /methodology, including the parts that are unflattering to measure.

Behavioral evidence over checklists. A credible score demonstrates what an agent can do with the documentation. Behavioral measurement means real models are quizzed against the live docs and their answers are graded mechanically against ground truth cited to specific locations in the documentation itself — human judgment, and with it human incentive, stays out of the grading loop.

Public task banks. The questions should be inspectable. A published task bank lets anyone verify that the tasks measure integration-relevant facts, and that they include trap tasks — requests for capabilities the API lacks, where the correct answer is a refusal. Hidden tasks can be tuned; published ones are accountable.

A control for prior knowledge. Models have read famous APIs' docs in training. A closed-book baseline establishes what a model already knew without the docs and excludes it, so fame earns nothing and obscurity costs nothing.

Receipts for every claim. Each result should ship with the URLs fetched, the tasks that failed, and the reason — enough for the graded party to reproduce the failure and for a skeptic to check it.

A grade that explains itself. A single unexplained number invites exactly the manipulation this article describes. The Discry Score decomposes into discovery and comprehension so a reader can see where a result came from and what would change it.

Every item on that list is checkable against Discry today. If we ever fall short of one, the shortfall is checkable too — that is the point of the list.

Discry's version of the conflict, and what contains it

Discry has a business attached to its benchmark. We help teams get agent-ready, and we sell fix artifacts derived from what the scan measured. Stated plainly: the rate-and-remediate conflict this article describes exists here too, one layer down, and pretending otherwise would fail our own standard.

What contains it is structure rather than promise. Scoring and artifact generation are separate functions, and the instrument never knows who is a customer. Scoring is identical regardless of payment — honest failing grades ship with receipts whether or not the graded party ever spends anything. Fix purchases never guarantee grade movement; artifacts go back through the same scan and get judged like anyone else's work. The methodology is fixed and published, the instruments are versioned, and the only thing that moves a grade is the documentation itself, verified by rescan. Independence, in other words, is engineered into the pipeline where it can be inspected, rather than asserted in the marketing where it can't.

What to do with any score you're handed

When an agent-readiness score lands on your desk — from anyone — ask for the methodology, the task bank, and the receipts, and check whether the grade decomposes into parts you can act on. A score that survives those questions is useful regardless of who produced it. A score that can't answer them is an ad.

Then do the work, which stays the same whoever grades you. Agent-readiness is a property of your public documentation surface: discoverability first — the machine-readable entry points that let an agent find a usable version of your docs at all — then comprehension. The highest-impact moves are well understood: a focused llms.txt, an AGENTS.md that briefs coding agents, a parseable OpenAPI spec, and the prioritized sequence that turns "make us AI-ready" into a sprint plan. None of those fixes belongs to any vendor, including us.

What grade does an AI agent give your API? Discry your API — free — 60 seconds, no signup.

See where your API stands.

Drop your docs URL. The scan probes the same signals this guide describes — in about a minute, free.

Discry your API — free