Glossary · DISCRY METHODOLOGY

Behavioral measurement

Behavioral measurement grades documentation by what models actually do with it: models are given the docs as their only source, quizzed on tasks a real integration depends on, and their answers are graded mechanically against ground truth cited to the documentation itself. It measures outcomes — whether a model could operate the API — rather than the presence of artifacts a checklist can count.

The approach exists because presence and function are different facts. A checklist can confirm an llms.txt file exists; it cannot confirm that a model reading it can build a correct request. Teams that shipped llms.txt or AGENTS.md files often did so with no way to verify the artifact worked — behavioral measurement closes that gap by testing the thing the artifact is for.

Mechanical grading is what makes the behavior a measurement rather than an opinion. Every ground-truth answer must trace to a specific location in the API's own documentation, and grading applies the same way to every API, so the result is reproducible and auditable rather than a judgment call about docs quality.

How Discry measures this

Behavioral measurement is the core of the Discry Score's comprehension axis. Real models are quizzed on the live, plain-fetched docs from versioned, published task banks; a separate closed-book pass establishes what they already knew, so only genuinely doc-dependent performance counts; and answers are graded mechanically, never by human judgment. The result publishes decomposed into discovery and comprehension, with the profile's coverage and failed tasks as receipts.

What does an AI agent make of your API?

Find out in about a minute — no signup.

Discry your API — free