# Crawler and Acceptable Use Policy

> Canonical: https://discry.ai/legal/acceptable-use · Markdown mirror: https://discry.ai/legal/acceptable-use.md

Effective 2026-09-22 · Version 1 · Version tag 2026-09-22.1

Operator: Discry LLC, a New York limited liability company. Legal contact: legal@discry.ai.

Operator: Discry LLC, a New York limited liability company, 418 Broadway, Suite 10855, Albany, NY 12207. Contact: legal@discry.ai.

Defined terms used: Discry, Operator, Site, Scan, Directory, Index, Band, Fix Kit, Account, Machine Access (defined in the Terms of Use); Takedown Notice, Correction Request, Rescan Request (defined in the Directory and Benchmark Policy); Index Data, Badge (defined in the API, MCP and Data License Terms). This policy defines one term of its own, **Exclusion List**, in section 9.1.

This policy is a notice incorporated by the Terms of Use, which control where the two differ.

## 1. What this policy covers

1.1 Part A (sections 2 to 9) describes what the Discry scanner does when it fetches other people's websites, as the code is, not as a well-behaved crawler is expected to be. Part B (sections 10 to 13) sets out what visitors and automated agents may not do on the Site.

## 2. How the scanner identifies itself

2.1 The scanner sends one of three user-agent strings.

- `DiscryScan/2.0`: the scan harness that produces graded results for the Index.
- `DiscryScan/2.0 (+https://discry.ai/methodology)`: the discovery-probe fetcher that runs when a visitor submits a URL to the free Scan.
- `discry-excerpt-fetch/1.0 (+https://discry.ai)`: the tool that refreshes the file excerpts described in section 7.2.

2.2 Only the harness string lacks a contact URL; this page and legal@discry.ai are the contact route for all three. Requests come from the Site's serverless functions (free Scan) and from a server Discry operates (scheduled Index scans). Discry publishes no IP range and offers no reverse-DNS verification; the user-agent string is the only identity signal.

## 3. What is fetched, and why

3.1 The scanner fetches public documentation URLs, starting from a URL a visitor submitted to the free Scan or from the documentation URL recorded for an API already in the Index.

3.2 From that starting point it requests the well-known files used to measure agent readiness, `robots.txt`, `llms.txt`, and `AGENTS.md`, plus documentation pages linked from the starting page; absolute links are extracted from fetched HTML for further probing. Discry does not run a site-wide crawl or enumerate a domain; no page-count limit is enforced in code.

3.3 Every fetch serves measurement of whether an AI agent can find and understand the documentation, and the evidence behind a published score. Discry does not fetch to build a search index, to republish documentation, or to train a model.

## 4. Public content only

4.1 The scanner sends no credentials, holds no session, submits no forms, and requests nothing an anonymous visitor could not request. It sends `Accept: text/html,application/json,text/plain,*/*` and follows redirects.

4.2 A 401, 403, 404, or 451 response is final. The scanner does not retry it or try a path around it, and records the target as blocked or unreachable instead of scoring it.

4.3 HTML is converted to text locally by removing scripts, styles, comments, and tags. No script on a fetched page is executed.

## 5. robots.txt: what the scanner does and does not do

5.1 The scanner fetches `/robots.txt` from every target and reads it to score it as a discovery signal: whether the `gptbot`, `claudebot`, and `google-extended` user agents are fully disallowed, converted to pass, partial, or fail.

5.2 The scanner does not currently use robots.txt to decide what to fetch. No `Disallow` rule, no `Crawl-delay` value, and no rule addressed to `DiscryScan` or to `*` changes which URLs it requests. This departs from RFC 9309, which expects a crawler to honour the rules a site publishes for it.

5.3 The reason is that robots.txt is one of the things being measured: an operator's choice to allow or disallow AI crawlers is a scored fact about that API's agent readiness. The `robotsTxt` signal weighs 3 of a 19-point discovery total. A scanner that stopped at a `Disallow` line could not report what the line says.

5.4 A robots.txt rule is therefore not the way to opt out; section 9 is. If the scanner changes to honour robots.txt, this policy is reissued under section 14.

5.5 Discry's own `robots.txt` allows every user agent across the whole Site.

## 6. Rate, timeouts, and retries

6.1 Each request has a 15-second timeout.

6.2 Only 429 and 503 responses are retried, after 2 seconds and then 6 seconds, honouring `Retry-After` up to a 30-second cap. A per-host circuit breaker stops all retries to a host that has exhausted the ladder three times.

6.3 There is no concurrency limit, request-per-second cap, or crawl-delay handling in the fetch layer, and Discry does not represent that its rate is capped. Volume is limited in practice by the fetch set in section 3.2 and an hourly Index refresh. An operator who observes a problematic rate has an immediate remedy that needs nothing from Discry: return 401, 403, 404, or 451 to the `DiscryScan` user agent, which the scanner treats as final and never retries (section 4.2). The operator should also write to legal@discry.ai, and Discry will place the domain on the Exclusion List under section 9.1, by hand until the list is built.

## 7. What is stored and what is displayed

7.1 Scan evidence. The evidence behind a score, including fetched text, is stored per API, and the traces and manifest behind each published report are stored with it. Scan reports are append-only by database rule: they cannot be edited in place, only removed at table level. No deletion job exists, so evidence is retained until Discry deletes it by hand or reissues the instrument.

7.2 Displayed excerpts. Discry caches the opening lines of certain third-party files, chiefly `llms.txt` and `AGENTS.md`, and shows them on its example and guide pages as illustrations: by default 8 lines, each cut at 160 characters. The cache is refreshed by hand, never at build time.

7.3 The basis for the excerpts is quotation for criticism, commentary, and evaluation: each is short, sits next to Discry's analysis, is limited to what the point requires, and links to the source. Every excerpt, and every piece of documentation text inside a published trace, is published as evidence, is attributed to its source, and is not licensed by Discry to anyone: it stays its author's property and sits outside the Index Data licence in the API, MCP and Data License Terms, as the Directory and Benchmark Policy sections 2.5 and 5.4 also state. Discry does not display full files, and removes an excerpt on request under section 9.4.

## 8. Processing by language-model providers

8.1 Fetched documentation is processed by third-party large language models to perform comprehension tasks from Discry's task bank. The enabled scan models are Claude models supplied by Anthropic, run under pinned identifiers so that a silent model change fails the scan rather than moving the score.

8.2 Discry trains no models and operates none of its own, so fetched content is never used to train a Discry model. Whether the provider retains what it processes is governed by Discry's agreement with the provider, and Discry makes no representation about it here.

8.3 A Fix Kit is generated by the Anthropic API from the scanned API's own evidence, with no personal data of the buyer in the prompt.

## 9. Opting out, exclusion, and removal

9.1 Two routes, and they do different things. The immediate route stops the fetch: a 401, 403, 404, or 451 response to the `DiscryScan` user agent stops the scanner at once, is never retried, and records the target as blocked or unreachable rather than scored (section 4.2). This is the route that stops Discry fetching your documentation, it needs nothing from us, and it takes effect as soon as your server is configured. The durable route stops the publication: the **Exclusion List** is the list of domains Discry does not publish. A domain owner or authorised representative asks by emailing legal@discry.ai with the domain and a way to verify authority over it (a reply from an address at the domain, or a file at its root). Discry verifies and adds the domain, and aims to acknowledge within ten business days.

9.2 Effect. Once a domain is on the Exclusion List, no Profile for it is published, and no row for it appears in the Directory, the public API, the machine exports, the MCP surface, or the Badge endpoint. An existing Profile stops being served. Exclusion does not stop a Scan running and does not delete what an earlier Scan recorded: it is a refusal to publish, not a refusal to measure. To stop the fetching itself, use the immediate route in section 9.1.

9.3 Exclusion removes the listing entirely. It is not a way to keep a listing while suppressing a low score; Correction Requests and Rescan Requests exist for a result you believe is wrong or out of date. Because exclusion does not stop collection, it is also not the route for erasing personal data or for a copyright complaint about material we hold: section 9.4 covers excerpt removal, section 9.7 and the Privacy Policy section 10 cover individuals, and the Directory and Benchmark Policy section 11 covers Takedown Notices.

9.4 Excerpt removal. The owner of a file quoted under section 7.2 may ask legal@discry.ai, giving the discry.ai page URL and the source URL, to remove the excerpt. Discry removes it and republishes the page. This does not require exclusion from scanning.

9.5 Copyright complaints about any content on the Site go through the Takedown Notice process in the Directory and Benchmark Policy.

9.6 Anyone may submit a public URL to the free Scan, including one for a site they do not own, and is responsible for that submission under the Terms of Use. Discry does not tell a domain owner who submitted a scan of their site.

9.7 Individuals. A person named or identifiable on a fetched page (an author name, a support address) may ask legal@discry.ai to erase their personal data from a displayed excerpt, a published report, or stored scan evidence, without authority over the domain. Discry removes it from the excerpt cache and the page, and deletes it from stored evidence by hand, because no deletion job exists; where a table is append-only, the record is anonymised or the rows removed at table level, as the Privacy Policy §9.1 and §10.5 describe. The Privacy Policy §10 sets out the full rights and timelines, and a request under this section is a request under that policy.

## 10. Prohibited technical conduct

Visitors, Account holders, and automated agents must not:

10.1 Send request volumes that degrade the Site for others. The public Machine Access surfaces carry no rate limit, authentication, or quota today ; the absence of a technical limit is not permission to exhaust the service.

10.2 Attempt to sign in as another person, guess or replay magic links or delivery tokens, or otherwise reach an Account, purchase, or Fix Kit that is not theirs.

10.3 Interfere with the Site's scheduled jobs, internal endpoints, or workers.

10.4 Probe or test the Site for vulnerabilities, or attempt to defeat any access control, except as section 12 allows.

## 11. Content, commercial, and anti-gaming restrictions

11.1 The restrictions in sections 8 and 9 of the Terms of Use (including what may be submitted to the free Scan, claiming scans, use of Badges and Bands, and the held-out task bank) apply to every use of the Site and are incorporated here without restatement.

## 12. Security research

12.1 Discry welcomes good-faith security research on the Site: test only your own Account and data, access or retain no one else's, no denial of service, no social engineering, and report to legal@discry.ai before any public disclosure. A researcher who follows those rules will not face a legal claim from Discry for that research.

12.2 Out of bounds: the internal worker endpoints, the payment flow beyond your own test purchase, third-party providers Discry uses, any site the scanner fetches, and the held-out task bank and canary strings, which the Terms of Use protect.

## 13. Enforcement and reporting

13.1 Discry may respond to a breach by warning, blocking a user agent, address, or Account, removing a claimed scan, refusing a purchase, or terminating an Account under the Terms of Use, and may act without notice for severe conduct such as section 10.2 or a denial-of-service attempt.

13.2 Discry has no automated enforcement today: no rate limiter on public surfaces and no suspension code path. Enforcement is a manual action by the Operator.

13.3 To report abuse, or a scanner request you believe breached Part A, email legal@discry.ai with the user agent or Account involved, URLs, UTC timestamps, and what you observed. Discry acknowledges reports and does not promise an outcome.

## 14. Disputes, changes, and archive

14.1 Any dispute about this policy is governed by section 14 of the Terms of Use.

14.2 Discry may change this policy. Each change increments the version and effective date at the top of the page, and the previous version is kept unchanged in Discry's legal archive and is available on request to legal@discry.ai. A change to the scanner behaviour in sections 5 or 6 is reflected here before or when it deploys.
