Web Crawlers
What Googlebot, ClaudeBot, GPTBot and PerplexityBot actually receive from this site — the policy verdict, the crawler document, and the headers, run in-process against the real page registry.
The problem, stated once
A Dash app ships an empty <div>. Everything a reader sees is assembled by JavaScript afterwards. Google runs JavaScript — eventually, on a second pass, budget permitting. Most crawlers do not run it at all, and neither do the AI fetchers that decide whether your package gets cited.
So the document a crawler stores for your page is a loading spinner, and nothing anywhere reports this. The page looks perfect to you.
Showcase A — what the crawler sees
Pick a page and a User-Agent. The verdict, the document and the headers below are produced by running the package's own pure handlers in this process against this site's real page registry — not by a mock, and not by a description of what they would do.
# Live component, rendered above on the browser lane.
# Source: docs/crawler_view/crawler_view.py
"""Showcase A — what a crawler actually receives from this site.
Runs the package's own PURE handlers in-process for a chosen page and User-
Agent, and shows the three things that are normally invisible: the policy
verdict, the crawler document's markup, and the headers that came with it —
next to what a browser gets for the same URL.
Exec-module rules (both tested elsewhere in this repo):
1. **Globally-unique id prefix** — everything here starts `crwv-`. Ids share
one namespace across every exec module on the site.
2. **No import-time registry walk** — this module is imported from inside
`pages/markdown.py`'s glob loop, so `dash.page_registry` is incomplete.
`component` is a placeholder; the page list arrives by callback.
"""
import html as _html
from dash import Input, Output, State, callback, dcc, html, no_update
import dash_mantine_components as dmc
from dash_iconify import DashIconify
ID = "crwv"
BROWSER_UA = (
"Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/537.36 "
"(KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36"
)
_VERDICT_COLOR = {"allow": "teal", "block": "red", "meter": "yellow"}
_CLASS_COLOR = {"training": "red", "search": "teal", "traditional": "blue"}
def _ua_options():
"""Real UA tokens grouped by class, straight from the vendor registry.
Runs at LAYOUT time, so it depends on `vendors` existing — which the
>=2.7.0 floor in requirements.txt guarantees. It carried a pre-2.7.0
fallback while that floor was still 2.6.1; the floor moved on
2026-08-23 and the fallback went with it.
"""
from dash_improve_my_llms.vendors import VENDORS
groups = {"training": [], "search": [], "traditional": []}
for vendor in VENDORS:
# Not every vendor publishes a robots token: `anthropic-legacy` is a
# UA-only identity (anthropic-ai, claude-web) with an empty
# robots_tokens tuple, so fall back to what the middleware matches on.
token = (vendor.robots_tokens or vendor.ua_tokens or (vendor.key,))[0]
groups.setdefault(vendor.cls, []).append(
{"value": f"{token}/1.0",
"label": f"{vendor.display} — {vendor.operator}"}
)
data = [{"group": "A human's browser",
"items": [{"value": BROWSER_UA, "label": "Chrome 120 (a person)"}]}]
for cls, label in (("training", "AI training crawlers"),
("search", "AI search & citation"),
("traditional", "Traditional search engines")):
if groups.get(cls):
data.append({"group": label, "items": sorted(groups[cls],
key=lambda o: o["label"])})
return data
component = html.Div(
[
dmc.Group(
[
dmc.Select(
id=f"{ID}-page",
label="Page",
placeholder="Loading the registry…",
searchable=True,
w=320,
),
dmc.Select(
id=f"{ID}-ua",
label="User-Agent",
data=_ua_options(),
value=BROWSER_UA,
searchable=True,
w=380,
),
],
gap="lg",
align="flex-start",
),
dmc.Space(h="md"),
html.Div(id=f"{ID}-verdict"),
dmc.Space(h="md"),
dmc.Tabs(
[
dmc.TabsList(
[
dmc.TabsTab("Rendered", value="rendered",
leftSection=DashIconify(icon="tabler:eye", width=15)),
dmc.TabsTab("View source", value="source",
leftSection=DashIconify(icon="tabler:code", width=15)),
dmc.TabsTab("Headers", value="headers",
leftSection=DashIconify(icon="tabler:list-details", width=15)),
]
),
dmc.TabsPanel(html.Div(id=f"{ID}-rendered"), value="rendered", pt="md"),
dmc.TabsPanel(html.Div(id=f"{ID}-source"), value="source", pt="md"),
dmc.TabsPanel(html.Div(id=f"{ID}-headers"), value="headers", pt="md"),
],
value="rendered",
id=f"{ID}-tabs",
),
dcc.Interval(id=f"{ID}-boot", interval=200, max_intervals=1),
]
)
@callback(
Output(f"{ID}-page", "data"),
Output(f"{ID}-page", "value"),
Input(f"{ID}-boot", "n_intervals"),
)
def _pages(_tick):
import dash
options = sorted(
({"value": e["path"], "label": e.get("name") or e["path"]}
for e in dash.page_registry.values()
if not e["path"].startswith("/admin/")),
key=lambda o: o["label"],
)
if not options:
return no_update, no_update
return options, "/"
@callback(
Output(f"{ID}-verdict", "children"),
Output(f"{ID}-rendered", "children"),
Output(f"{ID}-source", "children"),
Output(f"{ID}-headers", "children"),
Input(f"{ID}-page", "value"),
Input(f"{ID}-ua", "value"),
)
def _show(path, user_agent):
if not path or not user_agent:
return None, None, None, None
import dash
from dash_improve_my_llms.bot_detection import get_bot_type, is_any_bot
is_bot = is_any_bot(user_agent)
bot_class = get_bot_type(user_agent) if is_bot else "human"
from dash_improve_my_llms.bot_detection import get_bot_vendor
from dash_improve_my_llms.vendors import VENDORS, effective_policies
vendor = next((v for v in VENDORS if v.key == get_bot_vendor(user_agent)), None)
policy = "allow"
if vendor is not None:
try:
policy = effective_policies(
getattr(dash.get_app(), "_robots_config", None)
)[vendor.key]
except Exception:
# The fold reads a live config object; a showcase must degrade
# rather than break the page it is embedded in.
policy = "allow"
verdict = _verdict_card(is_bot, vendor, bot_class, policy)
if is_bot and policy == "block":
blocked = dmc.Alert(
[
dmc.Text("403 Forbidden", fw=700, size="lg"),
dmc.Text(
"This crawler is blocked on ordinary pages by the site's "
"vendor policy, so it never reaches the document below. "
"The documentation surfaces (/llms.txt and the tier "
"documents) stay open to it — that is deliberate, and it "
"is what `block_ai_training_docs` would close.",
size="sm",
),
],
color="red",
variant="light",
icon=DashIconify(icon="tabler:ban"),
)
return verdict, blocked, blocked, _headers_table(
{"Status": "403 Forbidden", "Cache-Control": "no-store"}
)
document = _crawler_document(path)
if not document:
empty = dmc.Alert(
"No prerendered document is available for this page in-process.",
color="yellow", variant="light",
)
return verdict, empty, empty, None
rendered = html.Iframe(
srcDoc=document,
style={"width": "100%", "height": "480px", "border": "1px solid var(--mantine-color-default-border)",
"borderRadius": "8px", "background": "white"},
)
source = dmc.Paper(
dmc.Code(_truncate(document), block=True,
style={"maxHeight": "480px", "overflow": "auto", "display": "block"}),
withBorder=True, radius="md", p="sm",
)
headers = _headers_table({
"Status": "200 OK",
"Content-Type": "text/html; charset=utf-8",
"X-Robots-Tag": "(none — this page is indexable)",
"Vary": "Accept (on /<page>/llms.txt)",
})
return verdict, rendered, source, headers
def _verdict_card(is_bot, vendor, bot_class, policy):
rows = [
("Classified as", "bot" if is_bot else "human (untouched by policy)"),
("Vendor", f"{vendor.display} — {vendor.operator}" if vendor else "—"),
("Class", bot_class),
("Effective policy", policy if is_bot else "n/a"),
]
return dmc.Paper(
dmc.Group(
[
dmc.Stack(
[
dmc.Text(label, size="xs", c="dimmed", tt="uppercase"),
dmc.Badge(
str(value),
color=_VERDICT_COLOR.get(value, _CLASS_COLOR.get(value, "gray")),
variant="light",
size="lg",
) if label in ("Class", "Effective policy")
else dmc.Text(str(value), size="sm", fw=600),
],
gap=4,
)
for label, value in rows
],
gap="xl",
),
withBorder=True, radius="md", p="md",
)
def _headers_table(headers):
return dmc.Table(
[
dmc.TableThead(dmc.TableTr([dmc.TableTh("Header"), dmc.TableTh("Value")])),
dmc.TableTbody([
dmc.TableTr([dmc.TableTd(dmc.Code(k)), dmc.TableTd(v)])
for k, v in headers.items()
]),
],
striped=True, withTableBorder=True,
)
def _crawler_document(path):
"""The crawler HTML for `path`, built by the package's own generator.
Calls the pure function rather than re-implementing it, so this showcase
cannot claim something the real route would not serve.
"""
try:
import dash
import dash_improve_my_llms as dimll
from dash_improve_my_llms.html_generator import generate_static_page_html
store = getattr(getattr(dimll, "_state", None), "page_metadata", None) or {}
meta = store.get(path)
if not meta:
return ""
all_pages = [
{"path": p, "name": (m or {}).get("name") or p}
for p, m in store.items()
]
app = dash.get_app()
app_config = {
"name": getattr(app, "title", "") or "",
"base_url": getattr(app, "_base_url", "") or "",
}
return generate_static_page_html(path, meta, all_pages, app_config)
except Exception as exc: # the showcase degrades, it never breaks the page
return (
"<!doctype html><meta charset='utf-8'>"
f"<p>Crawler document unavailable in-process: {_html.escape(str(exc))}</p>"
)
def _truncate(text, limit=12000):
return text if len(text) <= limit else text[:limit] + "\n… truncated …"
Three things worth trying:
- switch between Chrome and ClaudeBot on any page — the verdict flips
to block and the crawler gets a 403 before the document is ever built;
- switch to Googlebot and open View source — that is the prerendered
document, with real prose and real links, in the initial HTML;
- switch to PerplexityBot — allowed, because AI search citations are how
people find a package, while AI training is a different bargain.
The policy behind the verdict
One vendor registry drives both halves, so the site cannot say one thing and do another:
| Class | Default | Why |
|---|---|---|
| AI training (GPTBot, ClaudeBot, CCBot, …) | block | robots.txt was the published promise; 2.7.0 made the middleware keep it |
| AI search (Claude-User, ChatGPT-User, PerplexityBot, …) | allow | citations send readers; that is the point of docs |
| Traditional (Googlebot, Bingbot, …) | allow | ordinary search |
app._robots_config = RobotsConfig(
block_ai_training=True,
allow_ai_search=True,
allow_traditional=True,
crawl_delay=10,
)
Since 2.7.0 the coarse flags have a finer companion — vendor_policy, per vendor, allow / block / meter — and it takes a callable, read on every request. That is what the control board writes through.
The documentation surfaces stay open
Note what did not 403 above: /llms.txt, /llms-small.txt, /llms-full.txt and every /<page>/llms.txt answer a blocked training crawler with 200. That is deliberate. The corpus exists to get the package used, and an upgrade must not silently start refusing it.
RobotsConfig(block_ai_training_docs=True) closes that half if you want it closed. It is opt-in on purpose.
Live policy
This site's robots.txt and sitemap.xml are generated, never hand-written, and are always open to everyone — including crawlers the site blocks elsewhere. A crawler that cannot read the rules is a crawler that has no rules:
/robots.txt— generated from the vendor registry/sitemap.xml— with priority inferred from the page tree/llms.txt— the index, plus the cross-host network directory
Hidden pages
mark_hidden("/admin/control-board") removes a path from the sitemap, 404s its llms.txt, skips the MCP resource, and returns 404 to crawler requests on the page URL. It is not in robots.txt — a Disallow line publishes the path it is trying to protect.
Source: /audiences/web-crawlers
Note for AI agents: This is the static, prerendered view of an interactive Dash application served because we detected a non-JS user agent. Full prose docs:
- /audiences/web-crawlers/llms.txt — LLM-friendly documentation
- /sitemap.xml
- /robots.txt