Maturity model
The model is the spec.
Last Mile grades Agent Experience (AX) against the maturity levels below. This page is not a written summary of the model — it renders the same versioned data files the engine scores from, so the published spec and the running benchmark cannot drift apart.
Find. Understand. Operate. Transact.
Each maturity level asks one question about what an agent can do. A level is decided by measured signals — deterministic probes over a real browser-agent run — and never by what a site declares about itself. Open a level for its full spec.
-
Discoverability Do agents find you? Measured today
What the level means
Agents find your site and your capabilities — your content as well as the tools and actions you expose.
What decides it
A real agent is prompted for the task — does the airline (or its function) surface and get reached at all? If not, the journey never starts.
Measured signals 1
Each signal is produced by one deterministic probe. The agent drives the browser; the probe decides what the run showed.
- Discoverability probe: discoverability
Can the agent find and reach the relevant page or function at all (semantic search, structured nav, capability discovery).
Example tasks 20
Every measured task starts at this level — the agent is handed the request and has to reach the capability before anything else can be scored.
- Find a destination reachable from a hub within a stated budget Public
- Find the cheapest day to fly in the low-fare / flexible-dates calendar Public
- Get a destination idea matching a stated intent Public
- Read the size of the airline's route network from the where-we-fly surface Public
Showing 4 of 20 measured tasks that reach this level.
Mapped playbooks 2
The optimization ladder has three rungs. Tier 2 is the frontier — it earns its place only once Tier 0 and Tier 1 are clean.
- Add semantic markup and structured navigation to expose entry points Tier 1 · Publish the facts Discoverability
Mark up navigation landmarks with ARIA roles, use descriptive link text matching common task keywords, and add structured data (JSON-LD) so agents identify the right entry point without guessing.
- Publish an llms.txt capability index Tier 1 · Publish the facts Discoverability
Serve /llms.txt describing the site's key pages and capabilities in Markdown so agents can orient without crawling.
Tracked standards 8
Standards shape the ceiling of what a site can offer an agent. A declared standard never earns score credit by itself — only a measured signal does.
- robots.txt Established
Tells crawlers and AI bots which paths they may fetch. · IETF · RFC 9309
- Sitemap Established
An XML list of your URLs so crawlers can find every page. · sitemaps.org
- Link headers Established
Declares related resources in the HTTP Link header — no HTML parse needed. · IETF · RFC 8288
- DNS-AID Experimental
Publish SVCB DNS records so agents discover your services before any HTTP call. Early / experimental. Impossible on unsigned zones like *.workers.dev. · IETF draft · community N/A on *.workers.dev (unsigned zone) — record as N/A, never as a failure.
- MCP Server Card Experimental
A discovery card advertising the MCP tools your site exposes. Very early — few sites use it. · Anthropic · MCP (SEP-1649)
- API Catalog Emerging
A machine-readable index of your APIs at /.well-known/api-catalog. · IETF · RFC 9727
- OAuth discovery Established
Lets agents auto-discover your authorization server's endpoints. · IETF · RFC 8414
- Web Bot Auth Experimental
Agents cryptographically sign requests (built on RFC 9421) so you can verify and allow trusted bots. Draft, but shipping. · Cloudflare + Google · IETF draft
See what this level costs sites in the live Last Mile Index →
-
-
Understanding Can agents read & navigate you? Measured today
What the level means
An agent can read, parse and comprehend you — offers, rules, returned results — accessibility plus structure, with no DOM-wrangling or PDF-locked content.
What decides it
Can the agent parse offers and read the inputs and errors, or is content locked in JS / PDFs? Structured path, input legibility and error legibility are the tells.
Measured signals 3
Each signal is produced by one deterministic probe. The agent drives the browser; the probe decides what the run showed.
- Structured path probe: structuredPath Also under Interaction
Whether the site exposes AND verifiably executes a WebMCP / agent tool path (window.__webmcp or navigator.modelContext): the probe invokes a declared read-only tool and confirms it returns a correct, machine-readable answer — vs blind DOM driving. Verified is a site measurement, not a task grade. Spans two levels: a declared, readable tool surface credits Understanding, while a verifiably executed tool call is the agent operating the site and credits Interaction — where the WebMCP standard and practice are filed.
This signal spans two maturity levels. It is the same measurement in both places — counted once, credited under Understanding and Interaction, not as two independent signals.
- Input affordances probe: inputAffordances
Whether inputs are agent-readable controls vs hostile readonly/hidden date widgets and custom overlays that reject programmatic input.
- Failure transparency probe: failureTransparency
What failure surface the page showed at the end of the run: the signal is ON when a visible error (role='alert' / aria-invalid / an error class) or a 'no results' dead-end was present. Presence, not absence — a run that fails against a silent page leaves this off and is diagnosed as task completion. The probe observes; it never induces a failure.
Example tasks 15
Tasks whose verified answer is published information the agent has to read and interpret.
- Find a destination reachable from a hub within a stated budget Public
- Find the cheapest day to fly in the low-fare / flexible-dates calendar Public
- Get a destination idea matching a stated intent Public
- Read the size of the airline's route network from the where-we-fly surface Public
Showing 4 of 15 measured tasks that reach this level.
Mapped playbooks 3
The optimization ladder has three rungs. Tier 2 is the frontier — it earns its place only once Tier 0 and Tier 1 are clean.
- Replace hostile date widgets with agent-readable inputs Tier 0 · Remove the blocker Input affordances
Remove readonly/hidden date fields and custom datepicker overlays that reject programmatic input; use a standard <input type="date"> or publish the fact as readable text so an agent never needs to fight the widget.
- Surface real validation errors instead of silent dead-ends Tier 0 · Remove the blocker Failure transparency
When a form submission fails, return a visible, machine-readable error message — role="alert" or aria-invalid — so the agent knows the attempt failed and why.
- Publish core info as agent-readable content Tier 0 · Remove the blocker Task completion
Ensure the answer to the task is available as readable static content (a text table, a paragraph) so an agent can extract it without navigating a transactional flow.
Tracked standards 3
Standards shape the ceiling of what a site can offer an agent. A declared standard never earns score credit by itself — only a measured signal does.
- Markdown negotiation Emerging
Serve clean Markdown when an agent sends Accept: text/markdown (~80% fewer tokens). Beta, ~4% adoption. · Cloudflare · HTTP content negotiation
- schema.org / JSON-LD Established
Structured-data vocabulary that makes offers, fares and policies machine-readable. Established. · schema.org · W3C (JSON-LD)
- NDC Established
Airline XML standard for offers and orders — most carriers already have it. · IATA
See what this level costs sites in the live Last Mile Index →
-
-
Interaction Can agents operate you? (unpaid) Measured today
What the level means
The agent actively operates you — supplies inputs, runs searches, fills forms and performs account actions that DON'T move money. Above the money line.
What decides it
Does the agent complete a search, check-in or selection end to end — past any consent wall — with no human handoff? Unpaid task completion is the tell.
Measured signals 3
Each signal is produced by one deterministic probe. The agent drives the browser; the probe decides what the run showed.
- Structured path probe: structuredPath Also under Understanding
Whether the site exposes AND verifiably executes a WebMCP / agent tool path (window.__webmcp or navigator.modelContext): the probe invokes a declared read-only tool and confirms it returns a correct, machine-readable answer — vs blind DOM driving. Verified is a site measurement, not a task grade. Spans two levels: a declared, readable tool surface credits Understanding, while a verifiably executed tool call is the agent operating the site and credits Interaction — where the WebMCP standard and practice are filed.
This signal spans two maturity levels. It is the same measurement in both places — counted once, credited under Interaction and Understanding, not as two independent signals.
- Consent wall probe: consentWall
Whether a consent/cookie overlay intercepts agent pointer or keyboard focus and blocks the core task flow.
- Task completion probe: answerExtracted Also under Transaction
The verified binary outcome — did the agent surface an answer that passes the task's deterministic success check. The catch-all when no specific signal fired.
This signal spans two maturity levels. It is the same measurement in both places — counted once, credited under Interaction and Transaction, not as two independent signals.
Example tasks 5
Tasks where the agent has to operate the site — run a search, configure a flow, reach the review screen — without moving money.
- Check whether a sustainable fare or seat selection shows a clear price delta before payment Booking flow
- Compare the price of an add-on now versus later Booking flow
- Get a recommended bundle and its price in the booking flow Booking flow
- Reach a validated checkout total before payment Booking flow
Showing 4 of 5 measured tasks that reach this level.
Mapped playbooks 2
The optimization ladder has three rungs. Tier 2 is the frontier — it earns its place only once Tier 0 and Tier 1 are clean.
- Make consent dismissible without blocking agent interaction Tier 0 · Remove the blocker Consent wall
Ensure the consent/cookie banner can be dismissed via a standard accessible button (not a pointer-only overlay) and does not capture keyboard or programmatic focus from the core task flow.
- Expose WebMCP tools for true agent workflows Tier 2 · Offer the tool Structured path
Implement window.__webmcp (or a link[rel='mcp'] manifest) exposing typed tool definitions for core tasks (search, booking) so agents invoke structured operations instead of driving the DOM blindly.
Tracked standards 3
Standards shape the ceiling of what a site can offer an agent. A declared standard never earns score credit by itself — only a measured signal does.
- WebMCP Experimental
Web pages expose typed tools a browser agent calls directly, skipping the DOM. Experimental — Chrome origin trial, Gemini-only. · Google & Microsoft · W3C CG
- MCP (actions) Emerging
Open protocol connecting agents to your tools and actions. Broadly adopted; ~8 months old. · Anthropic
- OAuth Protected Resource Established
Metadata telling agents how to authenticate to your protected APIs. · IETF · RFC 9728
See what this level costs sites in the live Last Mile Index →
-
- The money line Everything above is unpaid. Below it, an agent moves money — the benchmark stops here today.
-
Transaction Can agents pay? (the highest bar) Beyond the money line
What the level means
The agent completes money-moving actions on the customer's behalf — with payment, agent auth and trust. Depends on Interaction. Below the money line.
What decides it
Does a verified agent complete a real paid booking end to end — payment cleared — with no human in the loop? Paid task completion.
Measured signals 1
Each signal is produced by one deterministic probe. The agent drives the browser; the probe decides what the run showed.
- Task completion probe: answerExtracted Also under Interaction
The verified binary outcome — did the agent surface an answer that passes the task's deterministic success check. The catch-all when no specific signal fired.
This signal spans two maturity levels. It is the same measurement in both places — counted once, credited under Transaction and Interaction, not as two independent signals.
Example tasks 0
No task in the current suite crosses the money line. The benchmark stops before payment, so this level has no measured task today.
Each run of a measured task returns one outcome: Completed, Failed or Blocked.
Mapped playbooks 1
The optimization ladder has three rungs. Tier 2 is the frontier — it earns its place only once Tier 0 and Tier 1 are clean.
- Support agentic payment rails (x402 · MPP · UCP · ACP) Tier 2 · Offer the tool Task completion
Adopt an emerging agent-payment protocol so an agent can complete a purchase under delegated authority. Frontier — earns its place only after Tier 0/1 are clean.
Tracked standards 6
Standards shape the ceiling of what a site can offer an agent. A declared standard never earns score credit by itself — only a measured signal does.
- ACP Emerging
Agentic Commerce Protocol — open checkout standard any ACP agent can pay. Beta; powers ChatGPT Instant Checkout. · OpenAI + Stripe
- UCP Emerging
Universal Commerce Protocol — merchants declare commerce capabilities agents discover and transact. Live in Google AI Mode / Gemini. · Google + Shopify
- x402 Experimental
Revives HTTP 402 for instant programmatic (stablecoin) payments. Young and volatile — optionality, not a baseline. · Coinbase · x402 Foundation
- MPP Experimental
Machine Payments Protocol — multi-rail agent payments over HTTP 402 (Stripe Agentic Commerce Suite). New. · Stripe + Tempo
- AP2 Emerging
Agent Payments Protocol — agents authorize payments via a cryptographically signed Mandate. Composes with UCP. · Google
- NDC OrderCreate Established
The NDC message that creates a booking / order programmatically. · IATA · NDC
-
How levels become published results Each task run returns Completed, Failed or Blocked. Blocked is an honest access wall, not a failure, and is excluded from the Index. A whole website is then graded on the scale below, derived from its measurable runs with an evidence floor of 2 measurable tasks.
- Fails
Measured tasks reach no verified completion.
- Struggles
Some tasks complete; material friction remains.
- Completes
The measured set completes, above the evidence floor.
- Gated
Every task hit an access wall; nothing measurable to grade.
Orthogonal axis
Access tiers are a gate, not a fifth level.
A maturity level asks what an agent can do. An access tier asks what gate it has to pass first. The two axes are independent: a site can be excellent at Understanding behind a login wall, and a public page can still break at Interaction. The tier is what makes an outcome Blocked instead of Failed.
- Public Measurable today
To pass the gate: nothing — drives end to end
- Booking flow Measurable today
To pass the gate: nothing — drives to the payment wall, then stops
- Booking ref Not yet measurable
To pass the gate: a real PNR + last name on the carrier (cheapest gated proof)
- Login Not yet measurable
To pass the gate: a test frequent-flyer account per carrier
- Eligibility Not yet measurable
To pass the gate: a real disruption/IRROPS booking — best-effort only, can't be staged on demand
- Hybrid Not yet measurable
To pass the gate: depends on the flow
- Staff Not yet measurable
To pass the gate: n/a — not customer-facing
The benchmark can measure 2 of 7 tiers today. The rest are named so the boundary stays visible instead of being quietly rounded down to a failure.
Where this page comes from.
Nothing on this page is authored copy about the model — it is rendered from the versioned core objects below at build time. If the engine's contract changes, this page changes with it.
| Data file |
|---|
data/maturity-levels.json |
data/dimensions.json |
data/access-tiers.json |
data/best-practices.json |
data/standards.json |
data/use-cases.json |
data/tasks/airline.tasks.json |