LAST MILE Where agents actually land

Maturity model

The model is the spec.

Last Mile grades Agent Experience (AX) against the maturity levels below. This page is not a written summary of the model — it renders the same versioned data files the engine scores from, so the published spec and the running benchmark cannot drift apart.

Find. Understand. Operate. Transact.

Each maturity level asks one question about what an agent can do. A level is decided by measured signals — deterministic probes over a real browser-agent run — and never by what a site declares about itself. Open a level for its full spec.

  1. Discoverability Do agents find you? Measured today

    What the level means

    Agents find your site and your capabilities — your content as well as the tools and actions you expose.

    What decides it

    A real agent is prompted for the task — does the airline (or its function) surface and get reached at all? If not, the journey never starts.

    Measured signals 1

    Each signal is produced by one deterministic probe. The agent drives the browser; the probe decides what the run showed.

    • Discoverability probe: discoverability

      Can the agent find and reach the relevant page or function at all (semantic search, structured nav, capability discovery).

    Example tasks 20

    Every measured task starts at this level — the agent is handed the request and has to reach the capability before anything else can be scored.

    • Find a destination reachable from a hub within a stated budget Public

      Inspiration & Planning · task id explore-destinations-within-budget

    • Find the cheapest day to fly in the low-fare / flexible-dates calendar Public

      Inspiration & Planning · task id airline-flex-fare-calendar

    • Get a destination idea matching a stated intent Public

      Inspiration & Planning · task id inspiration-cards-by-intent

    • Read the size of the airline's route network from the where-we-fly surface Public

      Inspiration & Planning · task id airline-network-destinations

    Showing 4 of 20 measured tasks that reach this level.

    Mapped playbooks 2

    The optimization ladder has three rungs. Tier 2 is the frontier — it earns its place only once Tier 0 and Tier 1 are clean.

    • Add semantic markup and structured navigation to expose entry points Tier 1 · Publish the facts Discoverability

      Mark up navigation landmarks with ARIA roles, use descriptive link text matching common task keywords, and add structured data (JSON-LD) so agents identify the right entry point without guessing.

    • Publish an llms.txt capability index Tier 1 · Publish the facts Discoverability

      Serve /llms.txt describing the site's key pages and capabilities in Markdown so agents can orient without crawling.

    Tracked standards 8

    Standards shape the ceiling of what a site can offer an agent. A declared standard never earns score credit by itself — only a measured signal does.

    • robots.txt Established

      Tells crawlers and AI bots which paths they may fetch. · IETF · RFC 9309

    • Sitemap Established

      An XML list of your URLs so crawlers can find every page. · sitemaps.org

    • Link headers Established

      Declares related resources in the HTTP Link header — no HTML parse needed. · IETF · RFC 8288

    • DNS-AID Experimental

      Publish SVCB DNS records so agents discover your services before any HTTP call. Early / experimental. Impossible on unsigned zones like *.workers.dev. · IETF draft · community N/A on *.workers.dev (unsigned zone) — record as N/A, never as a failure.

    • MCP Server Card Experimental

      A discovery card advertising the MCP tools your site exposes. Very early — few sites use it. · Anthropic · MCP (SEP-1649)

    • API Catalog Emerging

      A machine-readable index of your APIs at /.well-known/api-catalog. · IETF · RFC 9727

    • OAuth discovery Established

      Lets agents auto-discover your authorization server's endpoints. · IETF · RFC 8414

    • Web Bot Auth Experimental

      Agents cryptographically sign requests (built on RFC 9421) so you can verify and allow trusted bots. Draft, but shipping. · Cloudflare + Google · IETF draft

  2. Understanding Can agents read & navigate you? Measured today

    What the level means

    An agent can read, parse and comprehend you — offers, rules, returned results — accessibility plus structure, with no DOM-wrangling or PDF-locked content.

    What decides it

    Can the agent parse offers and read the inputs and errors, or is content locked in JS / PDFs? Structured path, input legibility and error legibility are the tells.

    Measured signals 3

    Each signal is produced by one deterministic probe. The agent drives the browser; the probe decides what the run showed.

    • Structured path probe: structuredPath Also under Interaction

      Whether the site exposes AND verifiably executes a WebMCP / agent tool path (window.__webmcp or navigator.modelContext): the probe invokes a declared read-only tool and confirms it returns a correct, machine-readable answer — vs blind DOM driving. Verified is a site measurement, not a task grade. Spans two levels: a declared, readable tool surface credits Understanding, while a verifiably executed tool call is the agent operating the site and credits Interaction — where the WebMCP standard and practice are filed.

      This signal spans two maturity levels. It is the same measurement in both places — counted once, credited under Understanding and Interaction, not as two independent signals.

    • Input affordances probe: inputAffordances

      Whether inputs are agent-readable controls vs hostile readonly/hidden date widgets and custom overlays that reject programmatic input.

    • Failure transparency probe: failureTransparency

      What failure surface the page showed at the end of the run: the signal is ON when a visible error (role='alert' / aria-invalid / an error class) or a 'no results' dead-end was present. Presence, not absence — a run that fails against a silent page leaves this off and is diagnosed as task completion. The probe observes; it never induces a failure.

    Example tasks 15

    Tasks whose verified answer is published information the agent has to read and interpret.

    • Find a destination reachable from a hub within a stated budget Public

      Inspiration & Planning · task id explore-destinations-within-budget

    • Find the cheapest day to fly in the low-fare / flexible-dates calendar Public

      Inspiration & Planning · task id airline-flex-fare-calendar

    • Get a destination idea matching a stated intent Public

      Inspiration & Planning · task id inspiration-cards-by-intent

    • Read the size of the airline's route network from the where-we-fly surface Public

      Inspiration & Planning · task id airline-network-destinations

    Showing 4 of 15 measured tasks that reach this level.

    Mapped playbooks 3

    The optimization ladder has three rungs. Tier 2 is the frontier — it earns its place only once Tier 0 and Tier 1 are clean.

    • Replace hostile date widgets with agent-readable inputs Tier 0 · Remove the blocker Input affordances

      Remove readonly/hidden date fields and custom datepicker overlays that reject programmatic input; use a standard <input type="date"> or publish the fact as readable text so an agent never needs to fight the widget.

    • Surface real validation errors instead of silent dead-ends Tier 0 · Remove the blocker Failure transparency

      When a form submission fails, return a visible, machine-readable error message — role="alert" or aria-invalid — so the agent knows the attempt failed and why.

    • Publish core info as agent-readable content Tier 0 · Remove the blocker Task completion

      Ensure the answer to the task is available as readable static content (a text table, a paragraph) so an agent can extract it without navigating a transactional flow.

    Tracked standards 3

    Standards shape the ceiling of what a site can offer an agent. A declared standard never earns score credit by itself — only a measured signal does.

    • Serve clean Markdown when an agent sends Accept: text/markdown (~80% fewer tokens). Beta, ~4% adoption. · Cloudflare · HTTP content negotiation

    • schema.org / JSON-LD Established

      Structured-data vocabulary that makes offers, fares and policies machine-readable. Established. · schema.org · W3C (JSON-LD)

    • NDC Established

      Airline XML standard for offers and orders — most carriers already have it. · IATA

  3. Interaction Can agents operate you? (unpaid) Measured today

    What the level means

    The agent actively operates you — supplies inputs, runs searches, fills forms and performs account actions that DON'T move money. Above the money line.

    What decides it

    Does the agent complete a search, check-in or selection end to end — past any consent wall — with no human handoff? Unpaid task completion is the tell.

    Measured signals 3

    Each signal is produced by one deterministic probe. The agent drives the browser; the probe decides what the run showed.

    • Structured path probe: structuredPath Also under Understanding

      Whether the site exposes AND verifiably executes a WebMCP / agent tool path (window.__webmcp or navigator.modelContext): the probe invokes a declared read-only tool and confirms it returns a correct, machine-readable answer — vs blind DOM driving. Verified is a site measurement, not a task grade. Spans two levels: a declared, readable tool surface credits Understanding, while a verifiably executed tool call is the agent operating the site and credits Interaction — where the WebMCP standard and practice are filed.

      This signal spans two maturity levels. It is the same measurement in both places — counted once, credited under Interaction and Understanding, not as two independent signals.

    • Consent wall probe: consentWall

      Whether a consent/cookie overlay intercepts agent pointer or keyboard focus and blocks the core task flow.

    • Task completion probe: answerExtracted Also under Transaction

      The verified binary outcome — did the agent surface an answer that passes the task's deterministic success check. The catch-all when no specific signal fired.

      This signal spans two maturity levels. It is the same measurement in both places — counted once, credited under Interaction and Transaction, not as two independent signals.

    Example tasks 5

    Tasks where the agent has to operate the site — run a search, configure a flow, reach the review screen — without moving money.

    • Check whether a sustainable fare or seat selection shows a clear price delta before payment Booking flow

      Book Travel · task id green-fare-or-seat-fee-transparency

    • Compare the price of an add-on now versus later Booking flow

      Book Travel · task id ancillary-timing-advice

    • Get a recommended bundle and its price in the booking flow Booking flow

      Book Travel · task id bundle-recommender-price

    • Reach a validated checkout total before payment Booking flow

      Book Travel · task id smart-checkout-validation

    Showing 4 of 5 measured tasks that reach this level.

    Mapped playbooks 2

    The optimization ladder has three rungs. Tier 2 is the frontier — it earns its place only once Tier 0 and Tier 1 are clean.

    • Make consent dismissible without blocking agent interaction Tier 0 · Remove the blocker Consent wall

      Ensure the consent/cookie banner can be dismissed via a standard accessible button (not a pointer-only overlay) and does not capture keyboard or programmatic focus from the core task flow.

    • Expose WebMCP tools for true agent workflows Tier 2 · Offer the tool Structured path

      Implement window.__webmcp (or a link[rel='mcp'] manifest) exposing typed tool definitions for core tasks (search, booking) so agents invoke structured operations instead of driving the DOM blindly.

    Tracked standards 3

    Standards shape the ceiling of what a site can offer an agent. A declared standard never earns score credit by itself — only a measured signal does.

    • WebMCP Experimental

      Web pages expose typed tools a browser agent calls directly, skipping the DOM. Experimental — Chrome origin trial, Gemini-only. · Google & Microsoft · W3C CG

    • MCP (actions) Emerging

      Open protocol connecting agents to your tools and actions. Broadly adopted; ~8 months old. · Anthropic

    • Metadata telling agents how to authenticate to your protected APIs. · IETF · RFC 9728

  4. Transaction Can agents pay? (the highest bar) Beyond the money line

    What the level means

    The agent completes money-moving actions on the customer's behalf — with payment, agent auth and trust. Depends on Interaction. Below the money line.

    What decides it

    Does a verified agent complete a real paid booking end to end — payment cleared — with no human in the loop? Paid task completion.

    Measured signals 1

    Each signal is produced by one deterministic probe. The agent drives the browser; the probe decides what the run showed.

    • Task completion probe: answerExtracted Also under Interaction

      The verified binary outcome — did the agent surface an answer that passes the task's deterministic success check. The catch-all when no specific signal fired.

      This signal spans two maturity levels. It is the same measurement in both places — counted once, credited under Transaction and Interaction, not as two independent signals.

    Example tasks 0

    No task in the current suite crosses the money line. The benchmark stops before payment, so this level has no measured task today.

    Each run of a measured task returns one outcome: Completed, Failed or Blocked.

    Mapped playbooks 1

    The optimization ladder has three rungs. Tier 2 is the frontier — it earns its place only once Tier 0 and Tier 1 are clean.

    • Support agentic payment rails (x402 · MPP · UCP · ACP) Tier 2 · Offer the tool Task completion

      Adopt an emerging agent-payment protocol so an agent can complete a purchase under delegated authority. Frontier — earns its place only after Tier 0/1 are clean.

    Tracked standards 6

    Standards shape the ceiling of what a site can offer an agent. A declared standard never earns score credit by itself — only a measured signal does.

    • ACP Emerging

      Agentic Commerce Protocol — open checkout standard any ACP agent can pay. Beta; powers ChatGPT Instant Checkout. · OpenAI + Stripe

    • UCP Emerging

      Universal Commerce Protocol — merchants declare commerce capabilities agents discover and transact. Live in Google AI Mode / Gemini. · Google + Shopify

    • x402 Experimental

      Revives HTTP 402 for instant programmatic (stablecoin) payments. Young and volatile — optionality, not a baseline. · Coinbase · x402 Foundation

    • MPP Experimental

      Machine Payments Protocol — multi-rail agent payments over HTTP 402 (Stripe Agentic Commerce Suite). New. · Stripe + Tempo

    • AP2 Emerging

      Agent Payments Protocol — agents authorize payments via a cryptographically signed Mandate. Composes with UCP. · Google

    • NDC OrderCreate Established

      The NDC message that creates a booking / order programmatically. · IATA · NDC

How levels become published results Each task run returns Completed, Failed or Blocked. Blocked is an honest access wall, not a failure, and is excluded from the Index. A whole website is then graded on the scale below, derived from its measurable runs with an evidence floor of 2 measurable tasks.

  • Fails

    Measured tasks reach no verified completion.

  • Struggles

    Some tasks complete; material friction remains.

  • Completes

    The measured set completes, above the evidence floor.

  • Gated

    Every task hit an access wall; nothing measurable to grade.

Orthogonal axis

Access tiers are a gate, not a fifth level.

A maturity level asks what an agent can do. An access tier asks what gate it has to pass first. The two axes are independent: a site can be excellent at Understanding behind a login wall, and a public page can still break at Interaction. The tier is what makes an outcome Blocked instead of Failed.

  1. Public Measurable today

    To pass the gate: nothing — drives end to end

  2. Booking flow Measurable today

    To pass the gate: nothing — drives to the payment wall, then stops

  3. Booking ref Not yet measurable

    To pass the gate: a real PNR + last name on the carrier (cheapest gated proof)

  4. Login Not yet measurable

    To pass the gate: a test frequent-flyer account per carrier

  5. Eligibility Not yet measurable

    To pass the gate: a real disruption/IRROPS booking — best-effort only, can't be staged on demand

  6. Hybrid Not yet measurable

    To pass the gate: depends on the flow

  7. Staff Not yet measurable

    To pass the gate: n/a — not customer-facing

The benchmark can measure 2 of 7 tiers today. The rest are named so the boundary stays visible instead of being quietly rounded down to a failure.

Where this page comes from.

Nothing on this page is authored copy about the model — it is rendered from the versioned core objects below at build time. If the engine's contract changes, this page changes with it.

Data file
data/maturity-levels.json
data/dimensions.json
data/access-tiers.json
data/best-practices.json
data/standards.json
data/use-cases.json
data/tasks/airline.tasks.json

Per-file history is unavailable in this build: it was produced from a shallow checkout, which carries no per-file commit dates. Rather than print the checkout's own commit against every file as if each had just changed, the page prints none — the repository commit below is the one verifiable fact about this build.

Built from repository commit b3ecc29. The model is intended to stay stable while the specifics under it evolve; the commit is how you tell which specifics you are reading.

How it is measured Read the methodology: a real agent drives, the harness decides. Open the methodology → What it produces See the graded Last Mile Index built from persisted run records. Open the Index →