Scans pass. Agents fail.
Last Mile runs real browser agents through real customer journeys, then verifies the outcome with a deterministic harness. The Index ranks completed tasks, not readiness theater.
- Measured airlines
- 7
- Task completion
- 9%
- Honest gates
- 3
Route 1 of 8: King Khalid International Airport to London Heathrow Airport
Three gates. One completed journey.
A site can be discovered and understood and still fail before the task is complete. Last Mile tests every gate with a real browser agent.
-
Discoverability
Can agents find you?
Surfaced and reached
-
Understanding
Can agents read you?
Structured path · legible inputs and errors
-
Interaction
Can agents operate you?
Unpaid task completion · failure recovery
Every gate receives one verified outcome
- Fails
- Struggles
- Completes
Ranked by completed tasks, with the static scan kept in view.
The public leaderboard preserves the inversion that matters: sites can present clean static signals and still fail the browser-agent run.
A real agent drives. The harness decides.
Lowest fare, check-in, policy extraction, or another measurable task with a known success check.
Run Browser agent attemptThe model drives the browser while Tarmac records friction signals and evidence.
Verify Deterministic outcomeThe agent never grades itself. The harness returns pass, fail, or blocked access.
Blocked access walls are separated from failures so the Index stays defensible. Read the scoring spine.
Friction Airways proves the levers move the score.
Real airlines show where agents break. Friction Airways shows why: same airline, same tasks, same scorer, three readiness profiles. Flip the levers and the Index swings from zero to perfect.
Open the PlaygroundEvery failing signal has a fix.
The recommendation ladder turns each friction signal into a tiered fix — the playbooks that move a site up the maturity model.
The report is generated from the same run records.
The public site, PDF report, and methodology pages all point back to the same task records and scorer. No parallel narrative, no self-assessment.
Get your site to the last mile.
Benchmarking, per-site diagnosis, and the fix roadmap — the engagement that takes a site from failing the agent run to completing it.
See what we offer