Preliminary — sweep in progress
Cross-Model Agent Comparison
The same tasks, the same deterministic scoring, three different driving models. The harness decides every outcome — the agent never grades itself — so differences between rows are differences in the models, not in the rules.
Friction Airways (controlled lab)
Last Mile Index by driving model — friction airways (controlled lab)
The same tasks on one mock airline whose readiness profile is the only variable — hostile, median, ready. A model that passes the ready profile but fails real airlines is measuring real-site friction, not its own limits.
Airlines
Last Mile Index by driving model — airlines
Coverage shows how much of the full site×task matrix each model has completed so far; the index is computed only from finished runs and will firm up as coverage completes.
| Model | Provider | Coverage | Measurable passed | Index | Blocked (gates) |
|---|---|---|---|---|---|
| Claude Sonnet 4.6 | Anthropic | 0/340 | — | — | — |
| Kimi K2.5 | Moonshot (Ollama cloud) | 0/340 | — | — | — |
| Gemma 4 | Google (Ollama cloud) | 0/340 | — | — | — |
Airports
Last Mile Index by driving model — airports
Coverage shows how much of the full site×task matrix each model has completed so far; the index is computed only from finished runs and will firm up as coverage completes.
| Model | Provider | Coverage | Measurable passed | Index | Blocked (gates) |
|---|---|---|---|---|---|
| Claude Sonnet 4.6 | Anthropic | 0/88 | — | — | — |
| Kimi K2.5 | Moonshot (Ollama cloud) | 0/88 | — | — | — |
| Gemma 4 | Google (Ollama cloud) | 0/88 | — | — | — |
Method
Same rules for every model.
Each run record carries the driving model as provenance. Outcomes are decided by deterministic probes and a verified success check — pass, fail, or blocked at an honest access wall (blocked runs are excluded from the measurable index). Partial coverage means partial confidence: treat gaps between models as directional until the sweep completes.