Signal Economics

OpenAI's CFO Says Cost-Per-Token Is the Wrong Way to Measure AI ROI — Hospitality Already Knew That

OpenAI CFO Sarah Friar's 'Useful Intelligence per Dollar' framework for enterprise AI ROI maps almost exactly onto the KPI-baseline-durability test Hotel Equities CEO Ben Rafter published three weeks earlier for hospitality operators evaluating AI spend.

OpenAI CFO Sarah Friar published a framework arguing that cost-per-token is the wrong yardstick for enterprise AI ROI, proposing “Useful Intelligence per Dollar” instead — measured by how much useful work AI completes, the true cost per successful task (full employee time, review, retries, and rework, not just token price), dependability, and whether returns compound as usage scales. It’s a vendor’s framework, published by the vendor selling the tokens, and that’s worth noting — but the underlying argument holds regardless of who’s making it: raw AI spend tells you nothing about whether the spend is working.

That’s precisely the corrective Ben Rafter, CEO of Hotel Equities, has been making about hospitality specifically. Rafter’s argument is that AI delivers measurable, provable returns fastest on the operational side — scheduling, procurement, revenue management, maintenance triage — while guest-facing chatbots and in-room assistants remain comparatively immature and harder to prove out financially. He lays out the questions he believes every owner should be asking any management company pitching AI: what specific KPIs moved, over what time period, compared to which baseline, and whether the gains are durable or a one-time novelty effect. That’s “Useful Intelligence per Dollar” in hotel-operator language, published a full three weeks before OpenAI’s version.

The overlap matters because it means hospitality operators don’t need to translate a generic enterprise-AI framework into property-level terms — the translation already exists, tested against real portfolios rather than theorized from the vendor side. Friar’s four questions (how much useful work, true cost per successful task, dependability, does it compound) map cleanly onto Rafter’s KPI-baseline-durability test, and both arrive at the same conclusion: the yardstick isn’t what a tool costs to run, it’s what it demonstrably changed.

For any operator being pitched an AI tool on token efficiency or model benchmarks, the actionable question is still Rafter’s, not Friar’s: which KPI moved, against what baseline, and is the gain still there in six months.

Auto-generated brief — verified before publishing.

← All signals

Meet the Founder

Want this kind of thinking applied to your portfolio?

A genuine conversation — no pitch, no deck. Twenty minutes with the person who'd do the work.

Book a 20-Minute Call