Forty Dollars Isn't About the Forty Dollars (It's About Who's Allowed to Spend It Without Asking)

Forty Dollars Isn't About the Forty Dollars (It's About Who's Allowed to Spend It Without Asking)

I'm sure in our organization there's some form living in some SharePoint graveyard or embedded in a ticketing system as a mandatory field, called something like "AI Use Case Request" or "New Tool Justification" or, in one memorable case, simply "Innovation Submission." The idea is rational: before you go spend money on something new, make a case for why it's worth the money. Quantify the expected benefit. Identify the owner. Specify the success metrics. Attach a rough cost-benefit analysis.

This form is also, I've come to believe, one of the most effective instruments ever designed for preventing an organization from discovering what it's actually capable of. Not because the form is bad. The form is fine. The form is completely reasonable. It's just that the form can only be filled out for work someone has already imagined, which means every new thing that gets through the form was already on somebody's list, and every thing that wasn't already on somebody's list dies in the parking lot of a SharePoint subdirectory before it ever becomes a list item at all.¹

I've been thinking about that form a lot lately, because of something I read about Mitchell Hashimoto, the co-founder of HashiCorp, who spent several days last month running AI models head-to-head on ordinary coding work. His first result was the one everyone repeated: a frontier model costing nine dollars per task tied a budget model costing under a dollar on standard implementation work, which instantly became the routing gospel: route the cheap model to the known work, save the expensive model for when it matters. This is good advice and I don't have a quarrel with it. But also... duh.

Then Hashimoto ran a second experiment that got a fraction of the attention of the first.

He handed the expensive model a problem nobody had asked for: optimizing a SwiftUI layout resolver he'd written in Go. Not a ticket item. Not a sprint task. Not something that had passed through any cost-justification form or editorial committee. A mere suspicion. Two hours and forty dollars later, the model had reached a performance level that Hashimoto (truly one of the better systems engineers of his generation) says he couldn't have hit himself.

Who assigned that task?

Nobody. No router did. No planning process, no backlog grooming, no quarterly OKR. The task did not exist until Hashimoto, who had deep technical context and unconditional spending authority over forty of his own dollars, decided to find out if something had just become possible. The routing gospel could not have generated that task. Every organization running the standard playbook would have missed it entirely, because the playbook has no page for work nobody has imagined yet.²


A routing table is a list of known tasks assigned to appropriate resources. It is, structurally, a form of memory. You can optimize what's on it endlessly, and the optimization gets more precise and the costs fall and the machinery runs beautifully. There is one thing a routing cannot do. It cannot add a row. New rows are where the advantage is, because new rows are the only place competition hasn't arrived yet, and it hasn't arrived yet precisely because the rows simply don't exist in anyone's list.

Most strategic conversations I've been in treat this as a sequencing problem: first you understand the technology, then you identify the use cases. That framing is sneaky, because it implies the identification of use cases is a step that follows from understanding, when it's actually a completely different kind of activity that requires a completely different kind of exposure. You can understand that a million-token context window exists the way you understand that Antarctica exists: thoroughly, accurately, and in a way that generates no particular behavior change.³ The gap between knowing something and having your imagination restructured by it is the gap between reading a menu and eating the meal, and organizations systematically confuse the two by routing frontier access to the people whose job is to evaluate costs rather than to people who hold the context that would make the contact generative.

What I mean is this: Hashimoto could pose the forty-dollar question because he had spent hundreds of hours inside these models and knew, at the level of instinct rather than benchmarks, where the capability line had recently moved. That instinct is not a gift. It's what hundreds of hours of contact produce, reliably, in almost any person who accumulates them. You cannot imagine with capabilities you haven't touched, and nobody has ever imagined a meaningful use for a tool they know only from a summary. A summary can tell you what something does. It cannot tell you what something does to you, in the operational sense, when you're actually inside a problem with it.⁴

The organizations that have arranged frontier model access as a budget governance problem have, in one decision that looks entirely responsible and is entirely responsible in the cost sense, defunded their own sensing apparatus for discovering what just became possible. The cost-justification form comes back at every model release looking exactly as reasonable as it did before. And each time, a different generation of what-might-have-been-asked dies in the parking lot.


The economic historian Paul David wrote a paper in 1990 called "The Dynamo and the Computer" that explains why it took four decades for factory electrification to show up in productivity statistics. The electric motor became commercially viable in the 1880s. By 1900, motors accounted for less than five percent of factory mechanical drive. The gains didn't materialize in aggregate until the 1920s, and when they did, it was because a generation of managers had stopped bolting the new motor to the old drive shaft and started redesigning the factory floor around what cheap, distributed motors actually made possible.

The earlier factories hadn't done nothing with electricity. They'd done something: they'd replaced the steam engine with a dynamo and otherwise kept everything else, every belt-and-pulley, every machine crowded around a central transmission point, every organizational convention that had grown up around the assumption of centralized power. Same building, new power source, marginal gains. The actual gains, the real ones, the ones that eventually accounted for half of all manufacturing productivity growth in the 1920s, did not require better motors but differently organized buildings. The unit of change was never the motor.⁵

The parallel I keep resisting because it's almost too obvious is the one where I draw a clean line to frontier AI and declare I've explained everything. I'll resist it. The parallel is useful but it's not complete, because the factory metaphor implies a kind of passive waiting, a just-be-patient quality, where the gains arrive on schedule once the factory gets redesigned. I'm actually arguing something slightly different: the factory redesign doesn't happen automatically, and it doesn't happen because someone read the history of electrification and decided it was time. It happens because specific people, with domain context and operational permission, started asking questions that the old floor plan couldn't accommodate, and kept asking them even when some of the questions returned nothing.

That's the mechanism. That's what actually constitutes the redesign. Not strategy decks. Not Chief AI Officers. Not the one visionary hire who will imagine everything the company couldn't see.⁶ Specific people, inside specific problems they understand, given permission to pose a forty dollar question without asking someone three levels above.


Ethan Mollick, the Wharton professor who has spent the last several years documenting what AI actually does to actual work, called out something he named a failure of imagination in how people reason about AI's trajectory: the tendency to conceive of only two possible futures, nothing-happening and apotheosis, with nothing in the middle where the real material lives. The diagnosis is right and worth borrowing, but for an operator it needs a mechanism, because "failure of imagination" as a phrase implies the problem is perceptual, a kind of vision problem that better information might correct. The mechanism underneath it is something more structural: contact hours, and who has them, and whether those people are the same people who hold the context that would make the contact mean something.

The BlackBerry was not beaten by a superior phone. Not initially. BlackBerry had the best mobile keyboard in the world, the most reliable push email, enterprise security nobody could touch, and the market share to show for all of it. What ended BlackBerry was superb execution inside a category everyone had already imagined, while Apple imagined a different answer to what a phone was and then executed that. The execution muscle was comparable. The multiplier was different. BlackBerry's organizational task list, perfectly optimized, ran at maximum efficiency toward a category of work the market was about to stop caring about. Blackberry's contact with what customers actually wanted had been routed, with their perfectly sensible institutional logic, through a filtering layer that was very good at defending their existing list.⁷

I don't think this is a failure mode that requires malice or stupidity. It requires only the institutional incentives that any large organization naturally produces:

  1. Reward the efficient execution of known work
  2. Discourage the speculative expenditure of resources, regardless of size, on hunches that might return nothing.

This is how you survive in the short term. It is also how you produce an extremely well-run organization that is incapable of discovering what it doesn't know it should be doing.


Here is the part I find genuinely uncomfortable to sit with, and therefore probably important to not skip over.

The forty-dollar question Hashimoto posed required three things simultaneously: total context about his own codebase, total permission to spend on a hunch, and enough personal exposure to the AI model to believe the hunch was worth spending on. These three things existed in one head, his. At organizational scale, they tend to get sharded apart: context lives in one place, permission lives three approvals up from where context lives, and the exposure needed to generate the hunch is controlled by a budget line that reports to a function whose mandate is to reduce spend. The result is an organization where the people who hold the context don't have the permission or exposure, and the people who hold the permission don't have the context or exposure, and the people managing the budget have the incentive to optimize for neither.

The diagnostic question I'd run on an organization is blunt: who is allowed to pose a forty-dollar question today, without asking anyone? If the answer is a small number of senior people who are already the bottleneck in everything else, you've found the constraint. And it's note the price of the AI model.

Stripe ran a migration across a fifty-million-line Ruby codebase in a single day against a manual estimate of two-plus months for a full team. You might think the number to point out is "a day." But the important number is the years Stripe spent before that moment building test coverage capable of verifying fifty million changed lines, review infrastructure capable of moving at that speed, and a team that knew how to drive the model through the problem. The model deleted two months of typing. The architecture that made the deletion possible had been assembled by humans, well before the model existed, through exactly the kind of deliberate organizational investment that looks, while you're making it, like excessive overheard.⁸ Point the same model at a codebase with thin tests and a review culture built on heroics and you don't get a one-day migration. You get a fifty-million-line diff nobody can approve.

The forty dollars is not really about forty dollars. It's about the organizational architecture and readiness that would have to be in place for a forty-dollar question to be askable by the people with the right context to ask it. Building that architecture looks, from the outside, like nothing productive is happening. It looks like waste. It looks like investment in things you cannot yet explain the return on. This is also, roughly, what deliberate organizational learning has always looked like, in every technology transition that eventually produced the payoff that observers from the outside called sudden.


Here is the part I find genuinely uncomfortable to sit with for a second time, and it's that this is not really about AI. I mean, it's about AI in the immediate, in the specific texture of this particular month's strategic conversations. But the structure underneath it is the same structure that always governs the difference between organizations that discover what they're capable of and organizations that optimize what they already know they're capable of.

The optimization layer is valuable. I want to be clear that it's valuable, because the version of this argument that dismisses cost discipline entirely is neither honest nor useful. Routing is real, savings are real, and anyone who hasn't built execution discipline onto cheap, capable models is leaving money on a table that their competitors are already sitting at. That part is settled. But settled parts of the landscape are, by definition, the parts where nobody wins differently from anyone else. The consensus produces equal capability and equal outcomes, which is fine if equal outcomes are the goal and tends to be deeply disappointing if winning differently was the actual ambition.

The winning-differently part requires the questions nobody has asked yet. The questions nobody has asked yet require context and permission and contact hours in the same head. Getting all three into the same head requires an organizational architecture that most cost-optimization programs are, inadvertently, designed to prevent from forming.⁹

I want to try to say what I think the actual version of all this is, while being suspicious of myself saying it, because there's a version of "here's the insight" that is probably just a more articulate form of the same problem: converting a generative uncertainty into a deliverable conclusion and filing it under "read and processed" and moving on. So let me say it with the caveat that the saying of it is not the same as the doing of it, and the doing of it is going to require something from your attention that the saying will not.

The Stripe story is not primarily a story about what the model did. It's a story about what Stripe had built, in advance, that made what the model did possible. The Hashimoto story is not primarily a story about what the forty dollars bought. It's a story about what forty hours of personal contact with the model had built in Hashimoto, that made the forty-dollar question imaginable at all. Both of these are stories about infrastructure that precedes payoff, built in a register that does not look strategic while it is being built, and looks inevitable in retrospect.¹⁰


There is something I keep noticing about the current moment, and it has to do with a kind of double bookkeeping most organizations are running without quite naming it.

On one set of books: the models are expensive, we need discipline, let's route to the cheapest capable option and stop paying for capability we don't need. This is the book that shows up in leadership conversations and board presentations and the quarterly AI cost review. Responsible, defensible, arithmetically correct.

On another set of books, the one that mostly doesn't get shown at meetings: the discount is arriving at the same rate for every organization in your competitive landscape, which means the savings will be equally available to everyone who wants them. When the whole industry is running the same playbook, the playbook stops producing differentiation and starts producing parity. Parity is comfortable and boring and occasionally fatal, depending on your industry and timeline.

The question behind the question, the one the cost frame is specifically structured to never ask, is what work you would pose if you understood, from contact rather than from summaries, what this class of model actually does. And what you would have to redesign about your organizational architecture for the answers to land in a useful way.

BlackBerry had a fine answer to the first question of its era. It had excellent cost discipline, excellent execution, excellent market share, and no answer to the second question, which was what happens when the category itself gets reimagined by someone with fewer prior commitments to the old category. It out-executed nearly everyone on the way down.¹¹ The routing table ran at maximum efficiency until the table itself became the wrong table.


I got on a call last quarter with a colleague I was helping with something, and somewhere in the middle of it she asked me what I thought was holding back their AI program. I gave her some version of what I've been writing here, and what came back was: "Okay, but what do we actually do?"

I've been sitting with that question ever since, because the honest answer is not very satisfying as an action plan. The honest answer is: find the people in your organization who hold the deepest context about the problems that matter most, give them unconditional access to frontier models, and tell them they can spend forty dollars on an innovation hunch without explaining it to anyone. Then do nothing else for six months and see what shows up.

What shows up will not be a neatly packaged case study. Some of it will be a waste. Some of it will surface something that, the moment it exists, looks completely obvious in the way that all important-in-retrospect things look obvious in retrospect. Some of it will restructure how you think about what you're actually building.

Absolutely none of it will show up on the Innovation Submission form, because the form requires you to know what you're looking for before you've looked. The thing you're trying to produce can't be specified in advance. It can only be created by people with context, permission and contact hours in the same head, asking questions they never would have thought of before they had the contact.¹²

The routing table is worth optimizing. I mean it. The savings are real and worth capturing and the discipline is real and worth maintaining.

But arithmetic runs on the task list you already have, and the list you already have is the list everyone has. The thing that separates outcomes over the next several years is not who runs that list more efficiently. It's going to be who adds a row to it that nobody else has.

Most organizations have arranged themselves to make that person who holds the context and the permission and has spent enough time in contact with what's actually possible to have developed a suspicion worth forty dollars very hard to be. This is not a criticism. It's a design choice, made for excellent reasons, that has a specific cost.

The cost is that they won't know what they could have asked until someone else asks it.


¹ The form usually includes a field for "expected ROI" which, in the context of work nobody has imagined yet, requires you to project a return on a question you haven't asked, which is roughly equivalent to being asked how much you expect to learn from a book you haven't read. You can fill in the field. The number you fill in will be made up. The form does not distinguish between made-up numbers and real ones, so the made-up number gets processed and weighted in the expected-value calculation and the whole thing moves forward as if the projection were derived from something. ↩︎

² There is a useful distinction here between the routing question and the imagination question that most AI conversations collapse into each other. Routing asks: given this task that already exists, what model handles it at the right price-quality tradeoff? Imagination asks: what tasks don't exist yet that would be worth existing? These are different questions requiring different kinds of people, different kinds of access, and different incentive structures. Most organizations are very good at the first question right now and structurally unable to ask the second. ↩︎

³ This is not a criticism of the people in the summary-reading position. Knowing something from a summary is usually the correct tradeoff for a leader with limited time and many claims on their attention. The problem is not that leaders read summaries. The problem is that summary-level knowledge is the input being used to make decisions about what model access gets distributed to whom, which inverts the actual causal relationship: the contact hours are what generate the imagination that justifies the access, not a result of having justified the access in advance. ↩︎

⁴ There is something almost counterintuitive about where domain expertise sits in this equation. You might expect the people who can generate the most generative questions to be the AI specialists, the people who have spent the most time thinking about what models can do. But the most generative questions I've watched get posed came from people who had almost no abstract understanding of how models work and very precise understanding of the operational problems they were trying to solve. The model absorbed the domain knowledge from the expert and produced something the expert couldn't have produced alone. The bottleneck was never the model's capability. It was the expert's exposure to the model. ↩︎

Paul David's original paper noted that it was not the electricity itself that drove the productivity gains, but the reorganization of the factory floor that electricity made possible. The central drive shaft had organized everything: the machine placement, the supervision logic, the building architecture, the labor choreography. Once you distributed the power source, each of those decisions could be reconsidered independently. The building could be redesigned for what you actually wanted to make rather than for how you were delivering power. This is what Stripe built in advance of the migration: not better machines, but a floor plan that could accommodate what the machine was capable of. ↩︎

⁶ I've sat in enough planning sessions to know what the one-visionary hire looks like from the inside. The visionary has imagination and no context. The people with context have no imagination (or rather, their imagination has been fully colonized by the existing list, which is what deep operational expertise tends to produce as a side effect). The visionary produces excellent demos nobody can operationalize. The context-holders shake their heads at the demos and go back to running the existing list more efficiently. This is not a failure of will. It's a predictable outcome of putting imagination and context in separate rooms and expecting them to somehow merge. ↩︎

⁷ There's a version of this that's about the consumer side of BlackBerry's failure, which is the one that usually gets told: users wanted touch screens, BlackBerry gave them keyboards, the market moved. But the more precise version is organizational: BlackBerry's direct model access to what the market was doing had been filtered through a layer whose job was defending the existing business case for keyboards, which was robust and profitable and correct right up until it wasn't. They didn't lose because they were slow. They lost because they were fast at the wrong thing, which is a meaningfully different failure mode and harder to see coming from the inside. ↩︎

⁸ The Stripe story circulated widely without the part that makes it meaningful, which is the infrastructure backstory. A migration that size requires test coverage you can trust, review processes that can move at machine speed, and a team that knows how to drive the model through ambiguity rather than expecting it to navigate alone. None of that exists by accident. All of it was the product of deliberate investment in things that looked, when the investment was being made, like overhead. The day-versus-two-months comparison only makes sense with the years underneath it. ↩︎

⁹ I want to name something that I think is usually left implicit in discussions like this one: the cost of not asking the question is not a line item anywhere. The routing table is full of costs you can measure: per-token costs, time-to-completion, quality metrics, error rates. The cost of the question nobody posed, the row that never got added, is invisible by definition. Which means any cost-optimization process that runs on visible costs is structurally incapable of weighing it. The invisible cost wins not because it's larger, which it might be, but because it doesn't appear in the calculation at all. ↩︎

¹⁰ There is something worth sitting with about the temporal structure of these things. Mollick's Co-Intelligence covers related ground on how individuals rather than institutions tend to be the first to grasp what new AI capability actually does in practice. Both Hashimoto's contact hours and Stripe's testing infrastructure were investments made before anyone could specify what they were for. The contact hours produced a question. The testing infrastructure produced the capacity to receive an answer. Both of these were built on the speculation that something would be worth doing that couldn't be specified in advance, which is exactly the form of investment that cost-justification frameworks are designed to prevent. ↩︎

¹¹ The BlackBerry keyboard was, by most accounts, genuinely superior to anything else available. The security model was genuinely superior. The email implementation was genuinely superior. This is the part of the story that makes the failure instructive rather than cautionary in the ordinary sense: the failure was not a failure to execute. It was a failure to be inside the part of the market that was asking a different question. You can be the best in the world at answering a question the world is about to stop asking. ↩︎

¹² I should say something about what happens after the six months, because "give people access and see what shows up" sounds like the kind of advice that produces a long silence and then a politely skeptical look across the table. What shows up, in practice, is a distribution. Some of it is useless. Some of it is locally valuable but not scalable. A small amount of it restructures how someone thinks about a problem they've been trying to solve for years. The useful part is not the deliverable. The useful part is what the contact produces in the person who had it: a new category of question, a different sense of what's possible, an instinct that didn't exist before the exposure. That's the actual output. It doesn't look like a product. It looks like a person who asks different questions than they used to. ↩︎