What do consultants get paid for?
The analysis is the easy part
A consultant I had lunch with recently is redesigning the loyalty program of a large airline. His team finished the analysis in two weeks. Months later the program still does not exist. This is because the purpose of the assignment is not to solve an analytical case study, but to figure out which redesign the parties will accept, and to get the people with authority to commit to implementing it.
I was not surprised to hear the story. In our just-published book Messy Jobs: The Work That AI Cannot Reach, Jin Li, Yanhui Wu, and I argue that a job is not a collection of independent tasks but a bundle of tasks and a position inside an organization. While many of the constituent tasks are clean, the job is messy because they must be combined under incomplete knowledge, conflicting objectives among the different parties and binding constraints on who has the authority to make decisions.
Hence we argue that automating the clean parts does not necessarily eliminate the job, because the remaining activities, tightly bundled with the rest, can remain the constraint. We argue that the bundle is strongest where separating the analytical/cognitive parts that can be automated would destroy local knowledge, trust, accountability or continuity.
Two objections
Critics of our argument raise two concerns. The first one has to do with advances in AI capabilities: models do some tasks extremely well and others badly. As they gain memory, use tools and acquire multimodal perception, critics would say, AI will get better at many other tasks like persuading, anticipating the objections raised in a meeting and adapting the tone. Hence even the interpersonal part by itself may not be a sanctuary for long. Just wait a bit for AI to get better, say the critics: Messy Jobs (in their view) describes the transition rather than the long run.
The second objection to our thesis is more radical. Maybe as long as we have humans in the loop, we need organizations. But if organizations really are a mess, slow, political, resistant to change, with a large role for humans precisely because someone has to hold meetings, build coalitions and learn the internal politics, why not get rid of the entire organization? What is the point of preserving the existing obsolete structures?
We believe that both objections fail, because some of the mess is substantive and necessary.
Where the “implementation” months go
To an outsider, my friend’s consulting project looks purely analytical.1 The team receives all the data, including all passenger records, redemption rates, customer-retention data, and so on. The team works out the key economic and financial trade-offs of the possible redesigns to figure out which redesign maximizes profits.
If doing this, given all the available data and the current AI tools, took two weeks, why has the project taken many months?
First, inside the airline, different parts of the business worry about different things. For instance, the salespeople have relationships with the hotel chains and do not want to disturb them, while the operations team worries about the staff at the airport counters who will have to deal with angry passengers who have grown used to certain privileges.
Second, there are the outside parties, from hotel chains to credit-card companies to the retailers that accept miles. Anything that improves the airline’s economics may reduce the value of the program to the hotels or to the card issuers. Each has a view on how card spending should count relative to flying, or how hotel nights should count relative to flying. There are winners and losers everywhere, and a reform that benefits the airline as a whole can hurt a particular business unit or a particular partner.
So the consultants spend weeks doing an enormous amount of work that looks peripheral to the problem. They repeatedly meet the head of the loyalty program, then the CFO, then the CEO. They also meet the commercial partners, and that means meeting the head of loyalty, then the finance team, then the chief executive. They revise the proposal. They redo the presentation.
And all of these people speak different languages. Organizations have different internal codes because they care about different things and deal with different problems. What the consultants are doing is a mix of analysis, translation and intermediation. The assignment of the consultants is to design a program that is an agreement that the relevant parties will authorize and implement.
Once the analysis is cheap, what remains is to learn what each party will actually accept, and to obtain commitments from those who are authorized to make them.
The real knowledge problem
An advocate of highly capable AI systems (“AGI-pilled”) would probably say this is a problem ready for AI. Have an agent redesign the program, have it meet the other constituencies, have it come back with a solution.
But what happens in those meetings deserves a closer look. There are four frictions in the room that make the meetings necessary.
First, knowledge is dispersed. This is the Hayek problem: different people hold different bits of information. AI can help with this. By making the emails, contracts, past redemption data and meeting transcripts searchable and available to the system, the model can extract part of this dispersed knowledge. But a lot of the dispersed knowledge remains in people’s heads as it is local and contingent, and will only emerge in the meeting.
Second, knowledge is also tacit: the Polanyi problem. We know more than we can tell. A senior partner in the consulting firm who knows the client well knows instinctively that a particular proposal will not fly. Again, AI may help reduce this problem, as it can learn from the actions people take on the basis of their knowledge. Brynjolfsson, Li and Raymond (2025) found that AI assistance diffused some of the communication and problem-solving practices of stronger customer-support agents to less experienced workers. The system had captured enough observable patterns in stronger agents’ behavior to reproduce some of their tacit knowledge advantage.
The third friction is that knowledge is not available to the AI system because people refuse to disclose it: strategic private information. The hotel knows how much it would cost to eliminate one feature of the loyalty program, and it will exaggerate that cost to extract more value in the exchange. Under specific assumptions about bilateral trade, Myerson and Satterthwaite showed that no mechanism can guarantee full efficiency while also inducing truthful revelation, respecting voluntary participation and balancing the budget. The theorem applies equally to humans and machines. Better models reduce the cost of drafting and bargaining, but don’t solve the problem of deciding who gets what.
The fourth friction is that the objective has not yet been formed or authorized. Think about the hotel chain. If you ask them initially what they want out of their program, they may not have an answer, since the organization does not know which features are critical or the cost of conceding a feature. That exploration only happens through iterative meetings, and what is happening in those meetings is that people are collectively discovering what the organization wants, what the different features are worth and which relationships they care about. The chief executive eventually makes the call, but the parts of the organization arrive at that point with different views and without a settled objective. The participants use those negotiations to discover the trade-offs, form their own view about their preferences among these trade-offs and authorize someone to bind the organization.
The economics of learning
Knowledge is expensive to acquire and cheap to use once acquired. In my work on knowledge hierarchies, that asymmetry makes hierarchy (“management by exception”) necessary: workers learn to solve the routine problems, which arrive constantly, while specialized problem solvers learn to deal with the exceptions and recover the cost of that knowledge because many workers refer their exceptions upward to them. A hierarchy, then, is a device for raising the utilization of expensive knowledge.
Machine learning relies on the same logic, taken to the extreme: a large language model is extremely expensive to train. In a recent post, Dwarkesh Patel described current models as one-millionth as sample-efficient as humans during training. He illustrates this with driving, where a teenager becomes a minimally competent driver after tens of hours of practice, three to four orders of magnitude less data than training autonomous-driving systems has required.
Epoch estimated that the datasets used to train language models grew at a historical rate of about 3.7 times per year. Why is this enormous data and training effort worth it? Because, as in the hierarchy problem, this knowledge, acquired at huge cost, is leveraged over a huge number of instances. You can spread the cost of learning over hundreds of millions of instances. The AI systems are hard to train, but cheap to run once trained.
Will algorithmic progress reduce the data requirement? This is controversial, but right now it does not look like this is happening. Gundlach et al. (2025) study training compute efficiency. They re-ran the main algorithmic innovations of 2012 to 2023 and could account for less than 100-fold out of the 22,000 fold efficiency gains previously attributed to algorithms, while their scaling experiments suggest most of the measured gains came from a handful of scale-dependent innovations. Ho argues that much of the rest is plausibly better data.
The problem with the renegotiation of the loyalty program is that the decisive local knowledge is used once, and it is only released strategically. The model can learn, from many other loyalty programs, the analytics and the usual sources of conflict. That part has high utilization: many airlines, many engagements. What is less portable is the local context: which objections are genuine and which ones are strategic, who has informal authority, which relationship is strong and can withstand pressure, what each party will concede. The local context has a utilization of one. And even perfect learning efficiency would not eliminate the incentive problem involved in eliciting truthful reporting.
True, some of the evidence accumulates in the “context window” from previous interactions. But even that is incomplete. Humans engaged in bargaining know that what they say will be used against them (whether by humans or models) and say different things on the record. That is why the conversations in the corridor are so important, as anyone who attends such meetings knows: a CEO or a minister may tell you in a brief aside something she would never share in public.
The argument is not that models cannot handle novelty. We have all seen them over the last few months prove theorems no human had managed to prove. But proofs can be verified, and hence the recombination of patterns learned in training can work. Nor is the argument that machines cannot bargain. In fact, we know that given fixed rules and a scoreboard, machines have reached human-level play in Diplomacy, a board game. But the loyalty renegotiation has neither fixed rules, nor an agreed objective scoreboard. What is missing is a cheap, portable data set that has not been strategically distorted and that would allow a model to predict whether a particular agreement will work.
This means jobs will be protected from automation by messiness when decisions involve more independent decision-makers and where the feedback is weaker. It also predicts that AI will reduce how long it takes to produce proposals by more than the time required to authorize and implement them. The result could be a “Jevons paradox”: less analytical effort per decision, but more contested proposals to be considered and authorized, and so more work for consultants and managers.
My book with Jin Li and Yanhui Wu, Messy Jobs: The Work That AI Cannot Reach, is now out on Amazon. If you have already read it, a one-minute rating on Amazon is the most useful thing a reader can do for it.
References
Brynjolfsson, Erik, Danielle Li, and Lindsey R. Raymond. 2025. “Generative AI at Work.” Quarterly Journal of Economics 140 (2): 889–942.
Crémer, Jacques, Luis Garicano, and Andrea Prat. 2007. “Language and the Theory of the Firm.” Quarterly Journal of Economics 122 (1): 373–407.
Garicano, Luis. 2000. “Hierarchies and the Organization of Knowledge in Production.” Journal of Political Economy 108 (5): 874–904.
Gundlach, Hans, Alex Fogelson, Jayson Lynch, Ana Trišović, Jonathan Rosenfeld, Anmol Sandhu, and Neil Thompson. 2025. “On the Origin of Algorithmic Progress in AI.” arXiv:2511.21622. https://arxiv.org/abs/2511.21622.
Hayek, F. A. 1945. “The Use of Knowledge in Society.” American Economic Review 35 (4): 519–30.
Myerson, Roger B., and Mark A. Satterthwaite. 1983. “Efficient Mechanisms for Bilateral Trading.” Journal of Economic Theory 29 (2): 265–81.
Polanyi, Michael. 1966. The Tacit Dimension. Garden City, NY: Doubleday.
The commercial details are confidential and I have changed some of them.



Very insightful. My one quick comment (also as a former partner at McKinsey and Bain :)) is that consultants are also paid for making trust incentive-compatible. Repeat business, references, and the firm’s reputation put future rents at risk—an implicit bond that AI does not yet replicate.
The issue of trust is studied in a recent paper by Piotr Dworczak and Alex Smolin, “Robust Trust” (2026): https://arxiv.org/abs/2602.09490
AI is still very bad at persuading. It can't communicate in the modes you need to communicate effectively. Like it can't keep an email chain going over a few days with multiple parties, it can't participate effectively in a Zoom call with many people, it can't have a "quick chat" to answer questions, it can only really answer slowly at length, it can't stay focused on any issue over a timeframe of weeks or months.
Maybe this will change eventually, but for now, the humans have to do all this stuff.