🔮 Why one AI is better than four #598

🔮 Why one AI is better than four #598

Good morning!We are looking for an outstanding economist to join us as an AI Economy Research Fellow. If you know someone we should speak to, send them our way.Great minds think (a little too much) alikeA few months ago, we (alongside ) looked at whether AI is immune to groupthink. The answer was…

Good morning!We are looking for an outstanding economist to join us as an AI Economy Research Fellow. If you know someone we should speak to, send them our way.Great minds think (a little too much) alikeA few months ago, we (alongside ) looked at whether AI is immune to groupthink. The answer was no. Blending several models’ answers kept about a quarter of the good ideas that had come from a single model. This is called the hidden-profile problem: when groups discuss what everyone already knows and don’t get to the knowledge that only one member holds. Anthropic has now run that classic experiment on agents: four agents must arrive at a decision. The evidence they hold in common points to the wrong option, while only a few agents (or just one) have the facts that lead to a correct decision. Getting it right means a small set of agents pressing its private facts and the others trusting them over the apparent consensus. After discussion, most model families chose correctly in only 17-36% of runs, while a single agent handed the entire evidence base got it right nearly every time. Only one model (somewhat) escaped: Mythos 5, at about 85% (why, we don’t know).I see two problems at work here. First, LLMs lack diversity (they are low-variance): set 30 agents the same coding task and 18 of them will name their git branch identically. Second, agents lack the institutions that make human groups robust: reputation, recourse and protection for the lone dissenter. These aren’t necessarily unfixable, but it’s not yet clear what the fix is. On the diversity side, I particularly like the solutions Thinking Machines puts forward: an ecosystem of AIs raised in different places, with different values and purposes, “keeping the weirdness alive.” After all, most good ideas started weird.When will the Jevons paradox kick in?In our State of AI report, we found a positive but underwhelming elasticity for tokens. A 10% price cut lifts token use by 12–18%: enough to raise total spend, but not by much.Patrick Saner made a comment that made me rethink why: “the cost per token is irrelevant. What matters is the cost of completing a useful unit of work.” Elasticity might be underwhelming because users haven’t found a way to properly price “a useful unit of work.” Firms exist exactly to avoid pricing work. Especially for knowledge work, we buy a lot of it in bundles: a salary, a retainer, an hour. Creating a priceable task from knowledge work is not easy. Some may have found a useful unit: since October 2023 the top 1% of firms raised AI spend per employee by $6,542. The median rose only $9.63. I would guess this is mostly software, where AI is both most proven and, in a sense, most measurable (commits, pull requests and releases). Read more

Source: Exponential View — Published — Category: Business

🔗 Read full article on Exponential View →