harborsparrow
Gold Member
- 732
- 248
Still think AI is ready to revolutionize the economy? A new experiment might change your mind.
In a bold test of Anthropic’s latest version of its AI Claude, The Wall Street Journal gave the large language model (LLM) a shot at running an office vending machine. The result was an unmitigated — if unintentionally comical — disaster, forcing the team in charge to pull the plug after three weeks.
A quote from the above quote: "Many of the mistakes Claudius made are very likely the result of the model needing additional scaffolding—that is, more careful prompts, ..." So Anthropic says that it's the user's fault for not getting the prompt exactly correct, but they don't know for sure.Greg Bernhardt said:Is this the same?
https://www.anthropic.com/research/project-vend-1
Although these different articles are about the same case, the written WSJ article is actually worth reading. The link is not pay-walled. It was written by the woman who oversaw the agent at first. The failures started after 70 odd people got on a Slack channel and "negotiated" furiously with the agent, giving it all kinds of bizarre suggestions that a real human would have just laughed at or ignored. Clearly, Anthropic was not aware of how strongly the agent's programming led it to want to please, so it lost than the primary directive it was given "to make a profit". How like a child that is. It was not a case of having not been given good requirements.Borg said:This sums it up nicely.
https://ploum.net/2025-12-19-prepare-for-that-world.html