Science Advisor
- 3,643
- 10,066
First, here's the "hot take" footnote from the reference (summed up with "He needs to see more before he gives his final appraisal."):gleem said:Gary Marcus, prominent cognitive psychologist and AI hype critic, on a "hot" take, says,
https://garymarcus.substack.com/p/hot-take-on-gpt-6-astra
He needs to see more before he gives his final appraisal.
- Looks to be pretty impressive. Multiple reports suggest it is a genuine advance.
- As someone who has campaigned for nearly a decade for (neuro)symbolic world models, often to exceptional hostility, it is extraordinarily vindicating to see that a product from OpenAI explicitly creates and manipulate symbolic world models in the course of some of its most impressive computations. . . .
Second, here are the 7 extra "hot takes" left out in the quote (summed up with ". . ."):This hot take is VERY tentative, pending more information about how the systems works and more detailed examination of what its limitations are.
I like the fifth one.
- What we don’t know is how robust that capability is. That is THE key question.
- Success on ARC-AGI is great and impressive, but not —despite the name of the task—proof of AGI; I suspect we will see loads of problems with open-ended real world tasks. As with other recent models I would suspect best performance in verifiable domains.
- And as a scientist, it’s disappointing that we don’t (yet?) know much about how the system actually works.
- Without a clearer sense of what’s under the hood, I feel less confident about both what it can and can’t do, and what new risks we may encounter. I doubt the world is ready.
- As ever, enthusiasts got an advance look; skeptics did not. That’s a sound marketing strategy, but it often turns out to be misleading. What we have often seen is initial enthusiasm that gets tempered over time. I suspect we will see that here as well.
- The new system appears to be less monitorable than prior systems, which is not great from a safety perspective. One really doesn’t want more capability in conjunction with less monitorability. But also more alignable, not sure why.
- Would be great to see whether Astra can make progress on any of the ten tasks that Miles Brundage and I bet on at the end of 2024. (No AI to date has succeeded on any, AFAIK.)