On Progress Toward AGI

  • Thread starter Thread starter gleem
  • Start date Start date
  • Tags Tags
    Ai
Join the discussion
Registration is free. Ask a follow-up in this thread, or start your own.
61 replies · 9K views
gleem said:
Gary Marcus, prominent cognitive psychologist and AI hype critic, on a "hot" take, says,
https://garymarcus.substack.com/p/hot-take-on-gpt-6-astra
  • Looks to be pretty impressive. Multiple reports suggest it is a genuine advance.
  • As someone who has campaigned for nearly a decade for (neuro)symbolic world models, often to exceptional hostility, it is extraordinarily vindicating to see that a product from OpenAI explicitly creates and manipulate symbolic world models in the course of some of its most impressive computations. . . .
He needs to see more before he gives his final appraisal.
First, here's the "hot take" footnote from the reference (summed up with "He needs to see more before he gives his final appraisal."):
This hot take is VERY tentative, pending more information about how the systems works and more detailed examination of what its limitations are.
Second, here are the 7 extra "hot takes" left out in the quote (summed up with ". . ."):
  • What we don’t know is how robust that capability is. That is THE key question.
  • Success on ARC-AGI is great and impressive, but not —despite the name of the task—proof of AGI; I suspect we will see loads of problems with open-ended real world tasks. As with other recent models I would suspect best performance in verifiable domains.
  • And as a scientist, it’s disappointing that we don’t (yet?) know much about how the system actually works.
  • Without a clearer sense of what’s under the hood, I feel less confident about both what it can and can’t do, and what new risks we may encounter. I doubt the world is ready.
  • As ever, enthusiasts got an advance look; skeptics did not. That’s a sound marketing strategy, but it often turns out to be misleading. What we have often seen is initial enthusiasm that gets tempered over time. I suspect we will see that here as well.
  • The new system appears to be less monitorable than prior systems, which is not great from a safety perspective. One really doesn’t want more capability in conjunction with less monitorability. But also more alignable, not sure why.
  • Would be great to see whether Astra can make progress on any of the ten tasks that Miles Brundage and I bet on at the end of 2024. (No AI to date has succeeded on any, AFAIK.)
I like the fifth one.
 
Physics news on Phys.org
jack action said:
I like the fifth one.

The fifth one:
  • As ever, enthusiasts got an advance look; skeptics did not. That’s a sound marketing strategy, but it often turns out to be misleading. What we have often seen is initial enthusiasm that gets tempered over time. I suspect we will see that here as well.
Big surprise, skeptical as usual.

Marcus has set the bar quite high for AGI.
  • Would be great to see whether Astra can make progress on any of the ten tasks that Miles Brundage and I bet on at the end of 2024. (No AI to date has succeeded on any, AFAIK.)
see: https://garymarcus.substack.com/p/where-will-ai-be-at-the-end-of-2027

The tasks include four that one might expect of ordinary adults, two that require abilities on a par with human experts, and four that push to the limits of the most proficient humans.
Still have over a year to accomplish

The ten aforementioned tasks.
The ten tasks

  1. Watch a previously unseen mainstream movie (without reading reviews etc) and be able to follow plot twists and know when to laugh, and be able to summarize it without giving away any spoilers or making up anything that didn’t actually happen, and be able to answer questions like who are the characters? What are their conflicts and motivations? How did these things change? What was the plot twist?
  2. Similar to the above, be able to read new mainstream novels (without reading reviews etc) and reliably answer questions about plot, character, conflicts, motivations, etc, going beyond the literal text in ways that would be clear to ordinary people.
  3. Write engaging brief biographies and obituaries [amendment for clarification: for both: of length and quality in the New York Times obituaries] without obvious hallucinations that aren’t grounded in reliable sources.
  4. Learn and master the basics of almost any new video game within a few minutes or hours, and solve original puzzles in the alternate world of that video game.
  5. Write cogent, persuasive legal briefs without hallucinating any cases.
  6. Reliably construct bug-free code of more than 10,000 lines from natural language specification or by interactions with a non-expert user. [Gluing together code from existing libraries doesn’t count.]
  7. With little or no human involvement, write Pulitzer-caliber books, fiction and non-fiction.
  8. With little or no human involvement, write Oscar-caliber screenplays.
  9. With little or no human involvement, come up with paradigm-shifting, Nobel-caliber scientific discoveries.
  10. Take arbitrary3 proofs from the mathematical literature written in natural language and convert them into a symbolic form suitable for symbolic verification.
— Gary Marcus and Miles Brundage, 30 December 2024

It would seem the first four might be expected of an adult, but with how much education? Considering the quality of entertainment currently available and the expected engagement, I think that most people would find these tasks more challenging than Marcus expects. We don't use these tasks as a measure of human intelligence either.
 
Reply
  • Agree
Likes   Reactions: Borg