- 14,792
- 7,438
By asking experts to give their educated guess estimates, say two years ago.PeterDonis said:Is it? How would one even determine that?
By asking experts to give their educated guess estimates, say two years ago.PeterDonis said:Is it? How would one even determine that?
I've just tried it:PeterDonis said:Please give a reference: where has this been done?
"Gamed" how? I honestly struggle to see what kind of evidence you would have to see to be convinced. Even if it had near 100% accuracy on whatever unambiguous question you threw at it, you would say "well it doesn't REALLY know the answer so it's basically just confusing you into thinking it gives correct answers". Well apparently it confused Terence Tao it gives useful answers in mathematics, pardon me for also being "confused"!PeterDonis said:This just shows that SAT tests can be gamed. Which we already knew anyway.
It is impressive because it can (sometimes) generate logical answers from text that it has never encountered before. This goes beyond parroting.russ_watters said:Impressive how? Doesn't this just tell you that it doesn't know the difference between fiction and reality, and more to the point, there's no way for you to know if it is providing you fictional or real answers*?
*Hint: always fictional.
Of course it is: noone is arguing that an LLM is not capable of frequently giving correct answers, or that a very well designed and trained LLM is not capable of giving correct answers within a large domain more frequently than many humans. The argument is that no amount of correct answers is equivalent to knowledge.AndreasC said:It seems like your argument is completely independent of whether or not it gives correct answers.
It is you that is wrong, and you are making claims for ChatGPT that its makers OpenAI don't make themselves.AndreasC said:you're just wrong and I'm imploring you to try it yourself.
Nobody is downplaying it, but in order to "recognize what is going on" it is necessary to understand what is actually going on. Noone can tell anyone else what to do but if I were you I would stop repeating my own opinions here and take some time to do that.AndreasC said:We can't be downplaying it like that because it's unfortunately going to become a significant part of the academic world, and people should recognize what is going on.
Well, I just tried it with Chat GPT. Its output was, alright. Not great, not terrible. It would be interesting to see what GPT-4 could do with it.PeterDonis said:Please give a reference: where has this been done?
Maybe you are not arguing that. But I don't think other people in this thread agree with you. Some people insist it is only "confusing" us into thinking the answers are correct. My argument is that the question of whether or not it "knows" is philosophical, and unrelated to practical considerations of whether or not it is reliable.pbuk said:Of course it is: noone is arguing that an LLM is not capable of frequently giving correct answers, or that a very well designed and trained LLM is not capable of giving correct answers within a large domain more frequently than many humans.
I will let others speak for themselves but I believe the only person that has used the term "confusing" in this thread is you.AndreasC said:Maybe you are not arguing that. But I don't think other people in this thread agree with you. Some people insist it is only "confusing" us into thinking the answers are correct.
Everyone is agreed on that, as @PeterDonis confirmed way back in #4:AndreasC said:My argument is that the question of whether or not it "knows" is philosophical, and unrelated to practical considerations of whether or not it is reliable.
PeterDonis said:The article is not about an abstract philosophical concept of "knowledge". It is about what ChatGPT is and is not actually doing when it emits text in response to a prompt.
The term "confusing" was not specifically used, but after I said it can give accurate answers to many questions, @PeterDonis in post #10 very specifically said it can't do that. They proceeded to say that it only sometimes gets "lucky" (I wouldn't call it luck exactly, it does it again and again for some subjects, you have to get UNlucky to get a wrong answer, then again it messes up more frequently on some other subjects) and gives an "answer". I don't know why "answer" was put in scare quotes but I believe it's probably due to scepticism that it even is an answer, and that it's not just me being confused. In the same post he argued that the only reason it passed tests was because of the "laziness and ignorance of the testers", presumably not because the answers were accurate.pbuk said:I will let others speak for themselves but I believe the only person that has used the term "confusing" in this thread is you.
I have very explicitly said I do NOT believe it is reliable multiple times. Specifically, in posts #3 (my very first on the thread), #5, and #13, plus in multiple other posts I have said again and again it often generates nonsense.pbuk said:I believe your misunderstanding is that because ChatGPT's answers are frequently correct that means that they are reliable.
Ah yes, I missed that. It seems we are in violent agreement.AndreasC said:I have very explicitly said I do NOT believe it is reliable multiple times.
Hahaha that is a very useful term online!pbuk said:Ah yes, I missed that. It seems we are in violent agreement.
Your take is weird to me, but it seems common, especially in the media. Consider this potential headline from 1979:AndreasC said:People often post more when it gets something wrong. For instance, people have given it SAT tests:
https://study.com/test-prep/sat-exam/chatgpt-sat-score-promps-discussion-on-responsible-ai-use.html
Sure, but the thing is, that it is able to do tasks that previous computer programs couldn't do. You couldn't copy and paste an SAT question into a program and get an answer before. It would require significant pre-processing, and in some cases you just wouldn't be able to get any help, because previous computer programs weren't good at, say, parsing natural language and taking into account context, subjective meaning etc. That is why it is impressive, because it accurately and quickly performs tasks that computers couldn't previously do, and were solely the domain of humans.russ_watters said:That's not impressive, it's a disaster. It's orders of magnitude worse than acceptable accuracy from a computer.
You could write that on the box of any new piece of software. Otherwise there's no reason to use it. But you're seeing the point now:AndreasC said:Sure, but the thing is, that it is able to do tasks that previous computer programs couldn't do.
Right. What's impressive about it is that it can converse with a human and sound pretty human. But now please reread the title of the thread. "Sounds human" is a totally different accomplishment from "reliable".AndreasC said:...previous computer programs weren't good at, say, parsing natural language and taking into account context, subjective meaning etc. That is why it is impressive, because it accurately and quickly performs tasks that computers couldn't previously do, and were solely the domain of humans.
What you show here is nothing like what AndreasC described.Demystifier said:I've just tried it
Exactly!PeterDonis said:What you show here is nothing like what AndreasC described.
But in post #13 you also said it can "repeatably" give accurate answers to questions. That seems to contradict "unreliable". I asked you about this apparent contradiction in post #15 and you haven't responded.AndreasC said:I have very explicitly said I do NOT believe it is reliable multiple times.
"ChatGPT Airlines - now 96% of our takeoffs have landings at airports!"russ_watters said:"New 'Spreadsheet' Program 'VisiCalc' Boasts 96% Accuracy - Might it be the New Killer App?"
ChatGPT is not parsing natural language. It might well give the appearance of doing so, but that's only an appearance. The text it outputs is just a continuation of the text you input, based on relative word frequencies in its training data. It does not break up the input into sentence structures or anything like that, which is what "parsing natural language" would mean. All it does is output continuations of text based on word frequencies.AndreasC said:previous computer programs weren't good at, say, parsing natural language and taking into account context, subjective meaning etc. That is why it is impressive
Or because the testers didn't bother writing a good test, that actually can distinguish between ChatGPT, an algorithm that generates text based on nothing but relative word frequencies in its training data, and an actual human with actual human understanding of the subject matter. The test is supposed to be testing for the latter, so if the former can pass the test, the test is no good.AndreasC said:the only reason it passed tests was because of the "laziness and ignorance of the testers", presumably not because the answers were accurate
See above.AndreasC said:the only reason it passed tests was because graders were "lazy"
Which, as I said, is already well known: that humans can pass SAT tests without having any actual knowledge of the topic areas. For example, they can pass the SAT math test without being able to actually use math to solve real world problems--meaning, by gathering information about the problem, using that information to set up relevant mathematical equations, then solving them. So in this case, ChatGPT is not going beyond human performance in any respect.AndreasC said:it only passed SAT tests because they can be "gamed"
"New from OceanGate: now 99% Reliable - Twice as Reliable as our Previous Subs!"Vanadium 50 said:"ChatGPT Airlines - now 96% of our takeoffs have landings at airports!"
I go back again to wondering what the creators are thinking about this...Vanadium 50 said:It's not just unreliable - we have no reason to believe it should be reliable, or that this approach will ever be reliable.
OpenAI's website is really weird. It is exceptionally thin on content and heavy on flash, with most of the front page just being pointless slogans and photos of people doing office things (was it created by ChatGPT?). It even features a video on top that apparently has no sound? All this to sell a predominantly text-based application (ironic)? The first section of the front page, though, contains one actual piece of information, in slogan form:pbuk said:Definitely not [AI], but they believe they are headed in the right direction:
If that's what a "good" test is, then it is tautologically true that GPT would be no good at them. The issue with tautologies is, of course, that they don't tell us anything new. What is new is that GPT can do many things that only humans with understanding could previously do. Of course it doesn't do them perfectly, but often it does them more accurately than most humans, and much faster. If what you want is the answer to an exercise, and it can give you the correct answer, say, 99% of the time, then that's good enough for many people and in many contexts, regardless of philosophical questions about understanding. And again, we are talking about things that computers previously just couldn't do. This is why it is significant and this is why I'm saying it should not be downplayed, because we will encounter this way too much in coming years.PeterDonis said:Or because the testers didn't bother writing a good test, that actually can distinguish between ChatGPT, an algorithm that generates text based on nothing but relative word frequencies in its training data, and an actual human with actual human understanding of the subject matter
Well, @Demystifier didn't do what I described. See my post where I tried it.PeterDonis said:What you show here is nothing like what AndreasC described.
I think they are planning to monetize this by first making a name for themselves and then selling a product where "close enough is good enough". For example, customer service chatbots.russ_watters said:go back again to wondering what the creators are thinking about this...
Is it?AndreasC said:If what you want is the answer to an exercise, and it can give you the correct answer, say, 99% of the time, then that's good enough for many people and in many contexts
But what if you want the answer as if given by Homer Simpson, or a Shakespearian Sonnet? Alpha cant do that ;)PeterDonis said:Is it?
Perhaps if my only purpose is to get a passing grade on the exercise, by hook or by crook, this would be good enough.
But for lots of other purposes, it seems wrong. It's not even a matter of percentage accuracy; it's a matter of what the thing is doing and not doing, as compared with what my purpose is. If my purpose is to actually understand the subject matter, I need to learn from a source that actually understands the subject matter. If my purpose is to learn a particular fact, I need to learn from a source that will respond based on that particular fact. For example, if I ask for the distance from New York to Chicago, I don't want an answer from a source that will generate text based on word frequencies in its input data; I want an answer from a source that will look up that distance in a database of verified distances and output what it finds. (Wolfram Alpha, for example, does this in response to queries of that sort.)
Exactly, here is the problem!PeterDonis said:Perhaps if my only purpose is to get a passing grade on the exercise, by hook or by crook, this would be good enough.
I think they already do it the Max Power way:BWV said:But what if you want the answer as if given by Homer Simpson, or a Shakespearian Sonnet? Alpha cant do that ;)
If I know that's what your business is doing, you won't get my business.AndreasC said:what happens when some business or gover does the math and figures it would rather risk being wrong than pay experts?