Thursday, July 23, 2026

Sam Altman's comment - IV

AI developers speak often about how their software “learns,” “reads,” or “creates” just like humans. This gives you a sense that current AI technologies are far more capable than they are. Often when we talk to ChatGPT or its cousins like Gemini, we imagine that we're in the presence of some incredible new mind. But the reality is that these tools have a habit of fabricating answers that are statistically plausible, but in fact patently false

Suppose you ask an LLM something about orange juice. How does it produce a response when it has no awareness of a world with orange juice? The model is trained on a very large set of texts (such as books, articles, or conversations) of human speech. From these massive piles of human language, the model is more likely to produce the word 'hotel' or 'sweet' than the word 'dog' or 'cancer' when it is asked about orange juice. It is more likely to construct new sentences about orange juice alongside sentences about lunch, summer or sugar, than it would alongside sentences about stars, or protons, or elephants. 

But the mathematical patterns linking sentences are the sole contents of its model. It doesn’t actually think about who is asking the question, or what factors might motivate your query. It doesn’t think about anything at all. It simply activates a complex statistical model learned from its vast language corpus of human sentences that mentioned orange juice and predicts a new string of words as a likely extension of the same pattern. It may match what you want or it may be garbage. It has no relationship to the truth. Probable and accurate are not the same thing. 

You may think that these tools are animating an eerily lifelike image of humanity but that is not true. It's only reflecting that small subset of largely English speaking, digitally connected humanity that is most represented on the digital corpus of the internet. If you have the privilege of having the bandwidth and the tools to be able to type on Reddit all day, you get to be reflected in this mirror, but lots of people don't, and that's important.

Champions of AI will say that it can tell us if a pattern in our mammogram indicates cancerous tissue. It can reveal whether you're likely to suffer sepsis so that your doctors can intervene before it's too late. It can reveal the structure of the proteins that make up our biology and the biology of every other living thing. It can reveal new classes of antibiotics that we desperately need in order to treat dangerous diseases like antibiotic resistant MRSA or Staph. 

We can use it to find lead in pipes faster so that we can prevent children from being poisoned. We can use it to clear landmines, figure out where those landmines are so that people don't have to perform such dangerous work. But none of these benefits actually came from training AI tools on us. They came from targeted uses of machine learning and data about the world to solve very specific kinds of problems. That's actually quite different from what something like ChatGPT does. 

Generative AI models like ChatGPT have the habit of “making shit up”. In 2023, The Washington Post reported that ChatGPT had named a law professor as a sexual predator, telling a vivid yet entirely fictional story about his attempted assault of a student while on a class trip, and citing a nonexistent Washington Post article from 2018 as a source. Such errors are not because of malfunction. They are a feature not a bug. 

A study by Purdue University showed that ChatGPT gave incorrect answers to coding questions over half the time. Ironically, users often preferred the wrong answers to more accurate ones generated by knowledgeable humans, in part due to the confident style of the tool’s answers, accompanied by statements like, “Of course I can help you!” or “This will certainly fix it.” You can get them to apologize for things they haven't got wrong. 

The paper that Altman was mocking, “On the Dangers of Stochastic Parrots” explains that, like a parrot that not only repeats its owner’s vocalizations but produces random (stochastic) variations on the owner’s familiar pattern, these models parrot back to us variations on our own speech. They do so with just enough coherence and familiarity to project the illusion of understanding, and just enough randomness to surprise us and make us think we are hearing something new.

They can be thought of as the ultimate bullshitters, a term coined by a philosopher called Harry Frankfurt. In his 2005 book On Bullshit, he distinguished between a lie and bullshit. When a liar tells a lie, the intent is to deceive. They have to keep very careful track of the difference between truth and falsehood, in order to make sure you don’t get near the former. In contrast, the bullshitter is indifferent to the distinction between truth and falsehood. They'll just say whatever is convenient to them and whatever gets them what they want with no concern for the truth.

No comments:

Post a Comment