Stubsack: weekly thread for sneers not worth an entire post, week ending 21st December 2025

BlueMonday1984@awful.systems · 3 months ago

Stubsack: weekly thread for sneers not worth an entire post, week ending 21st December 2025

blakestacey@awful.systems · 3 months ago

An academic sneer delivered through the arXiv-o-tube:

Large Language Models are useless for linguistics, as they are probabilistic models that require a vast amount of data to analyse externalized strings of words. In contrast, human language is underpinned by a mind-internal computational system that recursively generates hierarchical thought structures. The language system grows with minimal external input and can readily distinguish between real language and impossible languages.

corbin@awful.systems · 3 months ago

Sadly, it’s a Chomskian paper, and those are just too weak for today. Also, I think it’s sloppy and too Eurocentric. Here are some of the biggest gaffes or stretches I found by skimming Moro’s $30 book, which I obtained by asking a shadow library for “impossible languages” (ISBN doesn’t work for some reason):

book review of Impossible Languages (Moro, 2016)

Moro claims that it’s impossible for a natlang to have free word order. There’s many counterexamples which could be argued, like Arabic or Mandarin, but I think that the best counterexample is Latin, which has Latinate (free) word order. On one hand, of course word order matters for parsers, but on the other hand the Transformers architecture attends without ordering, so this isn’t really an issue for machines. Ironically, on p73-74, Moro rearranges the word order of a Latin phrase while translating it, suggesting either a use of machine translation or an implicit acceptance of Latin (lack of) word order. I could be harsher here; it seems like Moro draws mostly from modern Romance and Germanic languages to make their points about word order, and the sensitivity of English and Italian to word order doesn’t imply a universality.
Speaking of universality, both the generative-grammar and universal-grammar hypotheses are assumed. By “impossible” Moro means a non-recursive language with a non-context-free grammar, or perhaps a language failing to satisfy some nebulous geometric requirements.
Moro claims that sentences without truth values are lacking semantics. Gödel and Tarski are completely unmentioned; Moro ignores any sort of computability of truth values.
Russell’s paradox is indirectly mentioned and incorrectly analyzed; Moro claims that Russell fixed Frege’s system by redefining the copula, but Russell and others actually refined the notion of building sets.
It is claimed that Broca’s area uniquely lights up for recursive patterns but not patterns which depend on linear word order (e.g. a rule that a sentence is negated iff the fourth word is “no”), so that Broca’s area can’t do context-sensitive processing. But humans clearly do XOR when counting nested negations in many languages and can internalize that XOR so that they can handle utterances consisting of many repetitions of e.g. “not not”.
Moro mentions Esperanto and Volapük as auxlangs in their chapter on conlangs. They completely fail to recognize the past century of applied research: Interlingue and Interlingua, Loglan and Lojban, Láadan, etc.
Sanskrit is Indo-European. Also, that’s not how junk DNA works; it genuinely isn’t coding or active. Also also, that’s not how Turing patterns work; they are genuine cellular automata and it’s not merely an analogy.

I think that Moro’s strongest point, on which they spend an entire chapter reviewing fairly solid neuroscience, is that natural language is spoken and heard, such that a proper language model must be simultaneously acoustic and textual. But because they don’t address computability theory at all, they completely fail to address the modern critique that machines can learn any learnable system, including grammars; they worst that they can say is that it’s literally not a human.

Jayjader@jlai.lu · 3 months ago

Plus, natural languages are not necessarily spoken nor heard; sign language is gestured (signed) and seen and many, mutually-incompatible sign languages have arisen over just the last few hundred years. Is this just me being pedantic or does Moro not address them at all in their book?