More Information about the Origin of Languages
There have been a couple of threads more or less recently, dealing with the question of "How and Where did Languages Originate".
The following is an excerpt of the few first paragraphs of an interesting article published today in the New York Times, dealing with the subject.
It does not seem to address the Babel theory, I'm afraid, but it makes for an interesting read nevertheless.
Languages Grew From a Seed in Africa, Study Says By NICHOLAS WADE Published: April 14, 2011
A researcher analyzing the sounds in languages spoken around the world has detected an ancient signal that points to southern Africa as the place where modern human language originated.
The finding fits well with the evidence from fossil skulls and DNA that modern humans originated in Africa. It also implies, though does not prove, that modern language originated only once, an issue of considerable controversy among linguists.
The detection of such an ancient signal in language is surprising. Because words change so rapidly, many linguists think that languages cannot be traced very far back in time. The oldest language tree so far reconstructed, that of the Indo-European family, which includes English, goes back 9,000 years at most.
Quentin D. Atkinson, a biologist at the University of Auckland in New Zealand, has shattered this time barrier, if his claim is correct, by looking not at words but at phonemes the consonants, vowels and tones that are the simplest elements of language. He has found a simple but striking pattern in some 500 languages spoken throughout the world: a language area uses fewer phonemes the farther that early humans had to travel from Africa to reach it.
9 Answers
I still haven't read the article carefully, but I think it is partially nonsense, and I'll explain why. Indo-European languages only had, according to the best theories we have, only 5 vowels (with length distinction). French has 14 and English 15, and both are newer languages. A language like Urdu has even more phonemes. While some languages in South America have very few phonemes, there are others as phonetically complex as English (often less vowels, but more consonants). Navajo, from North America, has more sounds than English, for example. It is true that Khoisan languages in Africa have a ridiculous amount of sounds, a lot of them are clicks, but in the Caucasus there are languages with nearly 90 phonemes, even without clicks, and many of them do not even exist in Africa - I wonder whether they have considered this.
A lot of languages have increased their number of sounds as they evolved. The guy who wrote the article seems to be ignoring these and other facts. After all, he is a biologist, not a linguist.
I can see some counter-arguments here (being informed in this topic), but it sounds interesting nevertheless. Thanks for the link. Please post more like this.
Muy interesante tu punto de vista, Iza. Yo también lo he leído, y me pareció un poco prepotente hacer todas esas asunciones. Esperaba escontrar un trabajo de más categoría y la sensación con la que me quedé es que el trabajo es una mera conjetura, y su título y la manera de anunciarlo, bastante confusos y ambiciosos.
Para mí, el problema es que está utilizando métodos inapropiados desde el principio. Me llevó un rato entender por qué estaba utilizando métodos evolutivos genéticos para estudiar analogías entre fonemas. Luego vi que el hombre lleva casi toda su vida intentando demostrar esto, y por cierto, se cita a sí mismo unas cuantas veces.
El asunto es que al utilizar este tipo de análisis estadístico de similitud para hacer clases, ni las diferencias ni las apariciones escasas de sonidos tienen importancia para el programa. El programa va a tener en cuenta solo el grado de similitud de las palabras, si entiende que son muy parecidas (Clase 1:here, hier) ( Clase 2: aquí aqui qui) concluirá que esos idiomas provienen de la misma raíz, y a partir de las clases extrapolará una supuesta raíz que comparará con otras. Como el programa solo se centra en lo que es similar, no es extraño que tarde o temprano todas las palabras converjan en el mismo gruñido. Es decir, el método que está utilizando le va a conducir siempre a un origen común. Para mí esta es la principal debilidad de su método.
El hacer esto con genes tiene cierto sentido, porque aunque solo hay cuatro bases distintas, (muchas menos que fonemas, por eso a Atkinson no le queda más remedio que obviar la diversidad en su trabajo) uno trabaja con fragmentos de entre 2000 y 40000 pares de bases, de manera que cuando algún fragmento de 130 pares de bases, por ejemplo, es idéntico a otro, no es una locura pensar que son idénticos "por algo". Pero hacerlo con palabras simples teniendo en cuenta unos cuantos fonemas, como hace Atkinson, en mi opinión tiene bastante de "reinterpretar la evolución".
Having read through the article published in Science, (along with its supplementary article, it seems to me that Atkinson may have been a bit premature in presenting such bold and sweeping correlations as he has. One of the biggest questions I have is in regard to the following premise which appears near the beginning of the paper and upon which Atkinson seems to base all of his subsequent ideas:
The number of phonemesperceptually distinct units of sound that differentiate wordsin a language is positively correlated with the size of its speaker population (Phoneme Inventory and Population Size, Hay & Bauer) in such a way that small populations have fewer phonemes.
Based on this assumption, Atkinson then concludes that it is reasonable to expect that small founder populations would have exhibited a degree of phonemic reduction, and because of this effect, a comparison of phonemic diversity to population size might allow one to use statistical analysis to trace the origin of each language back to a single source.
One of the main problems that I have with this conclusion is that it seems to take the stance that there is indeed authoritative evidence to support a positive correlation between population size and phonemic inventory. This, however, is not reflected by either the tone or data presented in the single source (Hay, et al) from which Atkinson quotes. Moreover, due to certain design flaws and inconsistencies in the original study, it seems that he has built his entire line of reasoning upon a premise that, despite its statistical basis, appears mostly speculative in nature (due to several design flaws in the analysis conducted by Hay). I will enumerate some of these flaws below:
?Information on vowels was taken from the Bauer (2007) handbook which lists information on some 250 languages
?Original material on consonants was unavailable from this source, so less information was collected on consonants (total of 216 languages).
?Sample is not randomized (we think it is a reasonably representative one). Instead, the set of languages is biased towards big languages and toward Indo-European and Pacific languages.
?Excludes, i.e. lists separately, the phenomena of tonal length, nasalized form, and diphthongs (which can be taken as sequences vowel and glide or sequence of non-identical vowels.
Taking the above factors into consideration, the study appears to be too flawed to affirmatively declare a correlation between phonemic richness and population, especially in light of the fact that 1) samples were not randomized and 2) sample size was extremely small (less than 3% of the 7,000 or so purported languages in existence). It also does not take into account any dead languages from which each living language was said to originate, and no effort is made to compare phonemic deterioration or fluxes that may have historically occurred between a contemporary language and its parent language. Also, the fact that the set of languages studied is biased towards Indo-European and Pacific languages means that African languages (the main focus of the subsequent study), were either not even examined or barely examined to compare phonemic data with population size. In addition, suppression of diphthong, nasal and tonal length data skews results in favor of languages without such features. Despite this fact, Atkinson still uses this flawed analysis on vocalic phonemes as a jumping off point for his paper and presumes that all three of thesetonal nature, vocalic inventory and consonant inventorywill demonstrate a similar positive correlation with population size.
Even if one were to take it for granted that the overall vocalic inventory of a language were indeed related directly to population size, this still does not exempt Atkinson from going beyond the scope of the evidence presented in assuming that both consonant inventory (which, due to certain inconsistencies, even the authors of the original paper cautioned against) and the use of tones (which was not even presented in the original paper) should likewise demonstrate a positive correlation with population size.
Furthermore, Peter Trudgill, a legitimate sociolinguist, has argued in the past that the effects of population size, network structure, and language-contact situation need to be considered together; therefore, there would be
no reason at all to expect to find a simple correlation between the numbers of speakers in a language and the number of phonemes in that language.
Keren Rice, who studied languages of the Athabaskan (Na-Dene) family (languages traditionally spoken across North America and which include such languages as Apache and Navajo), found them to have typically large consonant inventories. Looking at both the WALS data (considering only data in which tonal nature, consonant and vowel inventories was available) as well as the Ethnologue data on these languages and their populations (the two sources from which Atkinson draws his own data), one finds the following:
?Chipewyan (Canada): ConsonantsLarge; Population: Ethnic Population: 6,000
?Hupa (U.S.): ConsonantsLarge; Population Data: Ethnic Population: 223; Nearly extinct
?Navajo (U.S.): ConsonantsMod. Large; Population Data: Ethnic Population: 178,000; 7,600 monolinguals
?Slavey (Canada): ConsonantsModerately Large; Population Data: Ethnic Population 5,200 total; Declining
Looking at this data, it should be evident that despite extremely small population sizes, these languages have maintained a large inventory of consonants (which should appear counterintuitive when one considers their distances from where language is purported to have originated by the author of this paper).
Aside from what I have already mentioned, I also have a few other questions (such as why should it be assumed that there was only one original source of language, and why should it be assumed that languages were already both phonemically rich and fully developed before migration occurred or why should it be assumed that migrations only involved small groups of an originally large and relatively integrated population group, and why should it be assumed that factors affecting genetic stability and evolution (in terms of molecules of DNA) should be equated with those factors affecting the stability and evolutionary qualities of a given language), but I suppose that I will leave these points alone.
One thing that is worth noting is that on the same day, the scientific journal Nature also published a study by Michael Dunn (which, not incidentally, was coauthored by two of Atkinson colleagues at the University of Auckland in New Zealand) that looks at the origin of language from an ontological perspective. While this is not remarkable in itself, what is interesting is that his ideas seem not only to be supportive of those of Atkinson but also to be a direct attack on theories on language acquisition originally posited by noted linguists Noam Chomsky and Joseph Greenberg that, although controversial in their own right, have enjoyed relative popularity. For this reason, if Atkinson and Dunns theories do end up holding up to scrutiny, I imagine that it will likely only be after having confronted several robust and possibly acerbic challenges on the part of those who support the ideas of Chomsky and Greenberg.
That's intresting they are doing with DNA. According to what the only thing differeniates human language with animal sound is a difference in structure in a specific gene.
I found this interesting, too.
Thanks, Gekko!
Question: words fascinate me and I would love to study linguistics, I am currently learning Italian so that I can apply to a graduate school....do all linguistic studies start with the formation of words, or a language? Or do some study just the evolution of the romance languages? That is more what I am interested in....starting with the basics of Latin and moving into Italian, and so forth...
El hacer esto con genes tiene cierto sentido, porque aunque solo hay cuatro bases distintas, (muchas menos que fonemas, por eso a Atkinson no le queda más remedio que obviar la diversidad en su trabajo) uno trabaja con fragmentos de entre 2000 y 40000 pares de bases, de manera que cuando algún fragmento de 130 pares de bases, por ejemplo, es idéntico a otro, no es una locura pensar que son idénticos "por algo". Pero hacerlo con palabras simples teniendo en cuenta unos cuantos fonemas, como hace Atkinson, en mi opinión tiene bastante de "reinterpretar la evolución".
It is not the analogy between allelic frequency and phonemic frequency that gives me pause but rather the manner in which such data has been collected and interpreted.
In studies of population genetics, there are two very broad approaches: Empirical studies which both measure and quantify genetic variation; and theoretical studies which build upon patterns and variations observed in the empirical studies by using mathematical models in an attempt to explain the evolutionary processes which underpin such variations.
Now as I said before, I can understand the analogy between the frequency and characteristics of phonemes with the frequency and characteristics of genes within a population. Certainly, it can be observed that various languages have undergone phonemic reduction due to phonemic mergers (blending of distinct phonemes to form a single allophone); however, the main problem I see with the approach taken by Atkinson in this regard is that not only does it ignore the fact of phonemic splits (which have been observed in the transition from Germanic to Old English, resulting in increased phonemic richness where a "founder effect" should have been expected) but his study was also severely lacking in empirical data upon which to base his mathematical models.
For example, he does even take into account the characteristic qualities of each type of phoneme but describes, instead, a cline based on an average of various phonemic attributes. Even more interesting is the fact that the author purposely excludes as an outlier the Taa (!Xóõ) language of the Khoisan family spoken in Southern Africa. With its extremely large phonemic inventory (with at least 112 distinct phonemes, not including tones, and as many as 163 according to some sources), the Taa language is considered by many sources to be the most phonemically rich of all languages. If Southern Africa is to be taken as the birthplace of all languages (as suggested by Atkinson) then this should not be surprising; however, it is when one considers that such an assumption is predicated on the idea that phonemic richness is positively correlated with the speaking population that this becomes problematic.
When you consider that the !Xóõ language is spoken by only about 4,200 people worldwide, it becomes difficult to support such claims. Compare this to the fact that one of its sister Khoisan languages, Khoekhoe, is spoken by over 240,000 people worldwide (mostly in Namibia) yet its phonemic inventory is less than half as rich. Such evidence does not appear to support the corollary that total phonemic inventory can be so simply linked to population size (the basis for Atkinson's thesis).
Consider that one balanced survey of 451 languages found a median of 29 phonemic segments per language (not including tones), with 70% of these having between 21 and 40. Given such statistics, I find it difficult to imagine that a cline of any significance could have been ascertained. Compounding the problem is that Atkinson, rather than examine phonemic characteristics individually (as is done in population genetics with gene loci), appears instead to have consider each phonemic segment holistically, so that deficiencies in one phonemic category might be masked by another entirely unrelated category, a phenomenon which does nothing to explain whether or not such outcomes are a result of related evolutionary processes.
From the viewpoint of a geneticist, this would be analogous to calculating variations in seed color and flower color by considering organisms which no longer used flowers and seeds as a means of propagation. Said another way, genetic drift is explained in terms of decreases in genetic variation rather than terms of complete absence of genetic phenomena (i.e. it is typically used to explain allelic variations rather than the complete loss of alleles). My initial impression is that the stochastic processes involved in language development are distinct from those described by traditional theories on population genetics and as such, any computational algorithm should be more rigorously defined and be based upon actual observed patterns in language evolution. Just my opinion.
10% interesting and 90% silly. ![]()