New blogs

Leherensuge was replaced in October 2010 by two new blogs: For what they were... we are and For what we are... they will be. Check them out.
Showing posts with label autosomal DNA. Show all posts
Showing posts with label autosomal DNA. Show all posts

Wednesday, August 25, 2010

A couple of new genetic papers I'd love to read


But by the moment just mentioning them briefly.


Both are published by the European Journal of Human Genetics (Nature) and both are mentioned by Diekenes today (link 1, link 2).


African genetic structure shows novel elements

There has been some negligence in mapping the genetic structure of African populations, with most papers taken a West African proxy (typically Nigerians), sometimes enriched with some of the last hunter-gatherers of the continent, to represent the whole complexity of the ancestral continent.

This paper seems to address this lack.

Martin Sikora et al., A genomic analysis identifies a novel component in the genetic structure of sub-Saharan African populations. EJHG 2010. Pay per view.

Abstract

Studies of large sets of single nucleotide polymorphism (SNP) data have proven to be a powerful tool in the analysis of the genetic structure of human populations. In this work, we analyze genotyping data for 2841 SNPs in 12 sub-Saharan African populations, including a previously unsampled region of southeastern Africa (Mozambique). We show that robust results in a world-wide perspective can be obtained when analyzing only 1000 SNPs. Our main results both confirm the results of previous studies, and show new and interesting features in sub-Saharan African genetic complexity. There is a strong differentiation of Nilo-Saharans, much beyond what would be expected by geography. Hunter-gatherer populations (Khoisan and Pygmies) show a clear distinctiveness with very intrinsic Pygmy (and not only Khoisan) genetic features. Populations of the West Africa present an unexpected similarity among them, possibly the result of a population expansion. Finally, we find a strong differentiation of the southeastern Bantu population from Mozambique, which suggests an assimilation of a pre-Bantu substrate by Bantu speakers in the region.

This Mozambican specificity was already spotted accidentally in Patin's paper on Pygmy Genetics last year but himself downplayed its importance because of his focus on Pygmy structure specifically.


West European R1b is distinct

Nothing really new for those who have kept a keen eye on the research of this Y-DNA haplogroup, the most characteristic of West Europe. But I'd still like to know more about the details. I doubt I could concur with the suggested timeline in any case.

Natalie M. Myres et al., A major Y-chromosome haplogroup R1b Holocene era founder effect in Central and Western Europe. EJHG 2010. Pay per view.

Abstract

The phylogenetic relationships of numerous branches within the core Y-chromosome haplogroup R-M207 support a West Asian origin of haplogroup R1b, its initial differentiation there followed by a rapid spread of one of its sub-clades carrying the M269 mutation to Europe. Here, we present phylogeographically resolved data for 2043 M269-derived Y-chromosomes from 118 West Asian and European populations assessed for the M412 SNP that largely separates the majority of Central and West European R1b lineages from those observed in Eastern Europe, the Circum-Uralic region, the Near East, the Caucasus and Pakistan. Within the M412 dichotomy, the major S116 sub-clade shows a frequency peak in the upper Danube basin and Paris area with declining frequency toward Italy, Iberia, Southern France and British Isles. Although this frequency pattern closely approximates the spread of the Linearbandkeramik (LBK), Neolithic culture, an advent leading to a number of pre-historic cultural developments during the past ≤10 thousand years, more complex pre-Neolithic scenarios remain possible for the L23(xM412) components in Southeast Europe and elsewhere.

M412 seems to be a novel SNP not yet reported at ISOGG. S116 (defined as its major subclade) used to describe R1b1b2a2, most diverse around the Pyrenees (unless this paper says the opposite). So I doubt the LBK hypothesis can hold, regardless of frequency.

Saturday, July 31, 2010

Korean genetics: between China and Japan with some unique personality


New paper on Korean genetics:


Jonshung Jung, Hoyoung Kang et al. Gene Flow between the Korean Peninsula and Its Neighboring Countries. PLoS ONE 2010. Open access.

The most relevant results are probably in these graphs that follow:

Figure 3: Genetic Structure

Legend and details for the various samples can be seen in figure 2. Notice please that CB (Cambodian) is mistakenly placed in Northern Asia, when it should be in Southern Asia (i.e. SE Asia) along with Vietnamese (VN) and Vietnamese-Koreans (VC probably though elsewhere tagged as VK).

Notice also that there are two K=5 in the global cluster analysis (C), one of them marked as RH, what means "recombination hotspot", which is a technical albeit interesting matter they deal with in extent in the supplementary materials. The conclusion seems to be that these recombination hotspots can induce distortions in the cluster analysis and that should be dealt with in order to prevent confusing results.

Anyhow the most valid run for the global dataset is surely K=4, showing four neatly distinct clusters: Africans, Europeans (or West Eurasians), Amerindians and East Asians. The minor "admixture" apparent levels among Amerindians and some East Asians (Mongolians, Cambodians) may or not mean admixture. In my opinion, based on comparison with other different studies, I think it does not but rather indicates a lesser degree of affinity with the main cluster and therefore some small degree of affinity with other Eurasians. Careful sampling strategies, such as the one done by Hui Li last year may be needed to discern what exactly means, if anything at all.

Most probably the "European affinity" apparent in these two cases just means some Central Eurasian (or Siberian) affinity of Mongols, as shown in Hui Li's paper, as well as in Amerindians, and some South Asian affinity of Cambodians as detected for neighboring Thais in the recent paper by Jinchuang Xing.

When comparing East Asians alone (B), three clusters are apparent (K=3): a "Mongolian" one (blue), a SE Asian one (green), more intense in Cambodians than Vietnamese, and the middle East Asian one including Chinese, Korean and Japanese (and largely Vietnamese too). K=4 seems to indicate a diffuse (low level) Japanese-Korean affinity but this is better seen in the middle East Asian comparison.

In panel A (with only Chinese, Koreans and Japanese), we can see three clusters again (K=3): Chinese (green), Japanese (red) and Korean (blue). However most Koreans are not clearly differentiated from their neighbors, specially Chinese, and only the insular population of Jeju shows a much stronger Korean-specific homogeneity (though still with some Japanese influence).

Figure 4: MDS and NJ Tree of Korean, Chinese and Japanese


MDS stands for multidimensional scaling, a way of visually presenting statistical data in two or more dimensions, somewhat similar to PCA. NJ stands for neighbor-joining, a method used to build affinity trees in genetics that you are probably familiar with.

According to the legend, in the MDS plot (A) the 1st (horizontal) dimension includes 90% of the variance, while the 2nd (vertical) dimension only 1%! It could perfectly be a linear plot with most East Asians clustering to the left and a small odd group to the right. This "odd group" is mostly made of SE Koreans, plus three Japanese Koreans (from Kobe) and one Kobe Japanese.

However this must be an error of the legend and the vertical axis is without doubt the 1st dimension. Notice the scale, which is of 1/1000 order of magnitude in the horizontal axis and only 1/10 in the vertical one. Notice also how a China-Korea-Japan, with some eccentricity for Korea, is fully consistent with all other data provided.

The NJ tree indicates an apparent first divergence in three branches:

1. A Japanese-only branch (most Japanese fall here)
2. A mostly Korean branch, including also one Japanese and three Chinese
3. The major branch including most Koreans and Chinese, as well as two Japanese. This one, in turn splits in:

3a. Including the two remnant Japanese and all the rest being Koreans.
3b. A large branch including only Koreans and Chinese. It divides in two:

3b1. Almost only Koreans (only one Chinese here I think), with high incidence of SW Koreans.
3b2. Many Koreans too (with high incidence of SE Koreans) plus most Chinese, concentrated in a particular sub-branch (bottom of graph).

Caution must be placed to Korean-centric readings in any case because Koreans are clearly oversampled, what is surely distorting the tree structure to at least some extent. However if this structure could find confirmation in further more balanced studies, it might well support a colonization process along the coast, which is pretty much mainstream these days.

Koreans anyhow mostly cluster with Chinese here too, which is consistent with the STRUCTURE analysis, showing some SW-SE (West-East?) internal polarity in the peninsula.

Notice that no sampling was undertaken, without doubt because of the political circumstances, in the northern half of the peninsula. Notice also that both Chinese samples are from the North (Beijing and Manchuria).

Thursday, July 22, 2010

New paper on human autosomal phylogenetics


A reader points me to a new quite interesting paper on human phylogeny from the viewpoint of autosomal DNA mainly.

Jinchuang Xing et al. Toward a more uniform sampling of human genetic diversity: A survey of worldwide populations by high-density genotyping. Genomics 2010. Pay per view.

A copy can be found at ZohoViewer and the supplementary material is also freely available.

Abstract

High-throughput genotyping data are useful for making inferences about human evolutionary history. However, the populations sampled to date are unevenly distributed, and some areas (e.g., South and Central Asia) have rarely been sampled in large-scale studies. To assess human genetic variation more evenly, we sampled 296 individuals from 13 worldwide populations that are not covered by previous studies. By combining these samples with a data set from our laboratory and the HapMap II samples, we assembled a final dataset of ~ 250,000 SNPs in 850 individuals from 40 populations. With more uniform sampling, the estimate of global genetic differentiation (FST) substantially decreases from ~ 16% with the HapMap II samples to ~ 11%. A panel of copy number variations typed in the same populations shows patterns of diversity similar to the SNP data, with highest diversity in African populations. This unique sample collection also permits new inferences about human evolutionary history. The comparison of haplotype variation among populations supports a single out-of-Africa migration event and suggests that the founding population of Eurasia may have been relatively large but isolated from Africans for a period of time. We also found a substantial affinity between populations from central Asia (Kyrgyzstani and Mongolian Buryat) and America, suggesting a central Asian contribution to New World founder populations.


Fig. 3 click to expand

The abstract already addresses which are the most important conclusions of the paper: (1) lower genetic distances with better sampling strategies, (2) claim of large distinct founder population at the origins of the Out of Africa migration and (3) claim of greater affinity of Native Americans with Central Asians than with East Asians senso stricto. Additionally they also emphasize (4) the finding that the West Eurasian component in South Asians is of West Asian origin rather than European.

In the graphs I have noticed a couple of other details worth of mention: (5) that Pygmies appear more distinct than Khoisan from the bulk of the species (which is somewhat contradictory with the haploid phylogeny) and (6) that the closest African populations to Eurasians are "Nilotic" groups of the Kenya-Uganda-Ituri area (neither the Horn of Africa nor the Nile Basin were sampled).

I will address some of these matters now.


The migrant Out of Africa population

The authors take some time to address the issue of the migrant population in pages 20-21:

The OoA hypothesis, proposing a single OoA bottleneck followed by an expansion into Eurasia approximately 50,000 years ago, has gained extensive support from the archaeological record and genetic studies. Nevertheless, many of the historical details of this diaspora remain unclear. A common interpretation is that the OoA bottleneck was the result of a migration of a small founding population into Eurasia. Given the difference in haplotype heterozygosity between African and non-African populations and the relationship between heterozygosity and effective population size, we can estimate the effective population size of such a founding population . Within Africa, the average 100-kb haplotype heterozygosity in our data is 0.91. Immediately outside of Africa in Europe, the Middle East, and Central Asia, the average haplotype heterozygosity is 0.82 (Figure 2). A reduction of heterozygosity from 0.91 to 0.82 in a one-generation bottleneck would require an effective population size of only 5.5 individuals. While a one-generation bottleneck is an oversimplification, these estimates indicate that an OoA bottleneck resulting from the migration of a small founding population would require an extremely small population size. However, given that the archaeological record indicates a rapid expansion of modern humans into Europe and Asia in just a few thousand years , it seems unlikely that Eurasia could be populated so quickly by a such a small founding population.

A more likely explanation for the OoA bottleneck is that Eurasia was populated by a larger population that had been relatively isolated from other modern human populations for tens of thousands of years prior to the expansion. The first fossil evidence for modern humans outside of Africa is in the Middle East at Skhul and Qafzeh between 80,000-100,000 years ago, which is at least 20,000 years prior to the Eurasian diaspora. If a population of modern humans remained in the Middle East until the expansion into Eurasia, there would have been sufficient time for genetic drift to reduce heterozygosity dramatically before the Eurasia expansion. This “Middle East isolation” hypothesis provides a robust explanation for the relative homogeneity of European and Asian populations relative to African populations (see Figures 3A-B) and is supported by a recent maximum likelihood estimate of 140,000 years ago for the time of Eurasian-West African population separation. Interestingly, a recent study of the Neandertal genome suggests that the non-African individuals, but not the Africans, contain similar amount of admixture (1-4%) with the Neandertals. The authors suggest that the admixture must have happened between the Neandertals with an ancestral non-African population before the Eurasian expansion. Given the fossil, archaeological, and genetic evidence, the Middle East isolation hypothesis warrants rigorous evaluation as whole-genome sequence data become available.

I must say that the real problem is to be talking of a mere depth of 50,000 years for H. sapiens colonization of Eurasia, when that must be the date of the reflux into West Eurasia. The archaeological record for Asia east of Iran is inconclusive (too poor) and the genetic data, including the one available here, strongly suggests that South and East Asia were colonized before West Eurasia.

Hence we must be talking of a quite greater time depth such as the 75-80,000 years ago or more, as has been suggested by most population genetic analysis as of late. Certainly nothing less than 60,000 years ago minimum.

The assumption the authors make is therefore wrong so it's likely that the conclusion is also wrong.

That doesn't mean that the considerations they make, specially those regarding a very small colonizer population do not make sense. This small group of adventurous colonists could perfectly have colonized Asia with much more time, leaving very few remains precisely because they were few and even when they grew up in numbers they were still not many. The relatively poor situation of Asian archaeology does not help to unravel the case in either direction but we must remember that the Jawalpurram remains have clear African MSA affinities (and hence are likely to be product of our species) and these date from before the Toba event, which could well have also helped in the reduction of Eurasian heterozygosity even more, some 74,000 years ago. And there are other archaeological clues that, while not clearly conclusive, may suggest an expansion into Asia since as early as c. 110,000 years ago.

Sure, it would be also a good idea to ponder carefully about the role of the Middle Paleolithic colonists of Palestine in the whole process if that is possible. I have nothing against that but I still don't like their reasoning in this point.


The branching out of Eurasians and the two South Asian components

The neighbor joining trees (see fig. 3 above and also fig. S1 at the supplemental materials, very similar) are one of the most interesting results of this paper and the authors are clearly proud of them.

I am going to ignore this "detail" hereafter but I must however mention that the tree produced in fig. S2, after the inclusion of a North African and two Palestinian populations is however very different. This is strange but I don't know how to handle this discrepancy. It might be a point of support for their hypothesis of a long separate coalescence in the Levant? Can't say.

The two other trees however really produce a result that is an almost perfect fit with haploid phylogenies, with Eurasians branching in two in Tropical Asia (South and East Asian branches) first of all.

Then the South/West Eurasian branch shows a division between South Indians and the rest, what I interpret as a split happening still in South Asia prior to the colonization of West Eurasia. Then Pakistanis and West Eurasians branch apart and then the same happens with Europeans diverging from the West Asian/Caucasus population.

Some of the branches' positions however may be caused by ulterior admixture so let's be careful with that.

The authors also emphasize the finding (consistent with what we have seen in other papers) that the second South Asian component, related to West Eurasians, is essentially of West Asian/Caucasus affinity and not European.

I agree with this and I think that it is an important point to make. It seems to imply that an important genetic flow has existed from West Asia into South Asia, specially the Northwest part of it. Of course the flow may have happened at different historical and prehistorical periods but it is important to realize that the Neolithic Age was surely when such migrations might have caused a greater impact.

In contrast some of "European" (darker orange) component is also visible, maybe originating in the Indoeuropean flows and maybe replaceable by a more specific Central Asian component (sadly Central Asia and Siberia is only sparsely sampled in this paper) if the findings of Hui Li are to be reproduced in the context of proper sampling strategies in this delicate area. Whatever the case the European input in South Asia is very minor, even if maybe slightly larger than among West Asians/Caucasians. We can safely infer, I understand, that it reflects the real Indoeuropean genetic input via Central Asia.

Most importantly a clearly distinct South Asian component (purple) has been detected and is strong enough to make up 50% of the Pakistani gene pool and almost the totality of some South Indian populations. Also notice the distinctive Irula component (blue), which may reflect the particular long isolation of these tribals, in the past tentatively classified as "Negritos".

Notice also the minor but significant presence of the Indian component in SE Asia, specially in Thailand. I have on occasion noticed that some Thais seem to have a distinctive phenotype and maybe this is the explanation.


East Asians and Native Americans

In this aspect I want to say that I am not totally persuaded by the authors' claim of greater Central Asian affinity of Native Americans. The main reason is that the "Central Asians" they mention such as Nepalese or Kyrgyzes are possibly admixed populations that owe their position in the NJ tree to that fact.

Even the Buryats appear to show some of that admixture. In this case (and maybe in the others too) it is probably a case of Central Asian specific components indeed but components that still may reflect a very ancient admixture event in the early Upper Paleolithic process of colonization of Central Asia and the Far North.

This is a limitation of this paper: they do some chest beating about a very throughout sampling (somewhat justified indeed) but in the case of Central Asia/Siberia they are lacking and the matter seems to be left unclear.

In any case, Native Americans or rather their ancestral founder population does look like having coalesced in a complex Central Asian and Siberian sparsely populated ancient landscape prior to their arrival to Beringia and subsequent colonization of America. Haploid genetics is very strongly supportive of such scenario.

It is difficult to ascertain however whether their high divergent location in the NJ tree, in the context of the East Asian branch, owes to them having diverged very early or rather (as I suspect) to their early admixture event, maybe partly shared with Central Asians and Siberians. We would need a much improved sampling strategy in those areas to be able to get some clear ideas.

Otherwise East Asians appear to show a first division between NE Asians and SE Asians, with the divide running across China. Not much more can be said, as the sample has not sufficient coverage, specially in SE Asia and Oceania.


African curiosities

One of the details of the trees that called my attention is that, in contrast to what happens in simplified haploid genetics, Pygmies are more distant from the rest of Humankind than Khoisan. This has an explanation, I believe, as the lineages more tightly associated with the Khoisan such as mtDNA L0 and Y-DNA A have representatives in NW Africa and even Arabia, indicating a protracted divergence (or repeated re-convergence) between the southern proto-Khoisanid branch and the main proto-Afro-Eurasian one. Instead when proto-Pygmies diverged they probably did for good, in spite of recent admixture with Bantus and some ancient lineages also shared with West Africans at minority levels.

Another such detail is that the populations most closely related to Eurasians are East Africans (Hema, Luhya, Alur). Overall the African branching process is coherent with the scenario I described here at Leherensuge some months ago.

Saturday, July 3, 2010

Central Eurasian genetic specifity detected


Found at
Dienekes and GNXP:

Hui Li et al., Genetic Landscape of Eurasia and “Admixture” in Uyghurs. AJHG 2009. Now freely available at PubMed.

The interesting stuff is in figure 1:


So we have now a cluster centered at the Khanty (an East Uralic population) but strong in Central Asians that is distinct from both West and East Eurasians.

I think this is the coolest discovery of autosomal genetics in quite a long time.

Dienekes protests about the inconsistence of this with Y-DNA but I fail to see the connection, because Y-DNA or even mtDNA, surely scattered in the early Upper Paleolithic when Eurasians were still very much undifferentiated, while these components are surely shallower and represent regional homogeneization processes that happened surely only after the LGM. Also the impact of Y-DNA flows may be really weak: Buriats and Finns share Y-DNA lineage, as do Polish and many Indians or West Europeans and some Central Africans but it's obvious that these don't correlate too well with autosomal genetic clustering which has its own processes of regional homogeneization, while haploid genetics and specially Y-DNA is subject to high odds of fixation by mere drift and founder effects.

Razib says that the authors seem to be arguing for greater number of (strategically chosen?) populations, what makes total sense and has served, when done, to add many shades to the tricomy of the old HapMap continental hyper-simplification.

However it's very possible that a deeper cluster analysis would have revealed futher clusters, maybe an Indian one (very small sample though) or a distinction between the Khanty and Central Asians or...

Where does this Central Eurasian cluster stem from. I'd say, based on haploid genetics, that it was probably created by the admixture of a West Eurasian and East Eurasian migration converging in that peripheral area and then coalescing into a rather homogeneous cluster on its own right. That would explain why the population appears almost exactly intermediate between the two major continental groups and also why there is no haploid lineage that specifically belongs to that region (all are shared with either West Eurasia or East Asia).

East Asian variations

It is noticeable that three quite distinct populations emerge in East Asia: inland, coastal and south. These are probably essentially comparable with the blue, yellow and one (or all together) of the SE Asian components detected in the HUGO Consortium paper.

Thursday, June 10, 2010

Jews are "Phoenicians", Palestinians are "Jews"





[Update: please read my 2012 independent analysis of Western Jews, Palestinians and other populations of West Asia, which adds some interesting nuances but in essence confirms what is said here].


This does not seem just another paper on Jewish genetics but more like The Paper. While it is pay per view and hence I haven't been able to read it in full , the material I could see at Dienekes' blog (same as in the supplementary material) is most revealing. I think this research will mark a before and after in Jewish genetics and also gives some interesting hints on other populations, specially in West Asia and North Africa.

Doron M. Behar, Bayazit Yunusbayev, Mait Metspalu et al. The genome-wide structure of the Jewish people. Nature, 2010. Pay per view but supplementary material freely accessible.

One of the good things is that finally comparison with Turks and other populations from that area where early Jewish Diaspora in the Hellenistic-Roman era is known to have lived, rather than in Palestine, in a time when Judaism (several sects including eventually Christianism) was still actively proselytizing.

Unlike what I used to think, Western Jews (Sephardi and Ashkenazi) do not cluster too well with the Turkish sample but they cluster almost perfectly with Cypriots and Lebanese. They seem to have no particular relation with Palestinians (nor Druze nor Negev Bedouins) but these also appear clearly different from other Arabs and in general any other sampled population.


This K-means analysis is for me the answer to all these endless discussions on the origin of Jews and Palestinians. Western Jews seem essentially to have coalesced in the Cypriot-Phoenician area probably by, essentially, conversion. Palestinians seem to be a uniquely distinct population, albeit somewhat admixed with their neighbors, which may well originate with the local Neolithic and certainly must have been there in early historical times.

In other words: Palestinians are most likely to be the true descendants from the Jews of the Biblical period, rather than modern Western Jews who seem more as originating from a Phoenician-Cypriot population which converted to Rabbinic Judaism for whichever reasons.

Other Jewish populations also seem to originate in conversion episodes but from different genetic pools. Hence Yemeni Jews appear as genetically Arab, Ethiopian Jews as Ethiopian, Indian Jews as Indian and Iraqi-Iranian Jews as Iranians. There may be some fine threading to enrich this overall picture but the essentials seem very clear in any case.

Anyhow, Moroccan Jews appear as just another branch of Western Jews (no particular relation with Moroccans is apparent) and I have not been able so far to identify the closest population to the small sample of Uzbek Jews. Also Turco-European Jews seem at least somewhat admixed with Europeans, something that was already well known.

As for other populations, I find interesting that an specific autosomal component of NW Africans has been detected (for the first time as far as I can tell). Many specific clusters have obviously not been detected because of the relative shallowness of the K-means analysis.

The supplementary materials have other interesting graphs, PC analysis of autosomal, Y-DNA and mtDNA genetics and a global K-means analysis of relevance when observing some peripheral Jewish populations specially but also providing some general hints on other populations (a very wide sample, specially in Eurasia).

I must cheer and congratulate the authors of this research for finally addressing the debate on Jewish (and Palestinian) origins with a worthy extended sample of many many populations, including key ones such as Turks, Cypriots and Lebanese (among others). The wording of the abstract doesn't say things as clearly as I do but their data speaks volumes.



Update: A reader, Joe, tells me that H.G. Wells already suggested this Phoenician true origin of Jews. In his book A Short History of the World, chapter XXII, he wrote, speaking of Semitic peoples, once all powerful but then defeated by the Indoeuropeans (Persians, Greeks, Romans):

Is it any miracle that in their days of overthrow and subjugation many Babylonians and Syrians and so forth, and later on many Phoenicians, speaking practically the same language and having endless customs, habits, tastes and traditions in common, should be attracted by this inspiring cult [Judaism] and should seek to share in its fellowship and its promise? After the fall of Tyre, Sidon, Carthage and the Spanish Phoenician cities, the Phoenicians suddenly vanish from history; and as suddenly, we find not simply in Jerusalem but in Spain, Africa, Egypt, the East, wherever the Phoenicians had set their feet, communities of Jews.


See also the discussion on Atzmon 2008.


Update Nov 21: a free copy of the paper is available here (PDF).

Wednesday, May 26, 2010

New paper on Spanish genetics


There's a new paper of some interest on Spanish population genetic structure:


J. Gayan et al. Genetic Structure of the Spanish Population. BMC Genomics. Open access.

The main aim seems to be to provide a Spanish dataset for population genetic research. And so far so good.

However the sampling strategy is awkward to say the least:

In the above map (my creation on the paper's data), red dots indicate sampling locations, while blue areas are regions (autonomous communities) not included in the sample. Most noticeable is that not just the Basque Country has been excluded but also all the surrounding area.

As I say, quite awkward.

Other surely distinctive unsampled areas are the Canary Islands and Galicia.

The whole design of the sample has a Castile-centric bias that is difficult to understand.

But, well, that's what they did. And, once we know that, we can go on to look at the results:

In this graph (fig. 6 annotated by me) we can see how Catalans and Andalusians tend to diverge from neutrality in orthogonal directions. To a less clear extent, the North Castilian samples (Arévalo, Segovia) also diverges somewhat.

We can say that PC1 describes a Catalan-other axis and PC2 an Andalusian-other one. The lack of distinctiveness of some geographically eccentric samples such as the Asturian one (Avilés)may well be caused by the small size of the sample. It is very possible that a PC3/PC4 graph would evidence some distinctiveness that is not apparent here.

Remember that PC graphs are merely bidimensional representations of some of the apparent structure, with all the limitations that this implies.


European comparison

When compared with other populations of European ancestry the PC graph is as follows (annotations by me on fig. 7):


Catalans appear to have some tendency towards Italy and NW Europe, while the less defined eccentricity of Andalusians only seems to tend towards Italy. There are also a couple of Castilian individuals who cluster with NW Europeans, maybe because the North Castile area sampled was the core of Visigothic settlement (just a hunch).


Global comparison

There's not much to say about Spaniards in the global scatterplot (fig. 8), really: all Europeans just cluster very tightly, the same as East Asians (Chinese and Japanese).

What I found intriguing and worth posting this graph is the curious coincidence of the scatter of Kenyan Maasai (MKK) and US African-Americans (ASW). Notice that the LWK sample (Luhya) are also from Kenya but Bantu and they cluster best with Nigerian Yoruba (YRI). Not really sure because it'd need further research but certainly the almost identical disposition of MKK and ASW samples is suggestive of the Maasai (and maybe other Nilotes) being somewhat admixed with West Eurasians. Alternatively the distribution might be reflecting some African-specific differences (just like Indians in the Eurasian context) and its overlap with African-Americans is to some extent an artifact of the limitations of PC analysis.

I do miss a comparison with North Africans, which seems to be a taboo in Iberian and European population genetic studies. However I do detect (and not only here) a slight "African" tendency among some Iberians which may well reflect a greater affinity with North Africans, in turn slightly more akin to ultra-Saharan Africans. This would be an interesting matter to analyze.


Update: I totally forgot to mention that just a few days ago I commented on another paper by a Catalan team that did compare Iberians and North Africans, which may serve for comparison. The results appear wildly different in the PC graph, with Catalans, Basques and Cantabrians clustering on one corner and the other Iberians scattered with a clear tendency towards the Eastern Mediterranean. No apparent North African affinity was detected.

Tuesday, April 27, 2010

East Asian autosomal genetics, second round


I already commented on the
HUGO paper (paywall, but full PDF here) before as a "working note". Recent discussion at another blog has got me working again on it.

I ended up with this map reflecting as well as possible the geographical distribution of the various components:


Names (arbitrary but descriptive) are my creation, as are the arrows illustrating possible gene flows. Question marks indicate where there's lack of data. Continuous lines indicate "pure" areas, dotted lines indicate greater or thinner presence of each component.

I suspect that the blue component represents the Neolithic stock, spreading first autonomously and then also (but more limitedly) in admixture with other components. If so, the other major components may have receded before it or may have just absorbed each other in mutual interactions.


Observations

The expansion of the Peninsular component to Indonesia did not carry the Neolithic component. So either is pre-Neolithic of was made by non-mixed Austroasiatic peoples after getting the "neolithic package" from further north without getting the genes.

The expansion of the Neolithic component to sub-Hymalayan South Asia did not carry other Eastern components, so it happened before or, in any case, in a separate way to admixture with other East Asian groups.

The South Maritime (Austronesian) expansion happened also independently from any genetic input, other than admixture with native Austronesians and Melanesians (West and East of Wallace Line respectively)

North and South Maritime populations did not interact, not even in the mainland, excepting Han expansion. This is quite curious and might explain the physiognomic differences between northern and southern "Mongoloids" much better than old hypothesis of admixture with Negritos/Melanesians.

Several distinct aboriginal groups seem to have receded before Austroasiatic and Austronesian expansion without significantly penetrating the settlers' genetic pools. These are mostly Negritos (Malay and Filipino) but also the Proto-Malay and the Mentawai. The Hmong, even if a larger ethnicity, can be argued to have suffered a similar destiny, absorbing a lot of Neolithic component and impacting other groups only minimally. In this they are similar to Austroasiatics, though in ISEA, the overlaying component is Austronesian, not Continental (and they have been absorbed linguistically).


Reconstructing the past

A plausible reconstruction of pre-Neolithic/Early Neolithic distribution could be this map:

Colors as above. I ignored the Mlabri and Proto-Malay. Please don't be too nit-picky with the necessary/convenient simplifications, thanks.

I understand that some quite reasonable ethnic interpretations can be made:

Sino-Tibetan: Neolithic Blue with various mixtures. Northern Neolithic may soon have evolved into a Blue-Yellow mix speaking proto-Sinitic.

Tai-Kadai (Kradai): Neolithic Blue with South Maritime Green (Southern Neolithic genesis).

Austroasiatic: Red, often with Neolithic Blue. Must have existed in Sundaland prior to Austronesian expansion and prior or simultaneous to Neolithic (Blue) expansion. Did they arrive there as Neolithic settlers, Epipaleolithic maybe, or were there "all the time"?

Austronesian: Bright Green, often in admixture. Must be original from the Taiwan-Philippines-SE China area.

Hmong-Mien: Along with their ethnic component (Light Blue), they display Neolithic Blue with minor Bright Green, suggesting that they adopted Neolithic before Tai-Kadai genesis. They have absorbed some Tai-Kadai blood but their genesis seems older than that (peripheral South Neolithic genesis).

Melanesian, Filipino Negrito and Malaysian Negrito are three different stocks.


Thursday, April 1, 2010

New paper on Chinese autosomal microsatellites


I'll comment more tomorrow maybe but here is the link anyhow.


Hongbin Lin et al., Genetic Relationships of Ethnic Minorities in Southwest China Revealed by Microsatellite Markers. PLoS ONE 2010. Open access.

Something that seems pretty obvious is that Guandong Chinese (Han) cluster apart from Northern Chinese (Han) and together with SE populations (Tai-Kadai and Mon-Khmer). However there is no absolute divide and depending on the method of analysis used populations cluster somewhat differently.

The most divergent groups are Tajiks (no surprise here) and the Drung a small Tibeto-Burman nation from Yunnan.



Update (Oct 17 2012): 


I never really got to review this paper properly but at the very least I must include here fig. 3:


... and fig. 4:


... which illustrate how Southern Han from Guangdong cluster with other SE Asians (specially Daic peoples) and not Northern Han (nor Tibeto-Burman, nor Mongol). Even more clearly than some other Southern ethnicities.

Monday, March 22, 2010

East Asian autosomal DNA (working note)


This is a highly simplified (approximate, tentative, very rough) geographical interpretation of
the HUGO consortium autosomal DNA clustering (paywall but someone hang it HERE - look at the details and not just this poor map, a mere working note, before assuming too many things, please), which produces five major components for East Asians and Melanesians at K=14. The rest are minority components (represented as circles) or South/West Asian ones (not shown here).

Continuous lines show the approximate areas with 50% or more of that component, dotted lines the areas (also approximate) with 20% or more.



Only three of the main components appear as majoritary in some populations: the yellow component, which reaches its highest frequency among Ryukyuans, the light green component which can be described as "Austronesian" but that is also important among Tai-Kadai and Southern Han and the dark green component, which is highest among Boungaville Melanesians and then in Eastern Lesser Sunda (Alor) and could be described as "Melanesian".

The red component is highest among Austroasiatic speakers, as well as Malaysian Malays, including Sea Dayaks (but not Proto-Malays nor Orang Asli), Javanese and Sundanese. The blue component is widespread among continental East Asian peoples (and Hymalayas) but shows no area nor ethnicity where is most concetrated.

Minor components (big dots) correspond to the Hmong-Mien (cyan, which share the blue and light green components too), the Mlabri (light purple, a tiny Austroasiatic hunter-gatherer group), Orang-Asli Negritos (dark red), Proto-Malay (purple-blue), Land Dayaks (grey here but white in the paper) and Filipino Negritos (dark purple).

Related posts:
- East Asians originated in SE Asia
- Indonesian Y-DNA is mostly Paleolithic
- Genetics of the Mlabri, Austroasiatic hunter-gatherers in Thailand

See also the supplementary materials: even more details!

Saturday, March 20, 2010

Genetics of the Mlabri, Austroasiatic hunter-gatherers in Thailand


Shuhua Xu et al., Genetic evidence supports linguistic affinity of Mlabri -- a hunter-gatherer group in Thailand. BMC Genetics, 2010. Open access.
Abstract (provisional)

Background

The Mlabri are a group of nomadic hunter-gatherers inhabiting the rural highlands of Thailand. Little is known about the origins of the Mlabri and linguistic evidence suggests that the present-day Mlabri language most likely arose from Tin, a Khmuic language in Austro-Asiatic language family. This study aims to examine whether the genetic affinity of the Mlabri is consistent with this linguistic relationship, and to further explore the origins of this enigmatic population.

Results

We conducted a genome-wide analysis of genetic variation using more than fifty thousand single nucleotide polymorphisms (SNPs) typed in thirteen population samples from Thailand, including the Mlabri, Htin and neighboring populations of the Northern Highlands, speaking Austro-Asiatic, Tai-Kadai and Hmong-Mien languages. The Mlabri population showed higher LD and lower haplotype diversity when compared with its neighboring populations. Both model-free and Bayesian model-based clustering analyses indicated a close genetic relationship between the Mlabri and the Htin, a group speaking a Tin language.

Conclusion

Our results strongly suggested that the Mlabri share more recent common ancestry with the Htin. We thus provided, to our knowledge, the first genetic evidence that supports the linguistic affinity of Mlabri, and this association between linguistic and genetic classifications could reflect the same past population processes.


While finding out about the Mlabri, an ethnicity I had never before heard of, and confirming that hunter-gatherers of Austroasiatic language do not just exist nowadays in Malaysia, is interesting in itself, what most has provoked my attention is the genetic comparison between various South East Asian ethnicities and the HapMap North Chinese and Japanese samples.

It seems that the paper has come right in time to illustrate the ongoing debate elsewhere on this blog, on whether there is any specific marker to Northern Han (or even Han in general) and other major East Asian populations such as Koreans and Japanese.

What I am realizing is that there is nothing really specific and that these northern populations almost invariably cluster with some of their southern neighbors, which are much more diverse. This really surprises me a bit because I would have expected that the Neolithic of the Yellow River, surely at the origin of the Chinese ethnicity (Han, Hui) would have been created by a somewhat distinct group that we should be able to identify by some genetic marker or set of markers, even if they spread southwards as the Chinese empire did.

Not at all. It seems that such specificity is almost invisible and that the Han (as well as genetically similar Korean and Japanese peoples) are very much unspecific in terms of genetic markers that could be easy to spot. They rather seem like melting pots of lineages which are traced to likely origins in the South of East Asia.

And, as I say this study comes very handy to illustrate this pattern within autosomal DNA as well, with Chinese and Japanese clustering with Hmong-Mien specially and then with Tai-Kadai peoples.

The neighbor-joining (NJ) tree:

The Htin and Mlabri are of course Austroasiatic, even if color-coded differently. JPT stands from Japanese from Tokyo and CHB as Chinese from Beijing (they are standard HapMap samples). The Karen are Tibeto-Burman speakers.

The maximum likelihood (ML) tree:

Careful here because the color-coding is not linguistic. As said before the Karen are Tibeto-Burman speakers and not Austroasiatic as would seem from their clustering and red color. CEU are Caucasoids of European ancestry from Utah and YRI are Yoruba from Nigeria (again standard HapMap samples for comparison). Cursive text is mine.

In this ML tree is maybe where the clustering is more apparent: the whole East Asia is first divided between the Mon and all the rest. They are neighbors of the Karen and linguistically most related to other Austroasiatics but they clearly cluster apart from all.

Then they branch into other Austroasiatics plus Karen (TB) and Tai-Kadai, Hmong-Mien and northern ethnicities. And this last group branches out into Tai (a branch of Tai-Kadai) and an amalgam of other Tai-Kadai (Yao), Hmong-Mien (Hmong) and the northern ethnicities. And only at this stage a North-South divide becomes apparent.

Wednesday, May 13, 2009

New paper on Mexican genetics


An interesting freely available study has just been published: Irma Silva-Solezzi et al. Analysis of genomic diversity in Mexican Mestizo populations to develop genomic medicine in Mexico. PNAS, 2009.

Results
We analyzed data from 300 nonrelated self-identified Mestizo individuals from 6 states located in geographically distant regions in Mexico: Sonora (SON) and Zacatecas (ZAC) in the north, Guanajuato (GUA) in the center, Guerrero (GUE) in the center–Pacific, Veracruz (VER) in the center–Gulf, and Yucatan (YUC)in the southeast. Considering that Zapotecos have been shown as a good ancestral population for predicting Amerindian (AMI) ancestry in Mexican Mestizos (16), we included 30 Zapotecos(ZAP) from the southwestern state of Oaxaca (Fig. 1). For comparative purposes, we included similar data sets from HapMap populations: northern Europeans (CEU), Africans (YRI), and East Asians (EA), including Chinese (CHB) and Japanese (JPT). A HapMap-like database with SNP frequencies in Mexicans and HapMap populations was generated (http://diversity.inmegen. gob.mx).

Some assorted images:

Fig.1. Map showing areas sampled and diversity compared to HapMap standard samples. All populations are Mestizos except ZAP who are Native American Zapotecos.

Fig. 2. Plots of the Mexican and HapMap samples: A includes YRI (Yorubans), B without YRI.


Fig. 3A. Poulation structure of Mexicans and HapMap samples.

No novel conclusions, I'd say. As is known (at least by Mexicans) Northwesterners (represented here by SON) are "whiter" than average. All Mestizo Mexicans are basically a mixture of Europeans and Native Americans with very small apportion of African ancestry (more notable in Guerrero and some Veracruz individuals).

It is interesting to see anyhow how when a true Amerindian sample is introduced, a clear distinction appears with the East Asian HapMap sample that, by itself, can only be a poor approximation to Native American genotype. To a lesser extent surely it could be argued that CEU may not represent well the SW European ancestry of Mexicans but guess the differences are much less extreme (Europeans in general cluster all very close to CEU when in intercontinental contexts).

For further references on medical genomics regarding Mexico check this article at Science Daily.

.

Friday, June 27, 2008

Larger samples better than larger databases in autosomal genetics


June 2008 PLOS Genetics is out and the article that caught my attention this time is:
Calibrating the Performance of SNP Arrays for Whole-Genome Association Studies by Ke Hao et al.

... according to our results a study employing N = 300 subjects and the Affy500K platform offers higher power than a study employs N = 250 subjects and the Ilmn650K platform. This 20% increase in sample size (N = 300 vs. N = 250) provides more power than the 90% increase in the number of SNPs genotyped (286K SNPs vs. 545K SNPs). In scenarios where funding becomes the constraining factor, our results suggest that genotyping larger sample size with cheaper SNP arrays might achieve better statistical power. On the other hand, if the constraining factor is the number of subjects, then it appears that SNP arrays offering the largest genetic coverage should be employed.


Saturday, February 9, 2008

Biased European genetics (again)

I am pretty sure that it's not really intentional but each time they study autosomal European genetic variation, they forget to study the central strip (French, Austrians, Hungarians, etc.). They seem to like going to the extremes and somehow "demonstrate" a fallacious discontinuity between Northern and Southern Europe.

This is the case of the new research by a US team led by Chao Tian. The overall results are most visible in figure 2C (excluding Askhenazis) that show both a distribution along the East-West geographical axis and along the North-South axis (maybe more apparent for lack of intermediate populations). Smaller components (figure 6) also seem to emphasize the E-W dominant cline, though they are dismissed by the authors.

This last is surely wrong. Not much older studies of the same kind evidenced (Bauchet et al, 2007) that when taken only two components the results are actually much distorted. Often smaller components in the overall, are very important and even dominant in one specific population. These locally dominant components become invisible when only the two or three more extended are considered, making large geographically defined populations to be classified by a minor component of their genetic pool.

European five main clusters

The above image is a five-pole diagram I drew some time ago based in the K=5 graph below from that other study (the study reached to K=6 but the 6th component was too diffuse to matter, maybe it is a Balcanic or Eastern European element, as these areas were not studied).

Considered only two components (K=2, not shown but corresponding to the red and blue ones), Spanish samples, for instance, fell almost completely in the red "Near Eastern" zone, while Basques resulted extremely ambiguous (due to near lack of either of these two components).

Instead, seen as a plot of five components, Spanish and Basques clearly cluster primarily with themselves and no one else. Some Spanish are somewhat intermediate with Eastern Mediterraneans while others are intermediate with Basques but mostly they cluster on their own.

Another find of the K=5 plot is a Central-Northern European cluster (green) distinct of the "Finnic" blue marker. Also it's noticeable that many Northern Europeans show tendencies towards not just Finns but also Basques or even Southern Europeans in some cases. Again the lack of representation of the intermediate strip (France is only represented by one sample, while the Danubian basin, the Balcans and Eastern Europe are totally absent) creates some distortion, enhancing N/S differences.

By the way, how do I read these clusters? In my opinion the Iberian (cyan), Basque (orange) and Central-North (green) clusters must represent late Paleolithic Magdalenian and/or Epipaleolithic populations: those of the Iberian, Franco-Cantabrian and Rhin-Danub regions respectively. The two principal components instead would represent two later arrivals: Neolithic for the red (Eastern Mediterranean) one and Uralic (Fino-Ugric) for the blue one - though this last one poses some difficulties of interpretation actually (it is very possible that this "Uralic" element has been distributed by Indo-European migrations as well, specially those linked to Scandinavia and the Baltic region, like Germanic peoples).