1)河北医科大学法医学院,河北省法医学重点实验室,河北省法医分子鉴定协同创新中心,石家庄 050017;2.3)公安部鉴定中心,法医遗传学公安部重点实验室,北京市现场物证检验工程技术研究中心,北京 100038;3.2)中国人民公安大学侦查学院,北京 100038;4)江苏师范大学生命科学学院,江苏省系统发育与比较基因组学重点实验室,徐州 221000
1)College of Forensic Medicine, Hebei Key Laboratory of Forensic Medicine, Collaborative Innovation Center of Forensic Medical Molecular Identification, Hebei Medical University, Shijiazhuang 050017, China;2.3)Key Laboratory of Forensic Genetics, Beijing Engineering Research Center of Crime Scene Evidence Examination, Institute of Forensic science, Beijing 100038, China;3.2)Institute of Criminal Investigation, People’s Public Security University of China, Beijing 100038, China;4)Jiangsu Key Laboratory of Phylogenomics and Comparative Genomics, School of Life Sciences, Jiangsu Normal University, Xuzhou 221000, China
This work was supported by grants from the Fundamental Research Fund for Institute of Forensic Science (2024JB026, 2024JB043), The National Natural Science Foundation of China (82171870), and the National Key Research and Development Program of China (2022YFC3341004).
目的 研究不同量级单核苷酸多态性(single-nucleotide polymorphism,SNP)位点组合,基于筛选SNP位点集进一步提高远亲缘关系的预测能力。方法 首先选取3种芯片中国基因分型芯片(Chinese genotyping array,CGA,Illumina)、全球筛查芯片(global screening array,GSA,Illumina)、23魔方V2版高密度SNP芯片(23MF_V2 high-density SNP array,Affy,Thermo Fisher Scientific (formerly Affymetrix))位点进行合并、质控,筛选得到一组高密度SNP位点集(1 180 k);从161份全基因组测序数据中获取1 180 k位点集,使用共祖片段(identity-by-descent,IBD)算法进行亲缘关系推断,通过IBD片段长度和预测准确性的变化趋势评估该位点集的亲缘关系推断能力。结果 经过筛选后,得到1 184 334个常染色体SNP位点集(下文称高密度SNP位点集)。与3种芯片位点集的平均结果相比较,高密度SNP位点集增加了总IBD片段长度以及1~9级的平均IBD片段长度;8级置信区间准确率为70.97%,提高了3.50%;1~8级平均置信区间准确率为91.39%,提高了1.00%;8、9级假阴性率分别降低2.42%、6.76%。高密度SNP位点集的1~8级亲缘关系推断系统效能达98.91%。通过随机减少位点结果发现,增加SNP位点数量能提升较远亲缘关系推断能力。结论 高密度SNP位点集显著增强远缘关系推断效能,可精准覆盖1~8级亲缘关系,且1~8级平均置信区间准确率稳定在90%以上。本研究发现,SNP位点数量增加,可以提高远亲缘预测能力。
Objective This study aims to explore the potential of different orders of magnitude single-nucleotide polymorphism (SNP) locus combinations for predicting distant kinship relationships. A high-density SNP locus set was constructed, and a comprehensive assessment of its inference capability was conducted.Methods Firstly, we selected three commercial chip panels, CGA (Chinese genotyping array, Illumina), GSA (Global screening array, Illumina), Affy (23MF_V2 high-density SNP array, Affymetrix) and merged them after quality control, forming a high-density SNP locus panel(1 180 k). Secondly, we selected 161 samples and collected their peripheral blood samples by using whole-genome sequencing technology. Within this sample population, the levels of kinship relationships fully covered the range from level 1 to level 9, and the number of kinship pairs at each level was consistently maintained at over 50 pairs. From 161 samples data of whole-genome sequencing, the 1 180 k locus set was extracted, which is referred to as the high-density SNP locus set in the following text. The kinship inference was conducted using the identity-by-descent (IBD) algorithm with the selected optimal parameters. To comprehensively evaluate the performance of the high-density SNP locus set in kinship inference, we compared it with the three commercial chip panels, the intersection of these three chip loci, and the control sets constructed by randomly reducing the number of the high-density SNP locus set. Based on the changes in the IBD lengths, as well as the dynamic trends in prediction accuracy, we conducted a scientific assessment of the kinship inference capability of the high-density SNP locus set.Results After screening, a set of 1 184 334 autosomal SNPs was obtained. During the process of screening the optimal IBD length threshold, the result revealed that 0 cM, 1 cM, and 2 cM all demonstrated good applicability. However, to avoid the issue of a large amount of redundant information caused by setting a too low IBD length threshold, this study ultimately selected 2 cM as the optimal threshold. Compared with the average results of three chip panels, the high-density SNP locus set increased the total IBD length and the average IBD length across levels 1-9; the accuracy of the confidence interval for level 8 was 70.97%, which represented a 3.50% improvement; the average confidence interval accuracy for levels 1-8 was 91.39%, representing a 1.00% increase; and the false negative rates at levels 8 and 9 were reduced by 2.42% and 6.76%, respectively. The system efficacy of the high-density SNP locus set for kinship inference of first to eighth degree relationships reached 98.91%. Through random reduction of the high-density SNP locus set results, it is found that increasing the number of SNPs with the panel, the detection efficiency of IBD length showed a significant upward trend. At the same time, the overall trend in the accuracy of kinship relationship prediction as well as the confidence interval accuracy also indicated that both metrics steadily increased with the addition of more loci.Conclusion The results show that the high-density SNPs panel significantly enhances the efficacy of distant kinship inference, accurately covering kinship degrees, with the average confidence interval accuracy for first to eighth degree relationships stably above 90%. The study finds that increasing the number of SNPs panel can improve the ability to predict distant kinship.
李晶,孙一杰,赵雯婷,汤子琛,刘京,李彩霞.高密度单核苷酸多态性的系谱推断效能研究[J].生物化学与生物物理进展,2026,53(3):740-753 LI Jing, SUN Yi-Jie, ZHAO Wen-Ting, TANG Zi-Chen, LIU Jing, LI Cai-Xia. Research on The Genealogical Inference Efficiency of High-density SNPs[J]. Progress in Biochemistry and Biophysics,2026,53(3):740-753
复制

扫码关注 生物化学与生物物理进展 ® 2026 网站版权 ICP:京ICP备05023138号-1 京公网安备 11010502031771号
