文章摘要
王凯歌,王炜凡,胡振涛,金勇.基于数据湖的非结构化医疗数据高效检索方法[J].高技术通讯(中文),2026,36(7):724~730
基于数据湖的非结构化医疗数据高效检索方法
Efficient retrieval method for unstructured medical data based on data lake
  
DOI:10. 3772 / j. issn. 1002 - 0470. 2026. 07. 006
中文关键词: 数据湖; 非结构化医疗数据; 高效检索; 高维索引
英文关键词: data lake, unstructured medical data, efficient retrieval, high-dimensional index
基金项目:
作者单位
王凯歌 (河南大学人工智能学院郑州 450046) 
王炜凡  
胡振涛  
金勇  
摘要点击次数: 53
全文下载次数: 40
中文摘要:
      随着智能医疗设备广泛应用,非结构化医疗数据呈现爆炸性增长,其高效检索与管理对医疗领域中数据挖掘、数据探索和人工智能应用具有重要意义,可辅助医生优化治疗决策。然而,现有数据检索方法注重数据实体语义表达能力,缺乏对检索算法优化,导致大规模数据检索效率低下。高维索引和分布式存储平台是提高检索效率的有效途径,但现存研究没有将二者有效结合,在检索效率和存储灵活性方面仍存在不足。针对上述挑战,本文提出一种基于数据湖的非结构化医疗数据高效检索方法EMUD(efficient retrieval method for unstructured medical data based on data lake),它使用新颖高维索引,结合分布式计算能力,显著提升非结构化医疗数据检索效率。同时,EMUD通过集成数据湖灵活数据存储特性,实现非结构化医疗数据高效管理,为医疗前沿应用提供灵活数据支撑接口。实验结果表明,在使用较小存储空间条件下,EMUD在检索精度、检索效率以及索引创建时间方面均优于其他高维索引。
英文摘要:
      With the widespread adoption of intelligent medical devices, unstructured medical data has grown explosively. Efficient retrieval and management of such data are essential for data mining, exploratory analysis, and artificial intelligence applications in the medical domain, thereby assisting clinicians in optimizing treatment decisions. However, existing data retrieval methods primarily focus on the semantic representation of data entities and lack the optimization of retrieval algorithms, resulting in low efficiency of large-scale data retrieval. High-dimensional indexes and distributed storage platforms are effective approaches to improving retrieval efficiency, but prior research has not achieved their effective integration, resulting in shortcomings in both retrieval performance and storage flexibility. To address the above challenges, this paper proposes EMUD(efficient retrieval method for unstructured medical data based on data lake), an efficient retrieval method for unstructured medical data based on data lake. It leverages a novel high-dimensional index combined with distributed computing capabilities to enhance the retrieval efficiency of unstructured medical data significantly. Meanwhile, EMUD integrates the flexible data storage characteristics of the data lake to achieve efficient management of unstructured medical data and provide a scalable data support interface for cutting-edge medical applications. The experimental results demonstrate that EMUD outperforms other high-dimensional indexes regarding retrieval accuracy, retrieval efficiency, and index construction time under the condition of using less storage space.
查看全文   查看/发表评论  下载PDF阅读器
关闭