| 李竟泽,贾克斌.基于纹理结构与CA-Transformer的深度图帧内快速编码算法[J].高技术通讯(中文),2026,36(7):713~723 |
| 基于纹理结构与CA-Transformer的深度图帧内快速编码算法 |
| A fast intra-frame coding algorithm for depth maps based on texture structure and CA-Transformer |
| |
| DOI:10. 3772 / j. issn. 1002 - 0470. 2026. 07. 005 |
| 中文关键词: 三维高效视频编码; 纹理图; 深度图; 编码树单元划分; 帧内编码 |
| 英文关键词: three-dimensional high efficiency video coding, texture map, depth map, coding tree unit partitioning, intra encoding |
| 基金项目: |
| 作者 | 单位 | | 李竟泽 | (北京工业大学信息科学技术学院北京 100124)
(先进信息网络北京实验室北京 100124) | | 贾克斌 | |
|
| 摘要点击次数: 64 |
| 全文下载次数: 44 |
| 中文摘要: |
| 对三维高效视频编码(three-dimensional high efficiency video coding,3D-HEVC)在深度图编码中复杂度过高的问题,本文提出了一种基于纹理划分结构与深度学习并行的算法实现深度图的快速编码算法。首先,提出了基于纹理图划分结构加速深度图编码的方法,提前识别并简化编码树单元(coding tree unit,CTU),得到快速划分结果。其次,针对无法使用纹理图划分结构快速划分的CTU,设计了一种名为CA-Transformer的神经网络。应用卷积神经网络对局部特征的敏感性与视觉Transformer对全局特征的敏感性,将结构复杂的深度图图像输入到本文的CA-Transformer网络中,得到对深度图CTU划分结构的预测。使用本文算法替换原始编码标准的CTU划分过程,实验结果表明,本文算法在保证率失真性能与合成视点质量几乎不受影响的前提下,平均减少了56.00%的编码时间。 |
| 英文摘要: |
| To address the high complexity issue in depth map coding of three-dimensional high efficiency video coding(3D-HEVC), this paper proposes a fast depth map coding algorithm that integrates texture partitioning structure and deep learning. First, a texture map partitioning structure based method is introduced to accelerate depth map coding, which identifies and simplifies coding tree unit (CTU) in advance to achieve rapid partitioning results. Second, a neural network named CA-Transformer is designed. Leveraging the sensitivity of convolutional neural networks to local features and the global feature modeling capability of vision Transformers, the proposed CA-Transformer network processes structurally complex depth map images to predict CTU partitioning structures. By replacing the original CTU partitioning process in the standard encoder with the proposed algorithm, experimental results demonstrate that the algorithm reduces encoding time by an average of 56.00% while maintaining nearly identical rate-distortion performance and synthesized view quality. |
|
查看全文
查看/发表评论 下载PDF阅读器 |
| 关闭 |