腾讯混元Hy ASR 3.0 preview:让语音识别理解上下文Tencent Hunyuan Hy ASR 3.0 Preview: Speech Recognition That Understands Context
腾讯混元Hy ASR 3.0 preview:让语音识别理解上下文 元宝已接入 8 月 4 日,腾讯混元正式发布新一代语音识别模型 Hy ASR 3.0 preview。基于最新一代大语言模型 Hy3 的语言理解能力,Hy ASR 3.0On August 4, Tencent Hunyuan released Hy ASR 3.0 preview, a speech-recognition model that understands context, now integrated into Yuanbao.
📌 核心要点
- 8 月 4 日,腾讯混元正式发布新一代语音识别模型 Hy ASR 3.0 preview中文8 月 4 日,腾讯混元正式发布新一代语音识别模型 Hy ASR 3.0 preview
- Hy ASR 3.0 preview 在多个开源评测集和自建评测集上整体表现领先中文Hy ASR 3.0 preview 在多个开源评测集和自建评测集上整体表现领先
- 目前Hy ASR 3.0 preview已上线腾讯云官网对外提供 API 服务,可广泛应用于智能客服、内容理解、语音搜索等场景中文目前Hy ASR 3.0 preview已上线腾讯云官网对外提供 API 服务,可广泛应用于智能客服、内容理解、语音搜索等场景
元宝已接入
8 月 4 日,腾讯混元正式发布新一代语音识别模型 Hy ASR 3.0 preview。基于最新一代大语言模型 Hy3 的语言理解能力,Hy ASR 3.0 preview 融合高精度语音识别与深度语义理解能力,在通用识别、上下文感知、多场景鲁棒性以及方言覆盖等核心维度实现全面提升,能够在更复杂的真实输入中给出准确、连贯且更接近用户意图的转写结果,从“逐字转写、单点优化”演进为“理解语境、兼容场景、一键直出”。
Hy ASR 3.0 preview 在多个开源评测集和自建评测集上整体表现领先。在开源评测集中,Hy ASR 3.0 preview 在多语种的WER(Word Error Rate, 词错误率)上都控制在3%左右,其中中文普通话 WER 3.34%、英语 WER 2.62%、粤语 WER 3.12%,整体领先竞品。
在自建评测集上,Hy ASR 3.0 preview在通用识别、方言识别、上下文Context理解、复杂声学场景(高噪、耳语)等场景化评测中,WER也保持较低水平,整体领先竞品。
从用户使用的角度,Hy ASR 3.0 preview重点提升了四类能力,包括:
- 通用识别更准确:提升通用语音、方言、中英混说等场景下的识别准确性,进一步减少错字、漏字问题。
- 更能理解用户意图:结合上下文Context精准捕捉用户语境,智能纠错同音词,消除语义歧义。
- 更容易适配专业场景:支持热词注入增强,帮助模型快速识别品牌产品名称、人名及行业术语,降低业务接入和持续维护成本。
- 复杂环境下更稳定:面向高噪、耳语、悄悄话等多种声学场景下进行专项优化,复杂条件下依然保持稳定表现。
架构、数据、后训练持续增强,提升语音识别能力上限
Hy ASR 3.0 preview的能力提升并非来自单一模块,而是模型架构、数据 Scaling 与后训练优化共同作用的结果。
在架构层面,Hy ASR 3.0 preview 采用兼顾效率与性能的 MoE 架构,并将基座模型升级至 Hy3,进一步增强语言理解、上下文建模和语义推理能力。在语音侧,团队自研无监督语音 Encoder,通过数千万小时级无监督语音数据训练,使其能够从复杂音频中提取高质量的声学表征。
为了持续提升模型的语音建模与理解能力,混元团队对语音 Encoder 和大语言模型进行联合训练,引入数千万小时级、多来源的语音数据,覆盖多种方言、口音和声学环境,并通过高质量数据管线进行精细化标注。在大规模联合预训练的基础上,Hy ASR 3.0 preview 通过多阶段能力注入,逐步获得上下文感知、复杂场景适应和方言识别等能力。
围绕通用识别、上下文理解和复杂场景鲁棒性,团队构建了高质量的 SFT recipe,覆盖上下文context、专业名词、不同声学环境以及多样人群语音。针对方言识别能力,SFT 数据体系进一步覆盖 10 大方言片区和 20 余个二级小片区。
在此基础上,团队进一步引入多阶段强化学习,分别针对通用转写准确性、Any-context 上下文能力和复杂长尾场景进行优化,降低模型在复杂声学环境中的误识别和漏识别问题。
目前Hy ASR 3.0 preview已上线腾讯云官网对外提供 API 服务,可广泛应用于智能客服、内容理解、语音搜索等场景。
元宝深度参与共研并已完成首发上线,用户打开元宝按住说话即可体验方言识别、上下文智能纠错与复杂环境稳定转写等能力提升,语音输入更准、更稳,且免费开放;WorkBuddy 等产品也在陆续接入中。
体验地址:
腾讯混元:
https://aistudio.tencent.com/visual
腾讯云:
https://cloud.tencent.com/document/product/1093/135476
- WorkBuddy重大升级,AI时代的Office来了2026-07-30
- 国产算力正在进入Token标准化时代2026-06-18
- 为什么最有价值的AI讨论总发生在知乎?2026-06-17
- 一个模型控制手脚腰身!机器人终于学会全身协同干精细活了2026-06-16
On August 4, Tencent Hunyuan officially released its new-generation speech recognition model, Hy ASR 3.0 preview. Built on the language-understanding capabilities of the latest large model, Hy3, Hy ASR 3.0 preview combines high-precision speech recognition with deep semantic understanding, delivering across-the-board improvements in general recognition, context awareness, multi-scenario robustness, and dialect coverage—producing transcripts that are more accurate, coherent, and closer to user intent in complex real-world inputs. It evolves from 'verbatim transcription with isolated optimizations' to 'understanding context, adapting to scenarios, and one-click output.'
On open benchmark sets, Hy ASR 3.0 preview leads overall. Its multilingual word error rate (WER) is around 3%: Mandarin Chinese 3.34%, English 2.62%, and Cantonese 3.12%—ahead of competitors. On Tencent's internal benchmarks covering general recognition, dialect recognition, context understanding, and complex acoustic scenes (high noise, whispering), WER also stays low and leads rivals.
From a user perspective, Hy ASR 3.0 preview focuses on four capability upgrades: (1) more accurate general recognition, reducing misrecognized and dropped characters across general speech, dialects, and Chinese-English mixing; (2) better intent understanding, using context to capture user meaning, intelligently correct homophones, and resolve ambiguity; (3) easier adaptation to professional scenarios via hot-word injection for brand names, people, and industry terms; and (4) greater stability in complex environments such as high noise, whispering, and soft speech.
The improvements come not from a single module but from joint gains in model architecture, data scaling, and post-training. Architecturally, it adopts an efficient MoE design and upgrades the base model to Hy3; on the audio side, a self-developed unsupervised speech encoder trained on tens of millions of hours of data extracts high-quality acoustic representations. Joint training of the encoder and LLM over multi-source speech data, plus multi-stage SFT covering 10 major dialect regions and 20+ sub-regions, and multi-stage RL, further cut errors in hard long-tail scenarios.
Hy ASR 3.0 preview is now available as an API on the Tencent Cloud site, usable in smart customer service, content understanding, voice search, and more. Yuanbao (Tencent's assistant) co-developed and launched it first—users can experience dialect recognition, context-aware correction, and stable transcription in noisy environments for free by long-pressing to talk; WorkBuddy and other products are being integrated.