公共语音AI基准测试表明达到了人类水平的性能,但可能针对测试进行了优化。
语音AI基准测试显示模型的表现达到人类水平,但这些分数可能是针对测试本身进行了优化,而不是反映实际应用中的表现。
语音AI基准测试显示模型的表现达到人类水平,但这些分数可能是针对测试本身进行了优化,而不是反映实际应用中的表现。
Reference disagreement (VoxPopuli case study) Masked Entity Retrieval Orthographic Switching Localizing the switches Conclusion Public voice AI benchmarks increasingly suggest that models are performing at human levels.
Reference disagreement (VoxPopuli case study) Masked Entity Retrieval Orthographic Switching Localizing the switches Conclusion Public voice AI benchmarks increasingly suggest that models are performing at human levels. Yet those scores don't always reflect how models work in th…
To test whether these behaviors generalize beyond the public benchmarks, we also collected fresh data from the same source domains but after the models' training cutoffs: recent European Parliament recordings for VoxPopuli and recordings from newly active LibriVox narrators for…