⚡ 本页包含 AI 生成的分析内容,仅供参考
本文提出了一款47nW的混合信号语音活动检测器(VAD),采用非易失电容ROM、短时CNN特征提取器和RNN分类器,解决了始终开启的VAD在超低功耗下的特征提取与分类问题,实现了高能效的语音激活检测。
Jinhai Lin1, Ka-Fai Un1, Wei-Han Yu1, Pui-In Mak1, Rui P. Martins1,2 University of Macau, Macau, China Instituto Superior Tecnico/University of Lisboa, Lisbon, Portugal 1 2 Real-time speech recognizers and translators rely on an always-on voice activity detector (VAD) to enable and disable the main system for effective power savings. A feature extractor and a memoryless classifier build the basic structure of the recent VADs, as depicted in Fig. 13.2.1 (upper-left). The feature extractor [1], [2] using Mel-frequency cepstral coefficients (MFCCs) regrettably occupies a substantial area and incurs a long latency for an extraction window of 16 to 25ms to cover a frequency down to ~100Hz. For example, the required filter bank in [1] consumes 1µW with a 25ms extraction window, leading to a 30ms latency. The mixer-based analog filter in [2] succeeds in squeezing the feature extractor’s power to 60nW by using a time-interleaved fashion,
a Non-Volatile Capacitor-ROM, a Short-Time CNN Feature, Extractor and an RNN Classifier