Always-on, voice-activated tinyML systems, like those implementing keyword spotting (KWS), demand low power consumption and a small footprint. In certain instances, subV energy-harvesting sources restrict the available supply voltage to below 0.5V [1]. Most KWS designs focus on optimizing the audio feature extraction (FEx) unit, which dominates the overall power and area. Analog FEx using multi-channel Gm-C bandpass filters (BPFs) and analog rectifiers [2], [3] can be as much as 10× more power efficient than digital FEx for a comparable silicon area [4]. However, analog FEx circuits have not demonstrated KWS with more than four keywords. They also suffer from a large footprint, challenging technology migration and limited dynamic range (DR) at low supply voltage, while speech signals have inherently a high DR. These limitations ultimately lead to the use of time domain (TD) [5], [6], or partial TD [7] alternatives. In [5], a 0.5V solarpowered TD-FEx employs a voltage-to-time converter (VT
Ali Mostafa, Emmanuel Hardy, Franck Badets
CEA-Léti, Grenoble, France