1,721,003 research outputs found
A low power 1024-channels spike detector using latch-based ram for real-time brain silicon interfaces
High-density microelectrode arrays allow the neuroscientist to study a wider neurons population, however, this causes an increase of communication bandwidth. Given the limited resources available for an implantable silicon interface, an on-fly data reduction is mandatory to stay within the power/area constraints. This can be accomplished by implementing a spike detector aiming at sending only the useful information about spikes. We show that the novel non-linear energy operator called ASO in combination with a simple but robust noise estimate, achieves a good trade-off between performance and consumption. The features of the investigated technique make it a good candidate for implantable BMIs. Our proposal is tested both on synthetic and real datasets providing a good sensibility at low SNR. We also provide a 1024-channels VLSI implementation using a Random-Access Memory composed by latches to reduce as much as possible the power consumptions. The final architecture occupies an area of 2.3 mm2, dissipating 3.6 μW per channels. The comparison with the state of art shows that our proposal finds a place among other methods presented in literature, certifying its suitability for BMIs
Low-Power Energy-Based Spike Detector ASIC for Implantable Multichannel BMIs
Advances in microtechnology have enabled an exponential increase in the number of neurons that can be simultaneously recorded. To meet high-channel count and implantability demands, emerging applications require new methods for local real-time processing to reduce the data to transmit. Nonlinear energy operators are widely used to distinguish neural spikes from background noise featuring a good tradeoff between hardware resources and accuracy. However, they require an additional smoothing filter, which affects both area occupation and power dissipation. In this paper, we investigate a spike detector, based on a series of two nonlinear energy operators, and a simple and adaptive threshold, based on a three-point median operator. We show that our proposal provides good accuracy compared to other energy-based detectors on a synthetic dataset at different noise levels. Based on the proposed technique, a 1024-channel neural signal processor was designed in a 28 nm TSMC CMOS process by using latch-based static random-access memory (SRAM), demonstrating a total power consumption of 1.4 μW/ch and a silicon area occupation of 230 μm2/ch. These features, together with a comparison with the state of the art, demonstrate that our proposal constitutes an alternative for the development of next-generation multichannel neural interfaces
CFPM: Run-time Configurable Floating-Point Multiplier
Approximate computing is a new approach that can help to reduce power consumption in error-resilient applications. Although many works have been proposed for fixed-point multipliers with predetermined levels of accuracy, they are not able to adapt to a wide range of applications, that need floating-point calculations with time-varying requirements. In this paper, we introduce an adjustable floating-point multiplier in which groups of partial products can be dynamically truncated, while the approximation error is reduced with the help of a simple rounding technique. In the proposed floating-point multiplier, precision and power can be adjusted at run-time based on the users' requirements. The developed circuits are synthesized in TSMC 28 nm CMOS technology. The comparison with the state-of-the-art shows a good trade-off between error and power consumption. Furthermore, we demonstrate the suitability and versatility of our multiplier through image processing applications, proving that it can be usefully employed in real-world scenarios
Novel Low-Power Floating-Point Divider With Linear Approximation and Minimum Mean Relative Error
Floating-point division involves the computation of the ratio (1 Mx)/(1 My), where Mx and My represents the mantissas of the input values. In this paper, we propose a new method for approximating this operation using a linear function of Mx, with coefficients that depend on My. The coefficients are calculated to minimize the Mean Relative Error Distance (MRED) of the approximation. To this end, the range of My is partitioned in N sub-intervals where the minimization of MRED is formulated as a linear programming problem, whose solution gives optimal coefficient values. The hardware implementation requires a small lookup table, two multipliers and an adder. An aggressive coefficients quantization is exploited to further optimize the design. Obtained MRED improves by increasing , ranging from 1.4% to 0.33%. Implementation results in a 28nm CMOS technology show that the proposed design outperforms the state-of-the-art, offering the best trade-off between hardware complexity and accuracy. Results for two image processing applications, change detection and JPEG compression, demonstrate remarkable performance, with SSIM very close to 1 and PSNR values exceeding 50dB
Low Power Spike Detector for Brain-Silicon Interface using Differential Amplitude Slope Operator
High-density implantable microelectrode arrays allow to study in-vivo a wide neuron population with the help thousands of integrated electrodes. Such a high electrodes count generates a large amount of data which poses severe challenges in the design of long-term implantable silicon interfaces that rely on a limited power budget and a narrow transmission bandwidth. In such a scenario, reliable spike detectors are needed, as they allow to transmit only the relevant neural information (the neuron action potential) instead of the whole raw recording. Spike detectors based on energy operators provide a good compromise between detection performance and hardware complexity. However, they require a suitable smoothing filter that affects both area occupation and power dissipation. In this paper, we propose a spike detector based on the cascade of two energy operators, without smoothing, and the use of a new simple adaptative threshold calculation. We show that this technique provides good detection metrics compared to previous approaches, for different SNR levels and with several noise models. The proposed system has been synthesized in TSMC 28 nm CMOS technology showing a per-channel area occupation of 0.0021 m m2 with a power consumption of 0.15 μ W, comparing favorably with the state of art of brain machine silicon interfaces
Low-power approximate multiplier with error recovery using a new approximate 4-2 compressor
In this paper we propose an energy-efficient approximate multiplier which uses a new approximate 4-2 compressor. The proposed compressor has a low error probability and its error conditions can be easily detected. This, as previously shown in the literature, makes it possible to implement error recovery, when the compressor is used in the partial product reduction phase of a multiplier. Simulation results show that proposed approximate multipliers exhibit a sensible reduction in Mean Error Distance and in maximum Error Distance, compared to previous art. Application to an image processing task shows an improvement of about 8dB in peak signal-to-noise ratio. Implementation results in 28nm CMOS show that the electrical performance of multipliers designed with the novel circuit are close to the one obtained with previously proposed approximate compressors
Approximate Recursive Multipliers Using Carry Truncation and Error Compensation
Approximate computing is a fast-emerging paradigm promising higher circuit performances in error tolerant applications. Binary multipliers are a common target for approximate computing due to their complexity and the multitude of their applications. In this paper, we investigate approximate recursive multipliers based on novel 4x4 multiplier blocks. We present three approximate 4x4 multipliers, with different error-precision trade-off, obtained by carry truncation and error compensation. These basic blocks are exploited to design 8x8 approximate multipliers. The proposed circuits are implemented in a 14 nm FinFET technology and show improved performance compared to the state-of-the-art
Quality-Scalable Approximate LMS Filter
Approximate Computing allows improving circuits performances by accepting inaccuracies in the calculations. Adaptive Least Mean Squares (LMS) filters can benefit from Approximate Computing, since they are inherently inexact and power-hungry. In this paper we propose a Quality-Scalable approximate LMS filter, in which it is possible to change the approximation level at runtime, by acting on an external quality knob. This allows to adapt the filter performance according to the target application and the input data. The proposed approach introduces approximations at algorithmic level. The filter is able to enter automatically in a low-power approximated-mode, by freezing the update of some of its coefficients. A simple auxiliary circuit monitors the error and the quality knob and drives the filter in the approximated-mode of operation when convergence has been reached. The proposed approximate LMS filter is implemented in TSMC 40nm CMOS technology. The implementation results show a power improvement in the range 5% to 32%, as a function of the desired quality loss
Real-Time Downsampling in Digital Storage Oscilloscopes with Multichannel Architectures
Digital Storage Oscilloscopes (DSOs) conjugate high performance with large number of features and flexibility. The basic structure, based on fast Analog to Digital Converter (ADC) and memory, is augmented with several components for channel matching and bandwidth improvement, and processors that provide visualization, frequency processing, jitter and stability measurement, etc. Unfortunately, fine resolution in sample rate selection is not available, such that for several applications the user must run complex measurement procedures that require data download and offline processing. The paper proposes a dedicated digital circuit that offers fine control of the time-base by real time downsampling the input stream at an almost arbitrary sampling rate. The proposed circuit implements a series of operations involving real-time data filtering, defragmentation and packing, which are not considered by alternative approaches, like those based on polyphase filters, that offer very limited choices for the sampling rate. The circuit is designed to work in conjunction with the highest performance DSOs that use a multichannel architecture. Design rules for circuit design are provided together with implementation results in 14nm FinFET technology. When designed for an architecture with 64 channels with 8bit input samples the circuit works in real time with a sampling rate of 220 GSps running at 3.42 GHz, with a silicon footprint of 0.17 mm2 and a power dissipation of 0.85W
Design of Generalized Enhanced Static Segment Multiplier with Minimum Mean Square Error for Uniform and Nonuniform Input Distributions
In this paper, we analyze the performances of an Enhanced Static Segment Multiplier (ESSM) when the inputs have both uniform and non-uniform distribution. The enhanced segmentation divides the multiplicands into a lower, a middle, and an upper segment. While the middle segment is placed at the center of the inputs in other implementations, we seek the optimal position able to minimize the approximation error. To this aim, two design parameters are exploited: m, defining the size and the accuracy of the multiplier, and q, defining the position of the middle segment for further accuracy tuning. A hardware implementation is proposed for our generalized ESSM (gESSM), and an analytical model is described, able to find m and q which minimize the mean square approximation error. With uniform inputs, the error slightly improves by increasing q, whereas a large error decrease is observed by properly choosing q when the inputs are half-normal (with a NoEB up to 18.5 bits for a 16-bit multiplier). Implementation results in 28 nm CMOS technology are also satisfactory, with area and power reductions up to 71% and 83%. We report image and audio processing applications, showing that gESSM is a suitable candidate in applications with non-uniform inputs
- …
