EN
Home > Technology & support > Technical article >

【Application Solution】In the Noise, Hear Clarity: Confronting the Pain Points and Engineering Challenges of Noise Reduction Scenarios, awinic's Self-developed Audio Solutions Break Through Across All Domains

2026-09-17


Preface

Have you ever had such an experience——

In a crowded subway, shouting into the phone three times"Hello, can you hear me", yet the other party still says"It's too noisy on your end”

During home meetings, keyboard sounds, air conditioning noise, and pet barks intertwine, and colleagues say"Is the signal bad on your end"

While recording during cyclingVlog, the whooshing wind drowns out other ambient sounds, stripping the original cycling memories of their auditory embellishment

In the mobile internet andAIoTera,"Being heard clearly"has become the baseline requirement for voice interaction. However, the acoustic environment of the real world is far more complex than that of a laboratory——The rumble of subways, keyboard clicks in offices, road noise inside vehicles, and wind outdoors constantly threaten the quality of voice communication.


Today, we will dissect the evolutionary logic and breakthrough strategies of audio noise reduction technology from four dimensions: user scenarios, core requirements, industry pain points, and solutions.


1. Noise is Everywhere: Four Scenarios, Four Types of"Unable to Hear Clearly"

Application scenarios for audio noise reduction have penetrated every corner of life, roughly categorized into four major types:


Mobile communication scenarios represent the most fundamental area of demand.

When users make phone calls or video chats in environments such as subways, streets, or shopping malls, background noise often reaches as high as70-85dB, directly masking voice signals and causing the other party to frequently ask"Hello, can you hear me".


Remote work scenarios have seen explosive growth in recent years.

Noise in home environments is dominated by non-stationary components——Keyboard typing (transient,2-5msduration), low-frequency hum from air conditioners (stationary,<500Hz), and family conversations (directional interference). These noises overlap highly with speech in the time-frequency domain, making it difficult for a single algorithm to address them all. Furthermore, multi-participant scenarios require noise reduction algorithms to support full-duplex communication, meaning that while speaking at the near end, echo leaking from the speaker and ambient noise must be suppressed in real-time. This places stringent requirements on the causality of the algorithm (causality).

Figure 1 Time-Frequency Domain: Comparison of Speech and Noise Features

In-vehicle scenarios are typical high-noise battlefields.

When vehicle speed reaches80km/h, interior noise can exceed75dB. The superposition of engine, tire, and wind noise, combined with voice interference from multiple occupants, poses significant challenges to in-car calls and voice assistants.

Figure 2 Extreme SNR Scenarios: In-Vehicle Calls

Smart wearables and home scenarios impose even higher demands on noise reduction.

TWSHeadphones need to switch seamlessly between noise cancellation and transparency modes; smart speakers need to extract wake words from TV sounds and conversations at a far-field distance of3-5meters; outdoor recording equipment needs to suppress wind noise to enhance the clarity of target audio signals.


Figure 3 Cycling Scenario: Wind Speed Noise Interference


The commonality among these scenarios is that users are no longer satisfied with merely"Being able to hear", but rather demand"Hearing clearly, hearing naturally, and hearing comfortably".

Figure 4 Noise Types by Application Scenario

II. Good noise reduction is not just about"making sounds quieter"

Based on the aforementioned scenarios, the core requirements for audio noise reduction can be summarized into four dimensions:

Figure 5 Four Dimensions of Core Requirements

First, speech intelligibility takes priority.


The primary goal in call and voice interaction scenarios is to ensure the other party can understand every word, rather than simply pursuing maximum noise suppression. Excessive noise reduction that results in muffled speech or missing characters actually reduces communication efficiency.

In terms of objective evaluation,STOI(Short-Time Objective Intelligibility) predicts intelligibility,PESQ/PLCMOSwhile assessing audio naturalness; subjective evaluation follows theITU-T P.835standard, separately scoringSIG(signal distortion),BAK(background noise interference), andOVRL(overall quality) on a5-point scale. High-end conferencing systems requireOVRL≥3.5, whereas hearing aid scenarios have even lower tolerance forSIG, demanding≥4.0.


Second, real-time performance and low latency.

Real-time calls typically require algorithm latency below20ms, while live streaming monitor feeds even demand below10ms. This means the algorithm must complete analysis and processing within extremely short audio frames, precluding the use of high-complexity offline models. Frame length and frame shift directly determine algorithm latency:20msFrame length+10msand frame shift are common configurations in speech processing, but they introduce at least20msof algorithm latency. If叠加AECof20-40msand codec20msare added, the total latency easily exceeds theITU-T G.114recommended upper limit of150ms.


Therefore, low-latency noise reduction often adopts aggressive configurations with10msframe lengths and50%overlap, at the cost of reduced spectral resolution (only400frequency bins@48kHz), increasing the difficulty of harmonic restoration.


Third, audio naturalness.

Video conferencing and live streaming scenarios demand extremely high speech naturalness. Noise-reduced speech must not exhibit"metallic sounds","tubular sounds", or obvious spectral holes; it needs to preserve the original timbre and emotion of the voice.


Fourth, balance between computational power and power consumption.


TWSDevices such as earphones and smartwatches have limited battery capacity. Noise reduction algorithms must find the optimal balance between effectiveness and power consumption, without sacrificing battery life for the sake of noise reduction.100msLatency@48kHz/16bitMono channel requires4800samples, with the circular buffer occupying approximately9.6KB; if adopting a32bitfloating-point deep learning model, the parameter count must be compressed to the hundreds ofKBlevel, necessitating reliance onINT8Quantization, weight sharing, or knowledge distillation. Taking the awinic Dijiang®audio solution as an example, the computing power required for front-end noise reduction in voice assistants is only about60MCPS, and in call scenarios about200MCPS,ROM+RAMtotaling less than1MB, to enable long-term operation on small-form-factor devices such asTWS.


III. Ideals Are Rich, Reality Is Harsh: Five Major Challenges in Noise Reduction

Despite clear demands, engineering implementation faces numerous thorny pain points:


The game between steady-state and non-steady-state noise.

Traditional spectral subtraction andWienerfiltering work reasonably well for steady-state noise like air conditioners and fans, but often prove helpless against non-steady-state noise such as keyboard typing, knocking, and clattering tableware. Such noise exhibits obvious time-domain characteristics and rapid spectral changes, making it difficult for traditional statistical models to track.



Figure 6 awinic Call Scenario Noise Reduction Solution


The challenge of wind noise processing.

Wind blowing across a microphone generates extremely high low-frequency energy that severely overlaps with the speech band. Traditional algorithms tend to misidentify wind noise as speech or over-suppress it, causing severe speech distortion. In outdoor scenarios, wind noise directly determines the baseline of call experience. Wind noise is not sound waves but random pressure disturbances on the microphone diaphragm caused by turbulence; its spectral energy concentrates in the low frequencies (<1kHz) and highly overlaps with the fundamental frequency of speech. The diaphragm of the feedforward microphone bears direct wind pressure, resulting in severe low-frequency saturation; even with windproof foam, physical attenuation remains limited.


At the algorithmic level, wind noise detection typically relies on a joint decision based on zero-crossing rate (ZCR) and low-frequency energy ratio. However, under high wind speeds (>5m/s), wind noise power exceeds speech by more than10dB, rendering traditionalVAD+gain control strategies ineffective. This necessitates reliable voice detection based on multi-microphone phase differences or bone conduction assistance. In this field, awinic has developed competitive solutions.


Taking the awinic Dijiang®AIwind noise reduction algorithm as an example, the overall subjective score of the solution improves by approximately 33%; advantages remain stable across different wind speed scenarios: improvement of about4m/s in 27%scenarios, about5-7m/s in outdoor scenarios, and about 35%in extreme strong-wind scenarios13m/s . Objective testing references the 43%standard, with significant improvements in theP.835metric across scenarios such as brisk walking, jogging, and motorcycle riding.SIGThe contradiction between noise reduction and speech damage.


The essence of noise reduction is

subtraction": subtracting estimated noise from the noisy signal. Insufficient subtraction leaves noticeable noise residue; excessive subtraction removes speech details along with the noise, generating"——musical noise"(") or spectral holes. Achieving balance between the two is the core difficulty in algorithm tuning.musical noiseFigure 8


Balancing Noise Reduction and Speech Damage The adaptation dilemma of multi-scenario switching.

As users move from a quiet office to a noisy street and then into a fast-moving car, the acoustic environment changes instantaneously. Fixed-parameter noise reduction solutions cannot adapt to such dynamic changes, while adaptive algorithms face trade-offs between convergence speed and stability.

The ceiling of hardware computing power.



High-end

chips can run complex deep learning models, but mid-to-low-end chips andDSPearbuds have extremely limited computing power. Achieving noise reduction effects close to those of high-end devices under computing constraints is the key bottleneck for mass production and deployment.TWSFigure 9

__TRANS_0150__ Computing power-Memory-Power consumption: The Impossible Trinity


IV. The Path to Breakthrough: From"Lone Operations"To"Collaborative Evolution"

Facing the aforementioned pain points, audio noise reduction technology is undergoing a paradigm shift from"lone operations"to"collaborative evolution":


The fusion of traditional algorithms and deep learning.

Pure deep learning models (such asRNNoise,DeepFilterNet) perform excellently in suppressing non-stationary noise but consume significant computing power. In engineering practice, a hybrid architecture of"traditional+AI"methods is often adopted: traditional algorithms quickly converge to handle stationary noise, while deep learning models are responsible for the refined suppression of non-stationary and transient noise. The two complement each other, balancing effectiveness and efficiency. This fused architecture has been validated in mass-production solutions. awinic Dihuang®is a typical representative of this approach——Relying on9+ years of independent algorithm R&D accumulation and verification across over20hundred million installed devices, it deeply combines the fast convergence of traditional signal processing with the refined suppression ofAImodels, achieving excellent noise reduction effects under the limited computing power of wearable devices.


Multi-microphone arrays and beamforming.

Single-microphone noise reduction can only utilize frequency-domain information, whereas dual-microphone or multi-microphone arrays can introduce spatial dimensions. Through beamforming (Beamforming) technology, the pickup beam is directed toward the target speaker while suppressing noise from other directions. This has become standard in scenarios such as automotive, conference pods, and smart speakers. For wearable devices, it is necessary to adoptINT8quantization (weights and activation values), structured pruning (removing unimportant channels), and knowledge distillation (large models guiding small models).


Furthermore, leveraging the spatial information of microphone arrays, single-channel noise reduction can be cascaded with beamforming (MVDR/GSC): beamforming provides6-12dBdirectional gain, reducing the difficulty for subsequent single-channel algorithms and allowing the use of lighter-weight networks. awinic Dihuang®supports2-8__TRANS_0182__8__TRANS_0183__


Multi-sensor fusion and scene awareness.

High-end solutions introduce bone conduction microphones or accelerometers as reliable references for voice activity——Bone conduction signals are immune to air-borne noise, picking up only speech transmitted through skull vibrations. Through cross-correlation analysis of bone conduction and air conduction, a high-confidenceVADcan be constructed to guide the timing of noise reduction gain application. In scenarios like translation that have special requirements for pickup distance, joint near-field and far-field pickup solutions combining multiple air microphones withVPU(Voice Pickup Units) have been implemented. These distinguish between near-field and far-field sound sources usingDM-BFdual-differential microphone arrays, cooperating with near-field/far-field suppression algorithms to achieve independent sound zone pickup.


Meanwhile, based on lightweightCNNscene classifiers (quiet/street/automotive/strong wind), noise reduction strategies are switched in real time: wind noise-specific filters are enabled in strong wind scenarios, and low-frequency suppression is strengthened in automotive scenarios, achieving true adaptability.


Conclusion

The ultimate goal of audio noise reduction technology is not to make the environment silent, but to guard that clarity of communication amidst the noise. From spectral subtraction to deep learning, from single microphones to arrays, from fixed parameters to adaptive intelligence, every evolution of technology narrows the gap between the ideal experience and the real environment.

In the future2-3years, uplink audio algorithms will continue to evolve towardsAI-ization, multi-modality, and all-scenario directions:2026will be implemented universally this yearAInoise cancellation, specific scenariosAInoise cancellation (indoor/outdoor/office/subway stations), directional beamforming (super-directional,8figure-eight, cardioid, near-field and far-field beams), sound source localization,AIecho cancellation, audio zoom, voice wake-up (VAD+KWS), etc.;2027will further advance multi-modal noise suppression (button press, hand gesture), voiceprint recognition (voiceprint activity detection+voiceprint noise reduction), large modelsASRthird-party integration, and other innovative applications. Through a lightweight scene classification network, it real-time determines whether the current environment is"quiet indoor","street","in-vehicle"or"windy environment", automatically switching to the corresponding noise cancellation strategy and parameter configuration to achieve true all-scenario adaptability.

Meanwhile, awinic's full-link layout at the hardware level——from audioPA(1W~300Wpower driver),Codec/ADC/DAC(high SNR>112dB, high sampling rate low latency<100uS), audio bus (100Mbpsin-vehicleSoundwire),ADSP/NPUprovides a complete system-level solution from chip to algorithm for algorithm implementation.