New Noise Reduction
RadioClear AI – A New Approach to Speech Noise Reduction
Improving Weak Signal Reception
One of the biggest challenges when listening to weak radio signals is separating speech from background noise. As signals become weaker, the noise can overwhelm the speech, making conversations difficult or sometimes impossible to understand.
Conventional noise reduction has been available in software-defined radios for many years. Most systems analyse the received audio, estimate which frequencies contain noise, and reduce their amplitude. This can work very well, but there is always a compromise. Too little suppression leaves distracting background noise, while too much can make speech sound metallic, watery or unnatural. Important speech information may disappear along with the noise.
For some time, I have wanted to develop a better approach for SDR Console, particularly for weak SSB reception. The objective was not simply to produce quieter audio, but to improve intelligibility while preserving the natural characteristics of the speaker's voice.
The result is RadioClear AI, a new noise-reduction system combining adaptive digital signal processing with a recurrent neural network and complex temporal filtering.
Comparison With RNNoise
RadioClear AI and RNNoise both use neural networks to improve speech intelligibility by reducing unwanted background noise, but they employ different approaches.
- RNNoise is a well-established, computationally efficient solution that combines traditional signal processing with recurrent neural networks to estimate frequency-dependent suppression gains. It performs well across a variety of speech applications and has the advantage of being lightweight and widely tested.
- RadioClear AI has been developed specifically with radio communications in mind, particularly weak SSB signals affected by receiver noise. It combines adaptive noise estimation and speech-aware spectral processing with a recurrent neural network that generates complex temporal filter coefficients. Unlike conventional gain-based suppression, this approach uses information from consecutive audio frames to adjust both the amplitude and phase of the enhanced signal.
Both approaches have strengths, and their relative performance depends on the type of interference, signal conditions and training material. RNNoise benefits from its maturity, efficiency and general-purpose design, while RadioClear offers a more specialised processing architecture tailored to weak radio reception.
In initial listening tests, RadioClear AI has produced encouraging results, with cleaner background noise and natural-sounding speech compared with RNNoise on the signals tested. However, these observations are subjective, and controlled comparisons across a wider range of recordings would be needed to establish a consistent performance advantage. Further improvements to RadioClear's training data are planned, particularly using cleaner speech recordings and a greater variety of speakers and receiver noise conditions.
Combining Traditional DSP with Artificial Intelligence
Rather than relying entirely on artificial intelligence, RadioClear AI uses two complementary processing stages.
- The first stage is based on conventional digital signal processing. It continuously estimates the background noise and analyses the incoming audio to identify frequencies likely to contain speech. It also examines harmonic relationships, which are particularly important for voiced speech.
- This information is used by an adaptive spectral suppressor to reduce background noise while protecting important speech components.
The result is an initial noise-reduced signal that provides a stable foundation for the neural network. This arrangement has an important advantage. The neural network does not need to learn everything about noise estimation and suppression from scratch. Instead, it receives a partially cleaned signal together with detailed information about the speech and noise characteristics. The neural network can then concentrate on improving the result.
Looking Beyond Individual Audio Frames
One of the limitations of conventional spectral noise reduction is that it generally operates by adjusting the amplitude of individual frequency components. However, speech is not a collection of unrelated frequencies. It is a continuously evolving signal with strong relationships between consecutive moments in time. RadioClear AI takes advantage of this.
The incoming audio is divided into overlapping frames, and each frame is converted into a frequency-domain representation using a Fast Fourier Transform (FFT).
The neural network receives 400 features describing the signal across 40 frequency bands. These include information about signal power, estimated noise, signal-to-noise ratio, speech probability, harmonic content and the relationships between consecutive complex spectra.
At the heart of the system are two Gated Recurrent Unit (GRU) layers. Unlike a simple feed-forward neural network, a GRU maintains information about previous frames, allowing the system to recognise how speech changes over time.
This temporal information is particularly valuable when the speech is partially obscured by noise.
Complex Temporal Filtering
The most significant difference between RadioClear AI and a conventional neural noise reducer is the way the network processes the audio. Many neural noise-reduction systems generate a set of gains that are applied to individual frequency bins. Although effective, this approach primarily controls amplitude.
RadioClear AI goes further by generating complex-valued filter coefficients. For each frequency bin, the system combines information from the current audio frame and the previous two frames. The neural network determines how these spectra should be combined to produce the enhanced signal.
Because the coefficients are complex, the system can modify both amplitude and phase. It can also exploit the relationships between consecutive frames rather than treating each frame independently. This gives the network considerably more flexibility than a conventional spectral gain mask.
The neural network produces coefficients for 40 frequency bands, which are interpolated across the 257 FFT frequency bins. A three-tap complex temporal filter then generates the enhanced spectrum, which is converted back into audio.
The complete process operates continuously in real time.
Training the Neural Network
The neural network is trained offline using examples of speech mixed with receiver noise. During training, the system is provided with both the noisy signal and the corresponding clean speech. It learns to generate complex filter coefficients that make the processed spectrum resemble the clean reference as closely as possible.
An important development decision was to train the network against the actual reconstructed complex spectrum, rather than simply asking it to reproduce a set of ideal filter coefficients. This means the training process directly rewards improvements in the enhanced signal.
The current network contains approximately 568,000 parameters and was trained using speech mixed with real receiver noise at signal-to-noise ratios ranging from +15 dB to −10 dB.
The training and evaluation process showed substantial reductions in complex spectral error compared with the initial deterministic noise-reduction stage. The improvements were particularly encouraging at the lowest signal-to-noise ratios, where additional noise reduction is most valuable.
Of course, mathematical measurements are only part of the story. A lower spectral error does not automatically guarantee better-sounding speech, so listening tests remain essential.
Real-Time Processing and Listening Results
The neural network was originally developed and trained using Python and PyTorch. The complete inference engine was then implemented in C++ for integration into SDR Console. To ensure that the C++ implementation behaved correctly, its results were compared directly with the Python reference. This included the recurrent network, complex coefficient generation, frequency interpolation and final spectral filtering. The results agreed to floating-point precision, giving confidence that the real-time implementation reproduces the trained model accurately.
Initial listening tests have been very encouraging.
On weak SSB signals, RadioClear AI produces natural-sounding speech with substantial background-noise reduction. In my own listening comparisons, the results have been noticeably better than RNNoise, particularly in preserving speech quality while reducing unwanted noise.
An additional aggressiveness control has also proved useful, allowing stronger suppression of residual noise where desired.
As always, the best setting depends on the received signal. Maximum noise reduction is not necessarily the same as maximum intelligibility, so the listener retains control over the amount of processing applied.
What Comes Next?
Although the current results are very promising, development is far from finished.
One area where I expect further improvements is the training material. The first model was trained using a relatively small amount of speech, and the supposedly clean reference recording contained a little background noise.
This is not ideal. If noise is present in the clean training reference, the network can learn to preserve some of that noise because it is treated as part of the wanted signal.
The next step will therefore be to create a much better training dataset using genuinely clean speech recordings, ideally from several different speakers, together with a wider variety of real receiver noise.
I intend to retain the existing neural architecture initially so that improvements from better training data can be evaluated independently.
There is also potential to extend the system to other receiving modes, including AM and music, while retaining the existing 8 kHz neural processing engine.
Final Thoughts
RadioClear AI represents a significant step forward in my work on speech noise reduction for SDR Console. The most interesting aspect is not simply that it uses artificial intelligence, but how conventional DSP, recurrent neural processing and complex temporal filtering work together.
By combining an adaptive noise estimator, speech-aware spectral processing and a neural network capable of exploiting relationships between consecutive audio frames, the system can perform noise reduction that would be difficult to achieve using conventional spectral attenuation alone.
The aim has always been straightforward: make weak radio signals easier and more pleasant to listen to, without destroying the speech we are trying to recover.
The initial results suggest that this approach has considerable potential, and I look forward to seeing how much further it can be improved with better training recordings.
RadioClear AI is still evolving, but it is already producing some of the best weak-signal speech noise reduction I have achieved in SDR Console.










