ICML May 2026 Conference paper
Active Continual Learning with Metaplastic Binary Bayesian Neural Networks
K. Cottart, T. Ballet, D. Bonnet, D. Querlioz
Proceedings of Machine Learning Research 306
BiMU (Binary Metaplasticity from Uncertainty) is a Bayesian continual-learning rule for binary neural networks: each one-bit weight keeps a probability of being +1 or −1, and its own uncertainty sets how far it may move, so the network keeps learning on long non-stationary streams without replay or task boundaries, and uses the same uncertainty to decide which samples are worth labelling.
Why
Always-on edge systems run inference continuously on a tight energy budget, yet conditions keep drifting: new users, sensor aging, changing environments. They see each data point once, rare events matter most, and they must notice when their own predictions are unreliable. They have to keep learning without forgetting, and they cannot afford to learn from everything. Binary neural networks fit the hardware, one bit per weight, but standard training forgets, and Bayesian binary networks trained for long enough stop learning altogether: their posterior saturates.
How it works
Each weight is a Bernoulli variable over ±1 with a single free parameter λ, whose mean is tanh λ and variance 1 − tanh² λ. BiMU updates λ with three terms, all derived from a bounded-memory variational objective:
- Plasticity: the gradient of the loss, the data term.
- Forgetting: a relaxation towards the prior, scaled by 1/N and gated by the weight’s own variance, so old evidence fades after a memory window of N samples.
- Stability: an uncertainty-dependent step size, so a weight that has committed resists flipping unless the data keeps asking for it.
Because the posterior never saturates, its uncertainty stays informative. BiMU reads it twice: to regulate how much each synapse learns, and to decide which inputs are worth a label and a backward pass.
Results
- 1000 tasks of Permuted MNIST, one sample at a time: 90.30% mean accuracy on the last five tasks, against 41.12% for BayesBiNN, 29.35% for the straight-through estimator and 10.27% for Synaptic Metaplasticity, with strong out-of-distribution detection throughout.
- OpenLORIS-Object, 12 tasks in one pass: 73.61% mean accuracy with features compressed to 1,024 dimensions (BayesBiNN 72.01%, STE 52.88%), up to 90.62% on the full features.
- Active continual learning on a long-tailed OpenLORIS-Object stream: 88.70% accuracy while labelling and updating on only 3.1% of the stream, a 32× reduction, against 87.76% when training on every sample.
- On an STM32 microcontroller: 12.4× less compute per data point (45.0 ms instead of 559.4 ms) for 1.7 points of accuracy.
Using it
BiMU is an Optax gradient transformation in the code repository:
import optax
from optimizers.bimu import bimu
optimizer = bimu(lr=1.0, lr_max=1.0, N=1000) # N: the memory window
opt_state = optimizer.init(params) # params: latent λ per binary weight
updates, opt_state = optimizer.update(grads, opt_state, params)
params = optax.apply_updates(params, updates)
The configurations in configurations/ reproduce each table and figure of the paper.
Abstract
Always-on edge systems must keep learning as conditions change under tight compute budgets and must detect unreliable predictions. Bayesian binary neural networks are attractive in this setting, but mean-field Bernoulli posteriors can saturate on long non-stationary streams, wiping out epistemic uncertainty and freezing plasticity. We propose BiMU, derived from a bounded-memory variational objective that balances stability, plasticity, and forgetting. BiMU combines a data term with controlled relaxation toward the prior and an uncertainty-dependent step size that prevents saturation and sustains informative uncertainty. This non-degenerate posterior enables fully online, buffer-free active querying via Monte Carlo disagreement, reducing label queries and backpropagation updates under imbalance. BiMU sustains learning and strong OOD detection on 1000-tasks Permuted-MNIST, and on OpenLORIS-Object achieves up to 32× label/update savings at matched accuracy under class imbalance and feature compression.
Related work
BiMU is the binary counterpart of MESU (Metaplasticity from Synaptic Uncertainty), derived for one-bit synapses instead of Gaussian ones. The same principle, that how certain a synapse is decides how much it may change, is being carried into the material itself with magneto-ionic synapses.
Questions
What is BiMU?
BiMU (Binary Metaplasticity from Uncertainty) is a Bayesian continual-learning rule for binary neural networks, published at ICML 2026. Each one-bit weight keeps a probability of being +1 or −1, and its own uncertainty sets how far it may move. The network keeps learning on long non-stationary streams without replay or task boundaries, and uses the same uncertainty to decide which samples are worth labelling.
Why do binary Bayesian neural networks stop learning in continual learning?
Because of posterior saturation. With a mean-field Bernoulli posterior, every new batch of evidence pushes each weight's probability further towards 0 or 1. On a long non-stationary stream the weights become certain, epistemic uncertainty disappears and the network freezes. BiMU prevents this with a bounded memory window N: evidence older than the window is relaxed back towards the prior.
How is BiMU different from BayesBiNN?
They approximate different posteriors. BayesBiNN (Meng et al., 2020) fits a mean-field Bernoulli posterior to all the data seen so far, p(ω | D1:t): every batch adds evidence, so the posterior keeps sharpening. That is the right answer in a stationary world, but on a drifting stream the weights saturate, and continual use needs task boundaries to reset the prior. BiMU targets the posterior over a sliding window of roughly the last N updates, p(ω | Dt−N:t). Its variational free energy weighs the previous posterior by (1 − 1/N), the prior by 1/N, and adds the likelihood of the current batch: old evidence is discounted geometrically and replaced by the prior, so certainty stays bounded without any task boundary. A second-order expansion of that objective gives BiMU's closed-form update, with its uncertainty-dependent step size. After 1000 tasks of Permuted MNIST, BiMU reaches 90.30% mean accuracy on the last five tasks, against 41.12% for BayesBiNN.
How does active learning work with BiMU?
The network draws K samples of its binary weights and runs K cheap forward passes on each incoming input. When the samples disagree, the input is uncertain: only then does the network ask for a label and run the expensive backward update. It is fully online and needs no replay buffer. On a long-tailed OpenLORIS-Object stream, BiMU reaches 88.70% accuracy while labelling and updating on only 3.1% of the stream (32× fewer), against 87.76% when training on every sample.
Can BiMU run on a microcontroller?
Yes. The paper measures it on an STM32 NUCLEO-64 board at 216 MHz. Learning only when unsure brings the expected compute per data point to 45.0 ms, against 559.4 ms when updating on every sample: a 12.4× reduction, for 89.30% vs 90.98% accuracy. At inference each weight is a single bit.
What does the memory window N do?
N sets how much evidence each synapse keeps, and so how many synapses stay uncertain enough to learn. Too small and the network forgets (27.09% mean accuracy over the last five of 1000 tasks at N = 100); too large and it saturates (83.70% at N = 100,000); in between it balances stability and plasticity (89.07% at N = 1,000).
Where is the code?
At github.com/kellian-cottart/active-continual-learning-bayesianbinn. BiMU is an Optax gradient transformation (optimizers/bimu.py, JAX and Equinox), with one script per table and figure of the paper.
BibTeX
@inproceedings{cottart2026active,
title = {Active Continual Learning with Metaplastic Binary
Bayesian Neural Networks},
author = {Cottart, Kellian and Ballet, Th\'eo and
Bonnet, Djohan and Querlioz, Damien},
booktitle = {Proceedings of the 43rd International Conference
on Machine Learning},
series = {PMLR},
volume = {306},
year = {2026},
eprint = {2605.30198},
archivePrefix = {arXiv}
}