Taylan Soydan, Miguel A. Bessa, Dirk Mohr, Rui Barreira
7 min
Abstract
Selective state space models (SSMs), such as Mamba, achieve strong per-token expressivity by making the time discretization step $\TildeΔ$ a learned function of the input. However, in doing so, $\TildeΔ$ no longer equals the physical time gap $Δ$ between consecutive observations, limiting the ability of these models to handle irregular time series. Continuous time SSMs, such as S5, keep $\TildeΔ\equivΔ$ and therefore handle irregular timestamps natively, but their dynamics remain linear time invariant (LTI), limiting per token expressivity. We propose \textbf{TIDES}, a selective SSM variant that reconciles selective and continuous architectures by moving input dependence off the step size and onto the diagonal state matrix. As a result, $\TildeΔ\equivΔ$ as in S5, allowing the model to handle irregular timestamps natively without sacrificing the per-token expressivity that makes selective SSMs effective. We show this on a novel \emph{Fading Flash} experimental benchmark, a compact controlled diagnostic for sequence models that jointly tests input dependence and extrapolation to out-of-distribution $Δ$ values, and isolates the distinct failure modes of current state-of-the-art architectures that TIDES avoids by construction. On large-scale benchmarks, TIDES sets the new best average rank on UEA time series classification and the Physiome ODE regression benchmark, and matches or exceeds the reference baseline model on 6 of 8 natively irregular datasets from astronomy, agriculture, neuromorphic sensing, and climate events. Code available at: \url{https://github.com/TaylanSoydan/TIDES}.
Alex: Why keep frequency fixed? Seems like an arbitrary line to draw.
Sam: Their reasoning is interpretive. A per-token decay rate means "forget faster at an event boundary, hold longer when something matters." That's the closest analogue to Mamba's selectivity. Changing frequency every token would redefine the model's basis of dynamics at every step. They test it, and making frequency input-dependent slightly hurt accuracy. They admit that argument is mostly conceptual.
Alex: You mentioned a controlled task. How do they show the two failure modes cleanly?
Sam: It's a toy they call Fading Flash. Picture a row of 40 detectors. Sparse flashes hit some, and each detector glows and then fades. The row is split into zones with slow, medium or fast fade rates. A global clock setting stretches or compresses the whole trajectory. The model sees the flashes and zone labels and must predict the glow.
Alex: So a fixed model can't produce three decay rates at once, and a learned gate might not handle a new clock speed.
Sam: That's the design. They trained on clock steps in a middle range, then tested well outside it. S5 stayed stable across clock speeds but couldn't capture the zone-dependent decays. The Mamba surrogate fit training well, then drifted outside the training range. TIDES learned the three decays and stayed robust. At the smallest test step, TIDES had a relative error around seven percent. The official Mamba-2 and Mamba-3 were around 140 to 150 percent.
Alex: A clean toy is nice, but it's built to make the point. What happens on real benchmarks?
Sam: Several tiers. On six classification datasets from the UEA archive, a standard multivariate suite, TIDES had the best average accuracy, about 66 percent versus about 65 for Mamba-3. It had the best average rank, but won outright on only two of the six. On Physiome-ODE, a forecasting benchmark of 50 irregular datasets from biophysical simulations, it again had the best average rank. But it tied a continuous-time baseline, LinODEnet, at 16 wins each. The confidence interval on their rank difference includes zero.
Alex: So on the big suites, it's competitive at the top, not clearly ahead.
Sam: That's a fair reading, and the authors call it statistically on par on Physiome. The most telling test is a random-drop experiment on EigenWorms, a worm-motion dataset. They trained with half the time steps randomly removed, then tested at different drop rates. They kept real timestamps so the gaps were visible.
Alex: And what happened when the test sampling got much sparser than training?
Sam: At the extreme, with 90 percent dropped, the official Mamba models fell from the low seventies to somewhere between forty and sixty percent accuracy. A Rough Transformer baseline collapsed similarly. TIDES stayed roughly flat, around 71 percent at that extreme. Variants with fixed input and output projections, including S5, sat near chance throughout. That's a small model, around thirty thousand parameters, with three seeds.
Alex: And on data that's irregular by nature, not artificially thinned?
Sam: Eight datasets across astronomy, crop satellite series, neuromorphic event sensors and climate records. It matched or beat the reference baseline on 6 of the 8. It was slightly behind on one variable-star set and on a gesture dataset.
Alex: What's the limitation you'd weigh most?
Sam: Scope. They didn't test language modelling or regular-grid long-range tasks, and they say plainly it isn't a general Mamba replacement. They also stress the toy task isn't evidence of real-world advantage. Beyond that, there are more knobs to tune than S5, and their scan isn't a hardware-optimized kernel.
Alex: So who should sit down with the full paper, and where should they start?
Sam: Anyone modelling irregular time series, or bolting timestamps onto Mamba. Start with the method section on where input dependence goes and why. Then read the random-drop ablation, which isolates each choice. The discussion's design principle is worth a skim.
Alex: And for everyone else, the line to carry?
Sam: Let time enter through the discretization, and put the flexibility in the dynamics. Mix the two, and your model forgets what a second is.
Alex: A small relocation with a clear payoff.