TIDES: Moving Input Dependence Off the Step Size to Give Selective State Space Models Native Time-Awareness

Published · AI Daily — AI-assisted deep research, methodology & disclosure

TIDES is a selective state space model (SSM) that fixes a quiet weakness in Mamba. In Mamba, the discretization step size depends on the input, so it must act as both a clock and a content gate. Real timestamps have nowhere to go. TIDES moves the input dependence off the step size and onto the diagonal state matrix. The step size then equals the true time gap, so irregular sampling is handled natively, and per-token selectivity is kept. The authors report the best average rank on the UEA classification and Physiome ODE regression benchmarks. On 8 natively irregular datasets, TIDES matches or beats the reference baseline on 6.

State space models (SSMs) are one of the most watched alternatives to the Transformer. S4 and S5 treat a sequence as samples of a continuous-time linear system, and their inference cost grows linearly with sequence length. Mamba added a selection mechanism: its parameters change with the input, so the model can decide what to remember and what to forget based on content. That selectivity has a side effect that gets little attention. The model's sense of real time becomes vague. The paper TIDES: Implicit Time-Awareness in Selective State Space Models targets exactly this. The authors are Taylan Soydan, Miguel A. Bessa, Dirk Mohr and Rui Barreira. It was submitted to arXiv on 10 May 2026 and revised to version 2 on 5 October 2026.

The problem sits in the discretization step size. In Mamba, the step size, usually written as delta, is produced from the current input. It plays two roles at once. It is the time gap between two observations, and it is also the content gate that sets how fast memory decays. With regular sampling, mixing the two roles does little harm, because every real gap is the same. When timestamps are irregular, the true gap has no entry point into the model. The network has to guess it from data. That is what 'implicit time-awareness' means in the title. Mamba does react to time, but only implicitly, with no physical constraint, and the time signal can tangle with the content signal. S5 does the opposite. Its step size follows the real time gap, so it handles irregular sampling natively, but its dynamics are linear and time-invariant, which limits per-token expressivity.

The core change in TIDES, in the words of the abstract, is to move input dependence off the step size and onto the diagonal state matrix. The step size stays equal to the physical time gap, so irregular timestamps are handled natively. Input dependence is carried by the diagonal entries of the state matrix, so each token can still be selective. In short: time stays with time, and content stays with content. Two conflicting demands that were squeezed into one variable now live in two. A short explanation helps, with one warning: this is an inference from the abstract. The exact parameterization, stability constraints and initialization should be checked in the paper and the released code. After zero-order-hold discretization, a diagonal continuous system has a decay factor of exp(delta times a) for each state channel. In Mamba, a is a fixed negative number and delta(x) changes with the input, so the factor is exp(delta(x) times a). In TIDES, delta is the real time difference t_k minus t_(k-1), and a(x) changes with the input, so the factor is exp(delta times a(x)). Under regular sampling the two forms look alike. Under irregular sampling they differ. In TIDES, decay scales with the true gap automatically: the longer the pause, the more is forgotten by default, while the content decides the rate. The model no longer has to infer time from data. Time becomes part of the structure.

On results, we quote only what the abstract states. The authors first built a small controlled diagnostic called Fading Flash, designed to isolate time-awareness so that different SSMs can be compared on that one skill. Then TIDES reached a new best average rank on the UEA time series classification benchmark and on the Physiome ODE regression benchmark. On 8 natively irregular datasets, covering astronomy, agriculture, neuromorphic sensing and climate events, TIDES matches or exceeds the reference baseline on 6. The abstract we read gives no accuracy figures, latency or parameter counts, and this article does not invent them. For the exact cost and quality trade-off, read the paper's tables and reproduce the runs. For developers and enterprises, the value lies in irregular sampling: intermittent lab tests in patient monitoring, tick-level events in finance, sensors that drop out or report asynchronously, event-camera output, and uneven astronomical exposures. The usual fixes are to interpolate onto a regular grid, which adds bias, or to concatenate the time gap as an extra feature and let the network learn its use. TIDES offers a third route: put the time gap straight into the discretization and guarantee it by structure. The code is on GitHub under TaylanSoydan/TIDES. Because the model is still a linear-time recurrent design, inference cost should be close to Mamba's, but that needs measurement and cannot be read from the abstract.

The work has limits. First, the abstract says 6 of 8 datasets, so on 2 the model did not match or beat the baseline, and the boundary of use depends on data traits. Second, the method relies on trustworthy timestamps; noisy or missing ones will reduce the gain. Third, placing selectivity on a diagonal matrix still leaves state channels uncoupled, and the ceiling on expressivity remains to be tested. Fourth, it is not yet known whether a parallel-scan kernel can match the heavily tuned Mamba implementation in speed. Future directions include hybrids with large language models, multi-timescale modeling, and continuous-time event streams. Overall, TIDES is a small, clear structural fix, and teams working on irregular time series should try it.

Sources