ReCAT: Remember, Count, And Time

Structured Recurrent Memory for Robot Manipulation

1Intuitive Robots Lab (IRL), Karlsruhe Institute of Technology (KIT), 2University of Edinburgh,
3NVIDIA, 4Robotics Institute Germany (RIG)
*Core contributors
66.7%
Real-robot memory tasks
vs. 8.3% for the strongest short-history baseline
62.4%
RMBench
best or tied-best on 6 of 9 tasks
95.3%
LIBERO
at 59 ms per policy step (16.9 Hz)

Same view, different action

Problem

Memory-dependent manipulation requires robots to make decisions using information that is no longer available to their current sensors: recalling an earlier visual cue, tracking task progress, counting repeated events, or estimating elapsed time.

After one scoop and after two, the robot sees the same scene, yet it must either scoop again or put the shovel back. Only memory of what already happened can decide.

Method

Method

We present ReCAT, a language-conditioned policy with structured recurrent memory. An instruction-conditioned encoder forms features from the current observation. A recurrent memory integrates the observation stream through Mamba-2 layers and one causal attention layer. A flow-matching Transformer decoder reads the current and the historical representation through separate cross-attention in every block.

ReCAT architecture: modality encoders, visual compressor, multimodal observation encoder, recurrent memory, and flow-matching action head.

Observation encoder

Compresses images, instruction and proprioception into one feature zt per step.

Recurrent memory

Mamba-2 layers plus one causal attention layer carry a compact state mt through the episode.

Flow-matching decoder

Reads zt and mt through separate cross-attention in every block.

Step-by-step walkthrough of the architecture.

Interchangeable update rules

Mamba-2 (default)

diag(at) ht−1 + vtkt⊤

Accumulate: repeated events build up in the state.

Mamba-3

diag(ateiθt) ht−1 + vtkt⊤

Accumulate + rotate: adds a phase that can track time.

Gated DeltaNet-2

αt(I − βtktkt⊤) ht−1 + βtvtkt⊤

Replace: a repeated cue overwrites its old content.

Three real-robot memory tasks

Franka Emika Panda (DROID setup), 20 rollouts per task. Each video compares GMP, X-VLA-H and ReCAT.

Count Add exactly two scoops of soil, put the shovel down, then plant the flower. 45 demos · ~1,200 steps

Plant: frame sequence with the memory-dependent decision point outlined.

Time Put the pot on the stove, wait ~30 s while the scene stays still, then add the pepper. 45 demos · ~980 steps

Pot Timer: frame sequence with the memory-dependent decision point outlined.

Remember Move the sponge to the plate, then return it to its unmarked start position. 48 demos · ~480 steps

Sponge: frame sequence with the memory-dependent decision point outlined.

Baselines fail exactly where memory begins

Result

ReCAT reaches 95.3% average success on LIBERO and 62.4% on RMBench, with the best or tied-best result on six of nine tasks. On three real-robot tasks probing spatial recall, event counting and interval timing, the best ReCAT variant reaches 66.7% average success, against 8.3% for the strongest short-history baseline. Controlled comparisons show that the observation encoder and every-block memory conditioning are needed for this performance, and that update rules behave differently as robot memory: additive updates have the highest observed success on counting and timing, and delta-rule updates on spatial recall.

Policies without persistent memory succeed in 0–8.3% of rollouts. ReCAT variants reach 50.0–66.7%. The 95% Wilson intervals for X-VLA + history ([3.6%, 18.1%]) and ReCAT-Mamba-2 ([54.1%, 77.3%]) do not overlap. Stage-wise columns show where each policy fails: the baselines usually finish the early manipulation and then fail at the decision that needs memory.

Method Plant Pot Timer Sponge Avg.
1 scoop2 scoops3 or moreFlower planted
Success
Pot on
stove
No waitWrong wait≈30 s
wait
Pepper in
the pot
SuccessPick sponge,
put on plate
Returned to
wrong position
Returned to
same position
Success
Baselines
Diffusion Policy–––0–––––0––00.0
Diffusion Policy + history151000100250025015–00.0
X-VLA–––0–––––0––00.0
X-VLA + history010851010003554057565108.3
Gated Memory Policy (GMP)652010159520105355505006.7
Ours
ReCAT-Mamba-201000851000010095952002066.7
ReCAT-Mamba-30950901000010010010000063.3
ReCAT-Gated DeltaNet-20700201000010090904004050.0

All values are percentages of 20 rollouts with randomized start positions. Red columns are failure modes, green columns are progress stages, and the darkest green is final success. Plant scoop columns are mutually exclusive. “No wait” and “Wrong wait” left the waiting state early or outside the 29–31 s window. Avg. is the mean success over the three tasks.

Simulation benchmarks

LIBERO

LIBERO is mostly Markovian, so it tests whether full-history recurrence keeps performance when the current observation is enough.

MethodSpatialObjectGoalLIBERO-10Avg.LIBERO-Plus
Diffusion Policy78.392.568.350.572.4–
X-VLA (0.9B, pretrained)98.298.697.897.698.171.4
ReCAT96.497.297.290.495.364.2

RMBench

Nine long-horizon bimanual tasks. In M(n) tasks, actions depend on several past events that are no longer observable. ReCAT has the best or tied-best score on six tasks, including three of the four M(n) tasks.

TaskTypeDPACTπ0.5*X-VLA*MEM-0EventVLA*MemoryVLA*MemERReCAT
90M80M3B0.9B10B4B7.3B10B374M
Observe and Pick UpM(1)1199421274
Rearrange BlocksM(1)02913138996531796
Put Back BlockM(1)0011189095810100
Swap BlocksM(1)112241667967614100
Swap TM(1)20215314879764
Average M(1)6.46.814.411.852.879.044.29.072.8
Battery TryM(n)101916262835332743
Blocks Ranking TryM(n)100611881531299
Cover BlocksM(n)00026897694050
Press ButtonM(n)000003006
Average M(n)5.04.85.57.328.554.038.819.849.5
Overall average5.85.910.49.842.067.941.813.862.4

* Uses large-scale pretraining. EventVLA has a higher overall average; it builds on a 4B backbone, about ten times larger than ReCAT, with an extra pretraining phase.

Efficiency

All policies measured on the robot's inference GPU (RTX 4060 Ti) with the same settings. Higher SPARC means smoother trajectories.

MetricReCATDP-HX-VLA-HGMP
Total / trainable params374M / 142M267.9M / 267.9M881.9M / 881.9M282.2M / 188.7M
Peak inference VRAM1.44 GB0.65 GB2.55 GB1.09 GB
Inference latency59 ms (16.9 Hz)110 ms (9 Hz)288 ms (3.5 Hz)93 ms (10.8 Hz)
Smoothness (SPARC)−5.16−6.19−6.76−13.06

Failure modes and limitations

Most remaining ReCAT failures happen before memory is needed. In Sponge, the start positions vary widely, and failed rollouts usually miss the grasp or placement; whenever the sponge reaches the plate, ReCAT returns it to the right place. In Plant, the failures come from inserting the flower stem.

Every decoder block must read the memory

Memory → decoder

Avg. real-robot success (%)

Cross-attn every block
66.7
Late fusion
38.3
AdaLN
31.7
Scale
28.3
Dropout
20
Gate
8.3

Encoder & components

Avg. real-robot success (%)

Hybrid Mamba–attn
66.7
Full Transformer
25
No multimodal enc.
0
No memory
0

State width (Plant)

Plant success (%)

dstate 32
15
dstate 64
20
dstate 128
85
dstate 256
0

Every-block memory cross-attention and the hybrid encoder matter most; removing memory drops all tasks to 0%.

BibTeX

@article{vanjani2026recat,
  title         = {ReCAT: Remember, Count, And Time: Structured Recurrent Memory for Robot Manipulation},
  author        = {Vanjani, Pankhuri and Hatab, Mostafa and Mizrakli, Can and Shaj, Vaisakh and
                   Li, Zhuoyue and Reuss, Moritz and Lioutikov, Rudolf},
  journal       = {arXiv preprint arXiv:2609.35200},
  year          = {2026},
  eprint        = {2609.35200},
  archivePrefix = {arXiv},
  primaryClass  = {cs.RO}
}