Reflect — making robot memory recoverable
I study how robots should react when their memory or plan stops matching the world. Reflect combines native vision-language-action policies with explicit progress tracking and recovery decisions. I tested the system in simulated manipulation, compared matched failure and recovery runs, and investigated why refreshing a plan can help in one situation and hurt in another.
Inspect the recovery study
Can a robot recover after its memory becomes stale or wrong?
A two-query hold expiry clears the external memory record and resumes the full mission: putting both moka pots on the stove. The baseline keeps holding indefinitely when that record is stale or corrupt. Every arm has a 65-call limit.
| Study | Memory | Population | Pairs | Expiry success | Indefinite success | Exact paired p |
|---|---|---|---|---|---|---|
| Confirmation | Stale | All planned | 10 | 9/10 | 0/10 | 0.0039 |
| Confirmation | Stale | Matched audit | 9 | 8/9 | 0/9 | 0.0078 |
| Confirmation | Corrupt | All planned | 10 | 7/10 | 0/10 | 0.0156 |
| Confirmation | Corrupt | Matched audit | 7 | 6/7 | 0/7 | 0.0312 |
| Pilot | Stale | All planned | 4 | 3/4 | 0/4 | 0.2500 |
| Pilot | Stale | Matched audit | 4 | 3/4 | 0/4 | 0.2500 |
| Pilot | Corrupt | All planned | 4 | 3/4 | 0/4 | 0.2500 |
| Pilot | Corrupt | Matched audit | 3 | 2/3 | 0/3 | 0.5000 |
Interpretation boundary. The all-pairs chart preserves the planned population, including simulator drift. Only the audited subset—with identical observations, prompts, seeds, actions, and controller decisions before intervention—supports the paired causal interpretation. This small, single-task study tests recovery from the router’s unproductive hold; it does not establish an advantage over a healthy direct policy or transfer to another mission.
Replay note. These videos sample one frame per policy query. They are replay illustrations of the selected conditions, not recordings at the controller’s real-time frame rate.
Is replanning always useful?
No simple rule survived both studies. FlexPi fresh replanning completed 12/15 runs versus 11/15 for stale chunks, but one seed reversed that aggregate result. On two released Unitree traces, changing correction timing affected action agreement in different directions.
Interpretation boundary. The Unitree comparison replays two released traces; it does not measure robot task success. The FlexPi count is seed-sensitive and should not be generalized beyond this matrix.
What comes next?
The next experiments are designed to separate recovery value from the advantages of having any external router at all.