Weekly Research Digest — 2026-07-20
9 new entries this week across 3 topic areas (July 14–20 new papers, plus belated additions from late June / early July not previously logged).
Vision-Language-Action (VLA) Models
| Release | Venue | Significance |
|---|---|---|
| reflex-real-time-vla-streaming-inference Reflex | arXiv 2607.14695 / ICML 2026 | Streaming inference for VLAs enables real-time control without sacrificing model capability; ICML acceptance validates the approach |
| harness-vla-memory-guided-agents-frozen-vla Harness VLA | arXiv 2607.08448 / IEEE RA-L Aug 2026 | Memory-guided agentic planner steers frozen VLAs through deployment shifts without any retraining |
| vla-corrector-detect-correct-inference-adaptive-action-horizon VLA-Corrector | arXiv 2607.01804 / ECCV 2026 | Lightweight LVM+OGG detect-and-correct layer closes the reactivity gap in action-chunked VLAs |
| drop-then-recovery-vla-redundancy-language-backbone Drop-Then-Recovery | arXiv 2606.27755 | Language backbone is >50% removable in VLAs with no performance loss; vision+action paths are the critical components |
World Models for Robotics
| Release | Venue | Significance |
|---|---|---|
| nvidia-cosmos-3-edge-ondevice-world-model-jetson NVIDIA Cosmos 3 Edge | blog / press release (Jul 16) | 4B on-device world model for Jetson Thor; moves Cosmos 3 physical-AI reasoning off the cloud and onto the robot |
| xiaomi-robotics-u0-embodied-synthesis-world-foundation-model Xiaomi-Robotics-U0 | arXiv 2607.11643 | 38B open world foundation model; 82× data-gen speedup and +26pp OOD manipulation improvement via synthetic augmentation |
| taco-tactile-world-model-vla-post-training TACO | arXiv 2607.02840 | Visuo-tactile world model converts real-world failures into self-supervised corrective VLA post-training signal |
Reinforcement Learning for Robotics
| Release | Venue | Significance |
|---|---|---|
| never-too-late-for-force-reactive-force-injection-vla Never Too Late for Force | arXiv 2607.14236 / RSS 2026 Workshop | Reactive force injection post-training accelerates contact-rich VLA adaptation without large-scale force pretraining |
| dexverse-modular-benchmark-multi-task-dexterous-manipulation DexVerse | arXiv 2607.08751 | 100-task modular benchmark across 3 arms × 6 hands with configurable visual variation for dexterous policy evaluation |
Generated automatically. All entries verified via web search.