Weekly Research Digest — 2026-08-16

25 new entries this week across 4 topic areas.


Vision-Language-Action (VLA) Models

ReleaseVenueSignificance
in-context-vla-agentic-tool-use-post-training In-Context VLA: Language via In-Context Post-Training and Agentic Tool UsearXiv⭐ HIGH PRIORITY: Reframes VLA “reasoning” as tool-mediated grounding instead of free-text CoT, SOTA across RoboCasa-GR1/SimplerEnv/LIBERO plus real hardware
explicit-language-memory-long-horizon-planning-vla Explicit Language Memory for Long-Horizon Planning in VLAarXivInterpretable textual memory alternative to latent-memory long-horizon VLAs
handedit-embodiment-aware-image-editing-dataset-dexterous-hands HandEdit: Embodiment-Aware Image-Editing Dataset for Dexterous HandsarXiv⭐ HIGH PRIORITY: 200M+ instance hand-to-robot image-editing dataset spanning 26 embodiments
xiaomi-robotics-1-xr1-scaling-real-world-manipulation-pretraining Xiaomi-Robotics-1 (XR-1): Scaling Real-World Manipulation PretrainingarXiv⭐ HIGH PRIORITY: 100,000+ hours real-world pretraining with auto-labeling pipeline, open weights
crosstracer-hierarchical-cross-embodiment-navigation CrossTracer: Hierarchical Cross-Embodiment NavigationarXivPixel-space waypoint interface beats Gemini-2.5-Pro by 28% relative on NaviTrace
hymes-skills-in-weights-memory-in-code HyMeS: Skills in Weights, Memory in CodearXivSplits motor skills (VLA weights) from memory management (coding agent) for non-Markovian tasks

World Models for Robotics

ReleaseVenueSignificance
lila-wam-lightweight-latent-reasoning-world-action-model LiLa-WAM: Lightweight Latent Reasoning World-Action ModelarXiv0.5B-param WAM trainable on a single consumer GPU, 90.5% on RoboTwin 2.0
lingbot-va-20-native-video-action-pretraining-generalizable-robot-control Native Video-Action Pretraining for Generalizable Robot ControlarXivTrains a video-action model from scratch rather than adapting a generic video generator
adawam-adaptive-multimodal-reasoning-world-action-models AdaWAM: Adaptive Multi-Modal Reasoning for World Action ModelsarXivRouter triggers expensive visual “dreaming” only when needed, cutting inference cost
vera-turning-video-models-into-generalist-robot-policies VERA: Turning Video Models into Generalist Robot PoliciesarXivDecouples embodiment-agnostic video planner from embodiment-specific inverse dynamics model
efficient-sim-to-real-transfer-world-action-models-synthetic-priors Efficient Sim-to-Real Transfer of World-Action Models from Synthetic PriorsarXiv⭐ HIGH PRIORITY: First zero-shot sim-to-real transfer of a WAM trained purely on synthetic data (Cosmos Policy + AnyTask)
joyai-sim-simulation-interconversion-toolchain-embodied-data-pyramid JoyAI-Sim: Simulation-Enabled Interconversion ToolchainarXivBidirectional robot-sim-human data conversion with human-validated evaluation
contactguard-pre-contact-execution-monitoring-latent-world-models ContactGuard: Pre-Contact Execution Monitoring with Latent World ModelsarXivUses a world model as a safety/verification layer to flag failures before contact
dreamx-phi-10-action-conditioned-video-world-model-manipulation DreamX-Phi 1.0: Action-Conditioned Video World ModelarXiv1st place on WorldArena 2.0 Track 1 leaderboard via geometric action-conditioning

Reinforcement Learning for Robotics

ReleaseVenueSignificance
bora-offline-online-residual-adaptation-dexterous-vla BORA: Offline RL and Online Residual Adaptation for Dexterous VLAarXiv⭐ HIGH PRIORITY: Frozen-backbone residual RL with human-in-the-loop correction for dexterous hands
vista-vision-grounded-physics-validated-umi-data-vla VISTA: Vision-Grounded and Physics-Validated UMI Data for VLA TrainingarXiv⭐ HIGH PRIORITY: 8M-pair VQA dataset plus physical-feasibility filtering for UMI-collected data
dexpie-stable-dexterous-policy-improvement-real-world-experience DexPIE: Stable Dexterous Policy Improvement from Real-World ExperiencearXiv⭐ HIGH PRIORITY: DAgger-style correction gives 37% success-rate gain on real dexterous tasks
midas-adaptation-generalist-robot-policies-minimal-data MiDAS: Adaptation of Generalist Robot Policies with Minimal DataarXiv⭐ HIGH PRIORITY: Single-demonstration BC anchor plus online residual RL on real bimanual hardware
scenesmith-agentic-scene-generation-robot-training-data SceneSmith: Agentic Scene Generation for Scalable Robot Training Datablog (MIT News)⭐ HIGH PRIORITY: Multi-agent pipeline generates diverse simulated scenes to address data scarcity

Humanoid Robotics

ReleaseVenueSignificance
human-as-humanoid-zero-shot-learning-ego-exo-videos Human-as-Humanoid: Zero-Shot Learning from Ego-Exo Human VideosarXiv⭐ HIGH PRIORITY: 4.8-7.2x demonstration throughput gain vs. teleoperation via FK-aware retargeting
skild-brain-10-omni-bodied-robot-foundation-model Skild Brain 1.0: An Omni-Bodied Robot Foundation Modelblog (Skild AI)Vault’s first Skild AI entry — single-network claim across ~200 hardware platforms
figure-ai-scaling-helix-logistics Scaling Helix: A New State of the Art in Humanoid Logisticsblog (Figure AI)Vault’s first Figure AI entry — deformable-object handling, 20% faster package throughput
figure-03-autonomous-ladder-climb Figure 03 Autonomous Ladder Climbpress/socialPublic demo of hard whole-body loco-manipulation; unverified/thin on technical detail
labimus-humanoid-dexterous-manipulation-chemical-laboratory Labimus: Humanoid Dexterous Manipulation in Chemical LaboratoryarXivFirst benchmark for humanoid manipulation in scientific lab settings
simple-humanoid-loco-manipulation-simulation-benchmark SIMPLE: Simulation-Based Policy Learning and Evaluation for Humanoid Loco-ManipulationarXivDual-engine (MuJoCo + IsaacSim) benchmark, 60 tasks across 50 scenes

Tooling note: This week’s research relied entirely on WebSearch snippets — direct WebFetch access to arxiv.org, huggingface.co, and most company blog domains was blocked by the environment’s network egress policy for every research agent, and several WebSearch budgets were exhausted before every lead could be chased. A handful of promising leads (SelfWAM, FlowPilot, RoboDojo, AthenaZero bimanual juggling, HumanoidMimicGen, an IMU-based humanoid teleoperation paper) were found but not included this week due to insufficient corroborating detail — worth a follow-up pass once fetch access or search budget is restored.

Generated automatically. All entries verified via web search.