Weekly Research Digest — 2026-08-09

36 new entries this week across 4 topic areas. This week’s sweep surfaced an unusually large cluster of VLA/policy post-training work — 17 of the 36 entries touch test-time adaptation, human-in-the-loop correction, preference optimization, or data synthesis for VLAs, flagged ⭐ HIGH PRIORITY below.


Vision-Language-Action (VLA) Models

ReleaseVenueSignificance
ego2robot-scalable-robot-data-synthesis-egocentric-human-data Ego2RobotarXiv⭐ HIGH PRIORITY: Largest ego-to-robot synthetic dataset yet (18.5k hrs, 15 morphologies) for VLA pretraining.
prigo-test-time-primitive-guidance-diffusion-flow-policies PriGoarXiv⭐ HIGH PRIORITY: Test-time primitive guidance steers diffusion/flow policies without retraining.
retrieve-then-steer-online-success-memory-test-time-adaptation-vla Retrieve-then-SteerarXiv⭐ HIGH PRIORITY: Online success-memory test-time adaptation for frozen generative VLAs.
wizard-robotic-policy-adaptation-weight-space-meta-learning WIZARDarXiv⭐ HIGH PRIORITY: Predicts LoRA weight updates in one forward pass — gradient-free VLA adaptation.
set-supervised-diffusion-policy-action-chunking-corrections Set-Supervised Diffusion PolicyarXiv / RSS 2026⭐ HIGH PRIORITY: Contrastive supervision from DAgger-style correction pairs, peer-reviewed at RSS.
vla-ad-offline-semantic-guidance-efficient-vla-distillation VLA-ADarXiv⭐ HIGH PRIORITY: 44x VLA distillation via offline VLM semantic supervision; headline numbers need independent checking.
retrieval-vla-training-free-in-context-adaptation Retrieval-VLACVPR 2026⭐ HIGH PRIORITY: Training-free episodic-memory in-context adaptation, peer-reviewed at CVPR.
dypes-vla-shared-dynamics-embodiment-specific-control-cross-embodiment DyPES-VLAarXivSplits shared dynamics priors from embodiment-specific control across arm/dual-arm/humanoid.
cross-embodiment-transfer-behavior-aligned-representations Cross-Embodiment Transfer via Behavior-Aligned RepresentationsarXivNew benchmark + representation study for transferring from action-free data.
cloak-zero-shot-cross-embodiment-masking-end-effector CloakarXivSimple end-effector masking trick enables zero-shot gripper/hand transfer.
primitivevla-reusable-motion-primitives-efficient-generalizable-manipulation PrimitiveVLAarXivPrimitive-based “disassemble & assemble” framing matches 100%-data baseline with 50% data on LIBERO.
instant-fold-in-context-imitation-learning-deformable-manipulation Instant-FoldarXivSingle-demo in-context learning for cloth folding, sim-to-real zero-shot claim.
generalist-ai-gen1-thousand-hands-cross-embodiment-end-effector Generalist AI — GEN-1 “Thousand Hands”blog⭐ HIGH PRIORITY: 9,000 gripper/tool variants from 500k+ hrs real data; unverified vendor claims.

World Models for Robotics

ReleaseVenueSignificance
feedback-world-model-precise-guidance-diffusion-policy Feedback World ModelarXiv⭐ HIGH PRIORITY: Closed-loop world-model guidance corrects diffusion policies online, no retraining.
wayve-gaia-4-multimodal-world-models-closed-loop-simulation Wayve GAIA-4blogClosed-loop multimodal (video+radar) driving world model for safety evaluation at scale.
deform360-massive-multiview-visuotactile-dataset-deformable-world-models Deform360arXiv / ECCV 2026Large visuotactile dataset for comparing 2D vs. 3D world-model paradigms on deformables.
how-should-world-models-be-evaluated-decision-making-position How Should World Models Be Evaluated?arXivPosition paper arguing for interventional-reasoning-centric evaluation over visual realism.
gauge-measurement-grounded-benchmark-physical-fidelity-simulation-video-world-models GAUGEarXivMeasurement-grounded physical-fidelity benchmark spanning simulators and video world models.
vlaflow-unified-training-framework-co-training-future-latent-alignment VLAFlowarXiv⭐ HIGH PRIORITY: Controlled ablation shows future-latent-alignment + language co-training beats either alone.
dream-tac-unified-tactile-world-action-model-contact-rich-manipulation Dream-TacarXivTactile world-action model, +31.7% action accuracy on contact-rich tasks over vision-only.
geniworld-generalizable-interactive-world-model-visual-actions GeniWorldarXivURDF-rendered visual-action conditioning reduces scene overfitting, doubles as policy evaluator.
omega-0-latent-predictive-world-action-model-concurrent-humanoid-loco-manipulation ω₀arXivConcurrent (non-hierarchical) humanoid loco-manipulation WAM + new 40hr real-world household dataset.

Reinforcement Learning for Robotics

ReleaseVenueSignificance
rtcf-retrieve-in-time-correct-in-frequency-test-time-correction-vla RTCFarXiv⭐ HIGH PRIORITY: Causal memory alignment corrects long-horizon VLA execution error, training-free.
unintervene-agentic-intervention-efficient-real-world-rl UniIntervenearXiv⭐ HIGH PRIORITY: Learns when to auto-intervene, cutting human supervision cost in real-world RL.
mapl-multi-objective-preference-learning-robot-locomotion MAPLarXiv⭐ HIGH PRIORITY: LLM-generated multi-criteria preferences replace hand-engineered locomotion rewards.
rarm-confidence-gated-progress-reward-modeling-rl-manipulation RARMarXivReward model from a single demonstration, confidence-gated to resist reward hacking.
densereward-dense-reward-learning-failure-synthesis-robotic-manipulation DenseRewardarXiv⭐ HIGH PRIORITY: Dense per-timestep reward model trained on 27k synthesized failure trajectories.
progress-reward-modeling-robotic-learning-survey Progress Reward Modeling SurveyarXivFirst unified survey of the fast-growing progress-reward-modeling literature.
torl-vla-tactile-guided-online-reinforcement-learning-contact-rich-manipulation TORL-VLAarXiv⭐ HIGH PRIORITY: Online RL refines tactile VLA actions at deployment for shifted contact conditions.
d-vla-high-concurrency-distributed-asynchronous-rl-framework D-VLAarXiv⭐ HIGH PRIORITY: Distributed async RL training infrastructure for billion-parameter VLAs.
sp3o-reinforcement-learning-segment-preferences-without-reward-modeling SP3OarXiv⭐ HIGH PRIORITY: Reward-model-free, critic-free segment-preference RL; robotics + LLM domains.

Humanoid Robotics

ReleaseVenueSignificance
teleopit-full-embodiment-humanoid-teleoperation-system TeleopitarXivFull-body VR teleoperation with hand-agnostic retargeting, 90-95% success from 96 demos.
thorarena-benchmarking-humanoid-physical-interaction-motion-force-demonstrations ThorArenaarXivFirst benchmark jointly scoring tracking, stability, and robustness under real external force.
heft-heavy-payload-full-size-humanoid-teleoperation-privileged-motion-guidance HEFTarXivTeacher-student control for heavy-payload (24kg on 65kg robot) full-size humanoid teleoperation.
humanoidarena-benchmarking-egocentric-hierarchical-whole-body-learning HumanoidArenaarXivBenchmark shows hierarchical humanoid control is fragile across different low-level trackers.
light-loco-parkour-versatile-perceptive-whole-body-locomotion-multi-skill-distillation Light-Loco-ParkourarXivEnd-to-end onboard-only perceptive locomotion with learned, gate-free skill transitions.

Generated automatically. All entries verified via web search.