22 new entries this week across 4 topic areas. This week’s cross-cutting sweep for VLA post-training / data-augmentation content turned up an unusually large batch — 17 of the 22 new entries are flagged high priority.
⭐ HIGH PRIORITY: Test-time-training “fast weights” scale VLA context to 8K timesteps, enabling one-shot in-context imitation and on-the-fly policy improvement without extra inference latency.
⭐ HIGH PRIORITY: “Action inversion” lets frozen flow/diffusion policies be DAgger-adapted from a handful of human corrections without full fine-tuning or online RL.
z1-efficient-rl-vla Z-1: Efficient Reinforcement Learning for Vision-Language-Action Models
arXiv
⭐ HIGH PRIORITY: GRPO-based RL post-training lifts π0.5 from 67.4% to 80.6% success across 24 RoboCasa tasks.
⭐ HIGH PRIORITY: Q-guided candidate selection plus human-in-the-loop corrections make RL fine-tuning of VLAs sample-efficient on precision/dynamic tasks.
⭐ HIGH PRIORITY: Uses synthetic robot video only for geometric auxiliary supervision (not pseudo-actions), avoiding the sim-to-real action mismatch of prior approaches.
Alibaba DAMO’s 2B/9B/122B-A10B embodied foundation model family adds contact-point prediction and native 3D grounding, evaluated across three distinct robot embodiments.
⭐ HIGH PRIORITY: Identifies why video-pretrained foundation models lose compositional generalization after being fine-tuned into VLAs, and fixes it with an inference-time guidance metric.
⭐ HIGH PRIORITY: Fei-Fei Li’s World Labs acquires a Stanford robotics-sim spinout to fuse generative 3D world models with simulation pipelines for synthetic robot training data.
Predicts embodiment-agnostic 3D interaction traces instead of pixels or actions, letting it pretrain from action-free video and transfer across embodiments.
⭐ HIGH PRIORITY: Multi-axis natural-language preferences (speed, safety, placement quality) give a denser reward signal than binary preference learning, +38pp over baselines.
⭐ HIGH PRIORITY: Test-time RL that uses a VLA’s own generation confidence as an intrinsic reward, approaching oracle RL performance with no ground-truth reward.
⭐ HIGH PRIORITY: Hybrid offline-online RL training retains RL’s OOD generalization advantage while cutting the costly online interaction needed to get there.
⭐ HIGH PRIORITY: Adapts generalist VLAs by running RL over language prompts instead of raw actions, composing existing skills for novel long-horizon OOD tasks.
⭐ HIGH PRIORITY: AI coding agents run the full policy-improvement loop on a real robot fleet, hitting 99% success on dexterous tasks and halving training time from 1→8 robots.
⭐ HIGH PRIORITY: Gaussian-splatting simulation generates teleop-free fine-tuning data for humanoid VLAs that matches or beats real teleoperated demonstrations.
⭐ HIGH PRIORITY: Portable VR/handheld gripper rig collects humanoid whole-body manipulation demonstrations with no robot or teleoperation hardware required.
GPU-parallel simulation RL lets Atlas lift 100+ lb loads using proprioceptive/force sensing rather than vision, generalizing zero-shot beyond its training weight range.
Generated automatically. All entries verified via web search.