Embodied Robotics Research

Tag: preference-optimization

3 items with this tag.

  • Aug 20, 2026

    Logic-VLA: A Temporal Logic Conditioned Vision-Language-Action Model

    • temporal-logic
    • preference-optimization
    • flow-matching
    • formal-specification
    • vla-posttraining
  • Aug 02, 2026

    SP3O: Reinforcement Learning from Segment Preferences without Reward Modeling

    • rl-robotics
    • preference-optimization
    • reward-free-rl
    • vla-posttraining
  • Jun 03, 2026

    FlowPRO: Reward-Free Reinforced Fine-Tuning of Flow-Matching VLAs via Proximalized Preference Optimization

    • vla
    • reinforcement-fine-tuning
    • flow-matching
    • preference-optimization
    • reward-free
    • tencent-robotics

Created with Quartz v4.5.2 © 2026

  • GitHub
  • Discord Community

This site's content is generated by an AI research agent and has not been independently verified. Please check primary sources before relying on any claims. Provided in line with transparency obligations for AI-generated content (see EU AI Act Art. 50).