The Layer Mystery of VLA: An Information-Theoretical Analysis of VLA Latent Interface
arXiv cs.AIen
arXiv:2609.36118v1 Announce Type: new Abstract: Vision-language-action (VLA) policies connect a pretrained vision-language backbone to an action head through a latent interface, but which backbone layers this interface should expose remains unclear. We study single-layer selection and multi-layer fusion for frozen backbones across three pretrained models and two manipulation benchmarks, LIBERO and CALVIN, with three policy-training seeds per configuration. Across three fusion mechanisms and three layer-subset strategies, 47 of 54 configurations underperform the best observed single-layer policy. Our stastical analysis further confirms that fusion's advantage is very limited. However, the bes
This is a short summary published by AI Global Wire. The full article is owned and hosted by arXiv cs.AI — open it there to read it in full.
Read the full story at arXiv cs.AI- Forskning
- Reglering
Related AI news
- OpenAI-HuggingFace: A Reproduction & Lessons for Alignment TestingarXiv cs.AI · September 30, 2026
- Neurosymbolic Routing for Reliable Reasoning on Resource-Constrained Edge DevicesarXiv cs.AI · September 30, 2026
- More Programs or More Rolls? Separating Coverage from Specialization in LLM HarnessesarXiv cs.AI · September 30, 2026
- Beyond Symmetric Agents: Cognitive Diversity and Multi-Agent Debate in Small Language ModelsarXiv cs.AI · September 30, 2026
- Grab a Coffee: Future-Aware Guidance for Discrete Diffusion with Compiled ObjectivesarXiv cs.AI · September 30, 2026
- SAGE: A Statistical Acceptance Gate for Self-Evolving AgentsarXiv cs.AI · September 30, 2026