SkillMemo | arxiv 2026.8.6 | Paper Reading
AtlasVLA | arxiv 2026.8.7 | Paper Reading

AtlasVLA | arxiv 2026.8.7 | Paper Reading

AtlasVLA: Persistent World-Ego State Modeling for Vision-Language-Action Models

这同样是一篇关于记忆VLA的文章,特色的地方是文章强调了优势在于只使用了第一视角作为图像观测的输入并通过设计的记忆机制使得在只有有限视角输入的情况下依然取得了超过很多优秀baseline的结果,比较值得我学习,毕竟实验室目前采用的就是单一腕部相机视角进行实验的。

Read more
RoboTTT | arxiv 2026.7.16 | Paper Reading

RoboTTT | arxiv 2026.7.16 | Paper Reading

RoboTTT: Context Scaling for Robot Policies

这是一篇由LiFeiFei团队最新发表的文章,第一次将robot的上下文扩展到8k用到的方法思想其中部分我之前和AI讨论时也涉及过值得去反思一下。

Read more
DAM-VLA | ICRA 2026 | Paper Reading

DAM-VLA | ICRA 2026 | Paper Reading

DAM-VLA: A Dynamic Action Model-Based Vision-Language-Action Framework for Robot Manipulation

这是一篇发表在ICRA 2026的文章,讨论了最近我看到VLA领域比较火的记忆和动态中的一个话题,不过这篇文章主要聚焦的是动态动作的生成。

Read more
MemoryVLA++ | arxiv 2026.6.8 | Paper Reading

MemoryVLA++ | arxiv 2026.6.8 | Paper Reading

MemoryVLA++: Temporal Modeling via Memory and Imagination in Vision-Language-Action Models

作者在上一篇MemoryVLA的基础上增加了世界模型预测未来的能力从而进一步提升了模型的能力。

Read more
RoboDojo | arxiv 2026.7.5 | Paper Reading

RoboDojo | arxiv 2026.7.5 | Paper Reading

RoboDojo: A Unified Sim-and-Real Benchmark for Comprehensive Evaluation of Generalist Robot Manipulation Policies

作者提出了一个面向通用型机器人操控策略综合评估的仿真与真实世界统一基准

Read more
EventVLA | arxiv 2026.6.18 | Paper Reading

EventVLA | arxiv 2026.6.18 | Paper Reading

EventVLA: Event-Driven Visual Evidence Memory for Long-Horizon Vision-Language-Action Policies

作者提出了一个基于稀疏视觉证据记忆概念构建的端到端框架用于解决多帧存储类型的记忆VLA的问题

Read more
AHEAD | arxiv 2026.6.01 | Paper Reading

AHEAD | arxiv 2026.6.01 | Paper Reading

Intercepting the Future: Latent-Space Predictive World Model for Dynamic VLA Manipulation

作者提出了“预测-然后-执行”封装模块,为冻结的VLA模型增强了运动感知的潜在世界模型。核心值得我借鉴的是不用动VLA只在外面训练一个模块同样可以做到性能的巨大提升。

Read more
RoboMemArena | arxiv 2026.5.11 | Paper Reading

RoboMemArena | arxiv 2026.5.11 | Paper Reading

RoboMemArena: A Comprehensive and Challenging Robotic Memory Benchmark

作者提出了一个全新的专为记忆VLA的benchmark,并且通过我的进一步了解发现它确实更加综合和多样,对于我目前的实验很有帮助.

Read more
MoLe-VLA | AAAI 2026 | Paper Reading

MoLe-VLA | AAAI 2026 | Paper Reading

MoLe-VLA: Dynamic Layer-skipping Vision Language Action Model via Mixture-of-Layers for Efficient Robot Manipulation

作者提出的模型主要用于解决机器人系统在部署时面临挑战,通过作者设计的时空感知路由器和自知识蒸馏方法,基于层混合的方式构建了全新的VLA.

Read more
WeChatQQGoogle ScholarWeeklyLogRSS