This release marks a key transition in robotics foundation models, shifting from repurposing digital world models to designing them natively for the physical world. Instead of relying on fine-tuned digital content generation models, LingBot-VA 2.0 is built from scratch to meet the original demands of dynamic modeling, causal prediction, and real-time execution in physical environments.
Integrating world models with embodied AI has been one of the major focuses of the AI industry. However, most mainstream approaches rely on video generation models designed for digital content, which are then fine-tuned for robot control. Because content creation prioritizes visual quality and creativity, while robot control requires execution efficiency and physical accuracy, this forced adaptation often leads to knowledge forgetting and reduced generalization.
LingBot-VA 2.0 takes a different approach. By pre-training from scratch using an autoregressive architecture, the model is designed to understand how an action will change the environment and to decide the next step based on that causal prediction.
Core Architectural Innovations
To achieve this, LingBot-VA 2.0 is built on four core designs:
· Semantic Visual-Action Tokenizer: A new visual encoder that aligns semantic and action information during visual compression, helping the model translate “understanding instructions” into “completing actions” more effectively.
· Strict Causal Pre-training: The model uses an autoregressive architecture from the beginning, ensuring that visual predictions and action generation strictly follow a one-way time sequence.
· Mixture of Experts (MoE): This architecture expands model capacity without sacrificing inference efficiency, balancing performance and speed.
· Enhanced Asynchronous Inference: This mechanism enables real-time closed-loop control, allowing robots to predict future states while executing actions and continuously corrects its next decisions using the latest real-world observations.
These designs solve the common industry challenge of low execution efficiency in embodied world models, delivering a real-time inference speed of 150 Hz on a single GPU. Furthermore, the model can generalize to new tasks using as few as 20 demonstrations through in-context learning without parameter updates.
A Complete Embodied-Native Full-Stack
LingBot-VA 2.0 serves as the capstone of Robbyant’s recent launch week, which introduced six models that together form a complete embodied-native full-stack for perception, world simulation, and action:
LingBot-Depth 2.0
LingBot-Vision
LingBot-VLA 2.0
LingBot-World 2.0
LingBot-Video
LingBot-VA 2.0
Zhu Xing, CEO of Robbyant, noted, “Robbyant will continue to explore new limits in embodied intelligence while accelerating the development of an open technology and application ecosystem to expedite robot deployment in industrial and real-world scenarios.”
About Robbyant
Robbyant is an embodied intelligence company within Ant Group, dedicated to advancing embodied intelligence through cutting-edge software and hardware technologies. Robbyant independently develops foundational large models for embodied AI and actively explores next-generation intelligent devices, aiming to create robotic companions and caregivers that truly understand and enhance people’s everyday lives and deliver reliable intelligent services across key use cases, such as elderly care, medical assistance, and household tasks.
To learn more about Robbyant, please visit: www.robbyant.com
View source version on businesswire.com: https://www.businesswire.com/news/home/20260709654440/en/
언론연락처: Ant Group Media Inquires Vick Li Wei
이 뉴스는 기업·기관·단체가 뉴스와이어를 통해 배포한 보도자료입니다.
ⓒ 주식회사 에이아이크리에이티브랩 & www.aifilmjournal.com 무단전재-재배포금지
BEST 뉴스
-
Esri to Debut the Power of Where Collection at 2026 Esri User Conference
Esri (https://www.esri.com/en-us/what-is-gis/overview), the global leader in location intelligence, will debut the Power of Where Collection (https://powerofwhere.com/?aduc=Public_Relations&aduca=2026-MUL-Esri_Press_Books&aduco=press-release&adum=Press_Release&adut=pow-series&... -
Statement on Terminating the Letter of Intent With AI Financial Corporation
Statement from Matthew Nicoletti, Chief Strategy Officer, Perpetuals.com (Nasdaq: PDC), on the proposed transaction with AI Financial Corporation. “Perpetuals has decided not to further pursue the acquisition of AI Financial Corporation’s subsidiary Alt5 Sigma Canada, Inc. and the earlier letter... -
디자인 툴 필요없다… 비즈뿌리오, 알림톡 전용 ‘이미지 메이커’ 출시
비즈뿌리오 ‘이미지 메이커’ 기능 출시 기업 메시징 서비스 비즈뿌리오를 운영하는 다우기술(대표 김윤덕)은 카카오톡 알림톡에 포함되는 이미지를 손쉽게 완성할 수 있는 ‘이미지 메이커’ 기능을 지난 1일 출시했다고 밝혔다. 최근 카카오톡 알림톡은 단순 텍스트 형태를 넘어 브랜드 로... -
MUUT, 신세계 강남점 입성… 롯데 잠실 이어 팝업 확대
신세계 강남 MUUT 팝업 전경 패션 아이웨어 브랜드 뭍(MUUT)이 7월 9일부터 22일까지 신세계백화점 강남점 5층 팝업 스테이지에서 팝업 스토어를 운영한다. 이번 팝업은 MUUT의 다양한 아이웨어 제품과 브랜드가 제안하는 스타일을 직접 경험할 수 있는 공간으로 꾸며졌다. 롯데월드몰 잠... -
글렌알라키 ‘15년 컬렉터스 에디션 PART I’으로 JPM 어워즈 금상 수상
글렌알라키 15년 컬렉터스 에디션 파트 1 프리미엄 주류 수입 유통사 메타베브코리아는 자사의 대표 싱글몰트 위스키 브랜드인 글렌알라키의 한정판 패키지 ‘글렌알라키 15년 컬렉터스 에디션 PART I’이 일본 마케팅 업계 최고 권위의 시상식인 ‘제54회 Japan Promotional Marketing Award... -
소비자는 부담 덜고, 어가는 판로 확대… GS더프레시 ‘ESG 장어덮밥’ 출시
GS리테일이 운영하는 슈퍼마켓 GS더프레시는 국내산 민물장어 소비 촉진을 위해 ‘국내산 통한마리장어덮밥’을 출시한다고 14일 밝혔다. 이번 상품은 GS더프레시가 한국어촌어항공단, 해양수산부와 함께 추진하는 ‘Co:어촌 프로젝트’ 일환으로 기획됐다. Co:어촌 프로젝트는 기업의 상품 개발·유통 역량과 국내 어가의 ...
