Research Projects

Our research primarily focuses on Multimodal Large Language Model (MLLM), Embodied AI, and AI Agent, with an emphasis on perception, reasoning, and decision-making in interactive environments:

Multimodal Foundation Models
Embodied Intelligence
AI Agent
  • Multimodal Large Language Models (MLLMs) aim to build and train models capable of understanding, reasoning, and generating content across multiple modalities, including text, images, audio, and video. Our research in this area explores the enhancement of multitask learning capabilities, the advancement of high-resolution perception, and the design of unified architectures for multimodal understanding and generation. Please refer to our GitHub Orgnization about JiuTian MLLM ("九天"多模态大模型) for more details. [bilibili] [bilibili]
  • Embodied AI studies agents capable of perceiving, reasoning, and acting within physical environments. We aim to build systems based on MLLMs that integrate multimodal perception, instruction comprehension, and continuous action planning to perform complex 3D tasks such as manipulation, and interactive behaviors.

    Task instruction: Open the drawer, put the toy inside, and then close it. [bilibili]

    Task instruction: Fold the T-shirt carefully and finish with precise placement. [bilibili]

    Task instruction: Put the bread into a bowl and heat it in the microwave. [bilibili]

    Task instruction: Place both the lemon from the bowl and the apple on the table onto the plate, then put the lid on the bowl. [bilibili]

    Task instruction: Grasp a plate and place it on the tablecloth. Both the plates and the tablecloths come in multiple styles. [bilibili]

    Task instruction: Grasp three different fruits and place them onto the plate. For each rollout, a different combination of fruits is randomly sampled. [bilibili]

    To further enhance our embodied AI research, our lab has recently acquired the R1 Lite robot from GaLaXea AI.

  • AI Agent focuses on sequential decision-making across various complex environments such as MineCraft, mobile device. We develop systems based on MLLM that span a wide range of directions, including framework-based agents, native agents, and RL-enhanced reasoning. [bilibili]

    Task instruction: Get the search results for stay tonight near 'wembley stadium' for 1 adult. Add one result to wishlist. Confirm that this item is in the wishlist.

    Task instruction: Search today's weather in Shenzhen on Chrome, then write the temperature into today.md using Markor.

主持科研项目:

  • 2026.01-2028.07 科学幻觉检测修复与机理对齐工具集研究及应用 新一代人工智能国家科技重大专项课题
  • 2027.01-2030.12 面向多模态具身大模型的交互知识学习与更新方法研究 国家自然科学基金面上项目
  • 2027.01-2031.12 跨形态具身智能自主学习理论与进化方法 国家自然科学基金重点项目课题
  • 2026.07-2027.07 具身智能多专家Agent集成框架技术 华为技术合作项目
  • 2026.08-2027.08 面向具身操作的算法优化技术研究 百度松果计划开放课题项目
  • 2025.12-2026.12 基于4D时空感知的高泛化技能模型 中兴通讯产学研合作基金项目(机器人专项)
  • 2026.01-2028.12 融合领域知识可泛化的PCB表面缺陷检测多模态大模型关键技术研究 深圳市自然科学基金面上项目
  • 2024.01-2026.12 多模态高层语义驱动的深度伪造检测算法研究 国家自然科学基金青年科学项目(C类)
  • 2024.01-2026.12 开放世界下可泛化的多媒体攻击检测 国家自然科学基金优秀青年基金(海外)
  • 2024.01-2026.12 预训练大型语言模型知识驱动的多模态深度伪造检测算法研究 广东省自然科学基金面上项目
  • 2023.01-2025.12 多模态深度伪造检测算法研究 深圳市新引进高精尖缺人才科研启动经费项目

Terms of Releasing Implementation:

Software provided here is for personal research purposes only. Redistribution and commercial usage are not permitted. Feedback, applications, and further development are welcome. Contact shaorui[AT]hit.edu.cn for bugs and collaborations. All rights of the implementation are reserved by the authors.


© 2025 OrionLab