Skip to content
@H-EmbodVis

H-EmbodVis

Embodied Vision Projects from Huazhong University of Science and Technology

H-EmbodVis

Embodied Vision · World Models · Autonomous Driving · 3D Scene Understanding

H-EmbodVis (Huazhong University of Science and Technology Embodied Vision Projects) is a research initiative. We primarily focus on Embodied AI, while also exploring Autonomous Driving and Generative Models.


🔬 Research Areas

We focus on building intelligent systems that can perceive, understand, and interact with the physical world. Key directions include:

  • Embodied AI & Agents: Integrating vision, language, and action planning.
  • World Models for Autonomous Driving: Developing end-to-end driving frameworks and simulators.
  • 3D Vision & Point Cloud Analysis: Efficient architectures for 3D representation learning.
  • Multimodal Foundation Models: Large-scale models for diverse data modalities.

🌟 Featured Projects

Autonomous Driving & World Models

  • HERMES (ICCV 2025) A Unified Self-Driving World Model for Simultaneous 3D Scene Understanding and Generation.
  • Orion (ICCV 2025) Holistic End-to-End Autonomous Driving via Vision-Language Instructed Action Generation.
  • Awesome-World-Model Curated collection of papers on World Models for Autonomous Driving and Robotics.

3D Vision & Efficient Computing

  • PointMamba (NeurIPS 2024) State Space Models (Mamba) applied to Point Cloud Analysis.
  • UniSeg3D (NeurIPS 2024) A Unified Framework for 3D Scene Understanding.
  • PointGST (IEEE TPAMI) Parameter-Efficient Fine-Tuning in Spectral Domain for Point Cloud Learning.
  • EasyCache Training-Free Video Diffusion Acceleration.

Multimodal & Embodied Agents

  • NAUTILUS (NeurIPS 2025) A Large Multimodal Model for Underwater Scene Understanding.
  • GRANT (AAAI 2026 Oral) Teaching Embodied Agents for Parallel Task Execution.
  • MERGE (NeurIPS 2025) Unifying Generation and Depth Estimation via Text-to-Image Diffusion Models.

Collaboration

We are always looking for passionate collaborators and students.

  • Connect: Reach out via email (dkliang@hust.edu.cn).
  • Reuse: Creating impactful open-source software is a core value. Please cite our papers if you use our code.

🌐 Website | 🎓 Google Scholar | 📂 Repositories

Pinned Loading

  1. VEGA-3D VEGA-3D Public

    [ECCV 2026] Generation Models Know Space: Unleashing Implicit 3D Priors for Scene Understanding

    Python 421 23

  2. TurboVLA TurboVLA Public

    TurboVLA: Real-Time Vision-Language-Action Model at 32 Hz on an RTX 4090 with <1 GB VRAM

    Python 481 61

  3. DOMINO DOMINO Public

    [ECCV 2026] Towards Generalizable Robotic Manipulation in Dynamic Environments

    Python 229 11

  4. HyDRA HyDRA Public

    Out of Sight but Not Out of Mind: Hybrid Memory for Dynamic Video World Models

    Python 278 14

  5. Orion Orion Public

    Forked from xiaomi-mlab/Orion

    [ICCV 2025] Official code of "ORION: A Holistic End-to-End Autonomous Driving Framework by Vision-Language Instructed Action Generation"

    Python

  6. PointGST PointGST Public

    Forked from jerryfeng2003/PointGST

    [IEEE TPAMI] Parameter-Efficient Fine-Tuning in Spectral Domain for Point Cloud Learning

    Python

Repositories

Showing 10 of 24 repositories
  • TurboVLA Public

    TurboVLA: Real-Time Vision-Language-Action Model at 32 Hz on an RTX 4090 with <1 GB VRAM

    H-EmbodVis/TurboVLA's past year of commit activity
    Python 481 Apache-2.0 61 11 1 Updated Sep 2, 2026
  • SimWAM Public

    SimWAM: A Simple World Action Model for End-to-End Autonomous Driving.

    H-EmbodVis/SimWAM's past year of commit activity
    Python 177 19 1 0 Updated Aug 27, 2026
  • DOMINO Public

    [ECCV 2026] Towards Generalizable Robotic Manipulation in Dynamic Environments

    H-EmbodVis/DOMINO's past year of commit activity
    Python 229 Apache-2.0 11 6 0 Updated Aug 18, 2026
  • ROAD Public

    ROAD: Reciprocal-Objective Alignment of Discriminative Semantics for 3D Shape Generation

    H-EmbodVis/ROAD's past year of commit activity
    Python 26 Apache-2.0 1 0 0 Updated Aug 4, 2026
  • HERMESV2 Public

    HERMES++: Toward a Unified Driving World Model for 3D Scene Understanding and Generation

    H-EmbodVis/HERMESV2's past year of commit activity
    Python 71 Apache-2.0 10 0 0 Updated Jul 28, 2026
  • HyDRA Public

    Out of Sight but Not Out of Mind: Hybrid Memory for Dynamic Video World Models

    H-EmbodVis/HyDRA's past year of commit activity
    Python 278 14 0 1 Updated Jul 23, 2026
  • VEGA-3D Public

    [ECCV 2026] Generation Models Know Space: Unleashing Implicit 3D Priors for Scene Understanding

    H-EmbodVis/VEGA-3D's past year of commit activity
    Python 421 Apache-2.0 23 5 0 Updated Jun 18, 2026
  • EasyCache Public

    Less is Enough: Training-Free Video Diffusion Acceleration via Runtime-Adaptive Caching

    H-EmbodVis/EasyCache's past year of commit activity
    Python 299 Apache-2.0 7 0 0 Updated May 12, 2026
  • NUMINA Public

    [CVPR 2026] When Numbers Speak: Aligning Textual Numerals and Visual Instances in Text-to-Video Diffusion Models

    H-EmbodVis/NUMINA's past year of commit activity
    Python 70 MIT 7 1 0 Updated Apr 11, 2026
  • PointTPA Public

    [CVPR 2026] PointTPA: Dynamic Network Parameter Adaptation for 3D Scene Understanding

    H-EmbodVis/PointTPA's past year of commit activity
    Python 35 MIT 1 1 0 Updated Apr 7, 2026

Top languages

Loading…

Most used topics

Loading…