About

Nanjing, China

[ Email | GitHub | Twitter | Google Scolar]

Short Bio·

Experience·

  • CMU: Zhang Lab, Research Assistant, Advisor: Ji Zhang (2024.9 - 2026.6)
    • Research on generalizable embodied navigation with vision-language models, focusing on the gap between general-purpose VLMs and the spatial reasoning required for navigation.
    • Developed hierarchical representations and navigation systems that ground VLM reasoning in objects, viewpoints, rooms, 3D geometry, and navigable frontiers.
    • Studied sim-to-real transfer of VLM-based navigation through hierarchical planning and real-robot deployment, bridging semantic reasoning with executable waypoint-based control.
    • Learned spatial-visual navigation intent from human demonstrations and explored grounded navigation through pixel-level goal prediction.
  • UC San Diego: Hao Su Lab, Visit Undergraduate, Advisor: Hao Su (2023.3 - )
    • Research on robotics and computer vision, especially works related to NeRF.
    • Modeling the physical properties of objects, including material and optical properties, from images captured from multiple perspectives.
  • Tsinghua University: Li Yi Lab, Undergraduate Researcher, Advisor: Li Yi (2022.9 - 2023.1)
    • Research on 3D computer vision, particularly the completion of CAD models with parameterizing methods.
  • MEGVII: Transformer Group, Undergraduate Intern, Advisor: Feiyang Tan (2022.7 - 2022.10)
    • Research on autonomous driving, particularly the detection of pedestrians, vehicles, and cyclists.
    • Winning 1st place in SSLAD2022 (the 2nd Workshop on Self-Supervised Learning for Next-Generation Industry-level Autonomous Driving, a workshop in ECCV 2022) Track 5 as a team.
  • Tsinghua University: BNRist, Undergraduate Researcher, Advisor: Gang Zhang, Xiaolin Hu (2022.7 - 2022.9)
    • Research on object detection problem in autonomous driving.
  • Tsinghua University: i-VisionGroup, Undergraduate Researcher, Advisor: Yueqi Duan (2022.3 - 2022.10)
    • Research on 3D computer vision, particularly detection and segmentation tasks of indoor scenes. Leveraging point cloud feature extraction backbones and transformers.