学术机构

首页 > 学术机构 > 工程系 > 师资队伍 > 副教授 > 正文

工程系

副教授

SMBU

杨硕

发布时间:2025-02-18    阅读次数:


Shuo YANG

Tenure- track Associate Professor of AI at SMBU


杨硕

深圳北理莫斯科大学工程系 预聘副教授


Dr. SHUO YANG received his Master's degree in computer science from the Institute of Software, Chinese Academy of Sciences in 2017, and his PhD degree in computer science from Beijing Institute of Technology in 2024.  He worked as an algorithm engineer at JD Finance from 2017 to 2018, and visited Megvii Research Institute from 2018 to 2019. He now primarily works on computer vision and embodied AI, with a focus on complex scene understanding and perception, decision-making, and interaction for intelligent agents such as robotic manipulators and humanoid robots. His research emphasizes multimodal perception and interaction, as well as vision-language-action (VLA) modeling. He has published 15 papers in leading international journals and conferences, including 11 as first or corresponding author. His work has appeared in IEEE T-PAMI, IEEE T-IP, IEEE T-MM, Pattern Recognition, as well as top-tier conferences such as CVPR, ICCV, AAAI, ACM MM, and IJCAI. He is a member of the CSIG Technical Committee on Multimedia, 3D vision, and the CSIG Guangdong Young Professionals Committee. He leads a key project funded by the Guangdong Provincial Department of Education on artificial intelligence (intelligent robotics) and participates in a Shenzhen Key Research Fund project.


杨硕博士于2024年在北京理工大学获得计算机科学与技术专业工学博士学位,2017年在中国科学院软件研究所获得计算机科学与技术专业工学硕士学位。2017年至2018年在京东金融任算法工程师2018年至2019年在旷视研究院交流访问。主要从事计算机视觉和具身智能(Embodied AI)研究,聚焦复杂场景理解、机械臂与人形机器人等智能体的感知、决策、与交互问题,重点探索多模态感知和交互、视觉-语言-动作建模(VLA)等关键技术。在国际权威期刊和会议发表学术论文15篇,其中以第一作者或通讯作者身份发表11篇,相关成果发表于IEEE T-PAMI、IEEE T-IP、IEEE T-MM、Pattern Recognition,以及ICML、CVPR、ICCV、AAAI、ACM MM、IJCAI等国际主流期刊与会议。现任CSIG多媒体专委会委员、三维视觉专委会委员CSIG广东省青年工作委员会委员,主持广东省教育厅人工智能(智能机器人)重点领域专项项目1项,并参与深圳市重点基金项目1项。


邮箱(Email): yangshuo@smbu.edu.cn

主页(webpage):https://shuoyang129.github.io/


Selected papers

[1] Rongjiang Zhu*, Wei Kang*,  Zeqi Liu, Junyu Chen, Shuo Yang†, Xinxiao Wu†, AmbiRefer3D: 3D Visual Grounding with Referential Ambiguity, International Conference on Machine Learning (ICML), 2026. (通讯作者,CCF-A会议)

[2] Shuo Yang†, Zirui Shang, Yongqi Wang, Derong Deng, Hongwei Chen, Xinxiao Wu, Qiyuan Cheng, Image-free Multi-label Image Recognition via LLM-powered Hierarchical Prompt Tuning, Pattern Recognition (PR), 2026.(中科院一区期刊)

[3] Yongqi Wang, Xinxiao Wu, Shuo Yang, Jiebo Luo, End-to-end Open-vocabulary Video Visual Relationship Detection using Multi-modal Prompting, IEEE Transactions on Pattern Analysis and Machine Intelligence (T-PAMI), 2025.(CCF-A,中科院一区期刊)

[4] Mengxiao Tian, Xinxiao Wu, Shuo Yang†, LLM-enhanced Action-aware Multi-modal Prompt Tuning for Image-Text Matching, International Conference on Computer Vision (ICCV), 2025.(通讯作者,CCF-A会议)

[5] Yongqi Wang, Xinxiao Wu, Shuo Yang†, METOR: A Unified Framework for Mutual Enhancement of Objects and Relationships in Open-vocabulary Video Visual Relationship Detection, International Joint Conference on Artificial Intelligence (IJCAI), 2025.(通讯作者,CCF-A会议)

[6] Zirui Shang, Yubo Zhu, Hongxi Li, Shuo Yang, Xinxiao Wu, Video Summarization using Denoising Diffusion Probabilistic Model, AAAI Conference on Artificial Intelligence (AAAI), 2025.(CCF-A会议)

[7] Shuo Yang, Xinxiao Wu, Zirui Shang, Jiebo Luo, Dynamic Pathway for Query-aware Feature Learning in Language-driven Action Localization, IEEE Transactions on Multimedia (T-MM), 2024.(中科院一区期刊)

[8] Shuo Yang*, Yongqi Wang*, Xiaofeng Ji, Xinxiao Wu, Multi-Modal Prompting for Open-Vocabulary Video Visual Relationship Detection, AAAI Conference on Artificial Intelligence (AAAI), 2024.(CCF-A会议)

[9] Shuo Yang, Zirui Shang, Xinxiao Wu, Probability Distribution Based Frame-supervised Language-driven Action Localization, ACM International Conference on Multimedia (ACM MM), 2023.(CCF-A会议)

[10] Shuo Yang, Xinxiao Wu, Entity-aware and Motion-aware Transformers for Language-driven Action Localization, International Joint Conference on Artificial Intelligence (IJCAI), 2022.(CCF-A会议)

[11] Guan'an Wang*, Shuo Yang*, Huanyu Liu, et al., High-order Information Matters: Learning Relation and Topology for Occluded Person Re-identification, IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2020.(共同一作,CCF-A会议)



关闭

地址:深圳市龙岗区大运新城国际大学园路1号

电话:0755-28323024

邮箱:info@smbu.edu.cn

深圳北理莫斯科大学版权所有 - 粤ICP备16056390号 - 粤公网安备44030702002529号

返回顶部