2026
6 citations
VLA model, Model Merging, Multi-task Learning
CORE A*
Y Fu, Zhizhen Zhang, Y Zhang, Z Wang, Z Huang, Y Luo
Co-first author.
IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2026
Explores cross-skill model merging toward more general vision-language-action behavior.
2026
New entry
VLA model, Mobile Manipulation, Self-Correction
J S Lim, Zhizhen Zhang, P Bohm, B Tidd, Z Huang, Y Luo
arXiv preprint arXiv:2604.01567, 2026
Introduces an anchored diffusion VLA policy for efficient end-to-end mobile manipulation with low-latency closed-loop control.
2025
2 citations
Vision-Language Pre-training for Robotics, Self-supervised Learning
CORE A*
Zhizhen Zhang, L Zhu, Z Fang, Z Huang, Y Luo
Annual Conference on Neural Information Processing Systems (NeurIPS), 2025
Studies temporal ordering and continuity objectives for stronger embodied agent pretraining.
2024
New entry
Video Understanding
Resolving Spurious Temporal Location Dependency for Video Corpus Moment Retrieval
Y Zhang, L Zhang, Zhizhen Zhang, Z Wang, X Xie, W Wang
2024 IEEE International Conference on Systems, Man, and Cybernetics (SMC), 2024
Addresses temporal shortcut behavior in moment retrieval over video corpora.
2024
3 citations
Networked Intelligence
CORE B
Hiercas: Hierarchical Temporal Graph Attention Networks for Popularity Prediction in Information Cascades
Zhizhen Zhang, X Xie, Y Zhang, L Zhang, Y Jiang
2024 International Joint Conference on Neural Networks (IJCNN), 2024
Uses hierarchical temporal graph attention to reason over information cascades.
2024
4 citations
Networked Intelligence
Contrastive Learning for Implicit Social Factors in Social Media Popularity Prediction
Zhizhen Zhang, R Qiu, X Xie
arXiv preprint arXiv:2410.09345, 2024
Models hidden social factors that influence popularity prediction in social platforms.
2023
11 citations
Video Understanding
CORE A*
View while Moving: Efficient Video Recognition in Long-Untrimmed Videos
Y Tian, M Yang, L Zhang, Zhizhen Zhang, Y Liu, X Xie, X Que, W Wang
Proceedings of the 31st ACM International Conference on Multimedia, 2023
Focuses on efficient recognition strategies for long and untrimmed video streams.
2022
32 citations
Systems
Edge-Assisted Real-Time Video Analytics with Spatial-Temporal Redundancy Suppression
Z Wang, X He, Zhizhen Zhang, Y Zhang, Z Cao, W Cheng, W Wang, Y Cui
IEEE Internet of Things Journal 10 (7), 6324-6335, 2022
Proposes real-time video analytics with redundancy suppression in edge-assisted settings.