1 code implementation • 20 Apr 2023 • Jun Zhu, Jiandong Jin, Zihan Yang, Xiaohao Wu, Xiao Wang
The averaged visual tokens and text tokens are concatenated and fed into a fusion Transformer for multi-modal interactive learning.
Attribute Pedestrian Attribute Recognition +1