TASK	DATASET	MODEL	METRIC NAME	METRIC VALUE	GLOBAL RANK
Zero-Shot Human-Object Interaction Detection	HICO-DET	DiffHOI	mAP (UC)	36.16	# 2
Zero-Shot Human-Object Interaction Detection	HICO-DET	DiffHOI	mAP (UO)	30.11	# 1
Zero-Shot Human-Object Interaction Detection	HICO-DET	DiffHOI	mAP (UA)	-	# 2
Human-Object Interaction Detection	HICO-DET	DiffHOI	mAP	41.50	# 4
Human-Object Interaction Detection	V-COCO	DiffHOI	AP(S1)	65.7	# 4
Human-Object Interaction Detection	V-COCO	DiffHOI	AP(S2)	68.2	# 4

Badge	Markdown
	`[![PWC](https://img.shields.io/endpoint.svg?url=https://paperswithcode.com/badge/boosting-human-object-interaction-detection/zero-shot-human-object-interaction-detection)](https://paperswithcode.com/sota/zero-shot-human-object-interaction-detection?p=boosting-human-object-interaction-detection)`
	`[![PWC](https://img.shields.io/endpoint.svg?url=https://paperswithcode.com/badge/boosting-human-object-interaction-detection/human-object-interaction-detection-on-hico)](https://paperswithcode.com/sota/human-object-interaction-detection-on-hico?p=boosting-human-object-interaction-detection)`
	`[![PWC](https://img.shields.io/endpoint.svg?url=https://paperswithcode.com/badge/boosting-human-object-interaction-detection/human-object-interaction-detection-on-v-coco)](https://paperswithcode.com/sota/human-object-interaction-detection-on-v-coco?p=boosting-human-object-interaction-detection)`

Boosting Human-Object Interaction Detection with Text-to-Image Diffusion Model

20 May 2023 · Jie Yang, Bingliang Li, Fengyu Yang, Ailing Zeng, Lei Zhang, Ruimao Zhang ·

This paper investigates the problem of the current HOI detection methods and introduces DiffHOI, a novel HOI detection scheme grounded on a pre-trained text-image diffusion model, which enhances the detector's performance via improved data diversity and HOI representation. We demonstrate that the internal representation space of a frozen text-to-image diffusion model is highly relevant to verb concepts and their corresponding context. Accordingly, we propose an adapter-style tuning method to extract the various semantic associated representation from a frozen diffusion model and CLIP model to enhance the human and object representations from the pre-trained detector, further reducing the ambiguity in interaction prediction. Moreover, to fill in the gaps of HOI datasets, we propose SynHOI, a class-balance, large-scale, and high-diversity synthetic dataset containing over 140K HOI images with fully triplet annotations. It is built using an automatic and scalable pipeline designed to scale up the generation of diverse and high-precision HOI-annotated data. SynHOI could effectively relieve the long-tail issue in existing datasets and facilitate learning interaction representations. Extensive experiments demonstrate that DiffHOI significantly outperforms the state-of-the-art in regular detection (i.e., 41.50 mAP) and zero-shot detection. Furthermore, SynHOI can improve the performance of model-agnostic and backbone-agnostic HOI detection, particularly exhibiting an outstanding 11.55% mAP improvement in rare classes.

PDF Abstract

Code

Add Remove Mark official

IDEA-Research/DiffHOI official

Tasks

Add Remove

Human-Object Interaction Detection

Zero-Shot Human-Object Interaction Detection

Datasets

MS COCO

HICO-DET

V-COCO

Results from the Paper

Edit

Ranked #2 on Zero-Shot Human-Object Interaction Detection on HICO-DET (using extra training data)

Get a GitHub badge

Task	Dataset	Model	Metric Name	Metric Value	Global Rank	Benchmark
Zero-Shot Human-Object Interaction Detection	HICO-DET	DiffHOI	mAP (UC)	36.16	# 2	Compare
			mAP (UO)	30.11	# 1	Compare
			mAP (UA)	-	# 2	Compare
Human-Object Interaction Detection	HICO-DET	DiffHOI	mAP	41.50	# 4	Compare
Human-Object Interaction Detection	V-COCO	DiffHOI	AP(S1)	65.7	# 4	Compare
Human-Object Interaction Detection	V-COCO	DiffHOI	AP(S2)	68.2	# 4	Compare

Methods

Add Remove

CLIP • Diffusion

Edit Social Preview

Boosting Human-Object Interaction Detection with Text-to-Image Diffusion Model

Code Edit Add Remove Mark official

Tasks Edit Add Remove

Datasets Edit

Results from the Paper Edit

Methods Edit Add Remove

Code

Add Remove Mark official

Tasks

Add Remove

Datasets

Results from the Paper

Edit

Methods

Add Remove